Lead Software Engineer, AI Platform Reliability
$61k - $101kJ.P. Morgan
Salary: $61,000 - 101,000 per year Requirements:
- We require formal training or certification in software engineering concepts, plus 5+ years of applied experience.
- We need strong hands-on Python development experience, with a background in delivering production-grade services.
- We value demonstrated experience leading the effective use of approved AI-assisted software development tools, with the ability to set expectations for verifying correctness, performance, and security.
- We look for a strong understanding of responsible AI practices in engineering workflows, including data sensitivity, secure handling of inputs and outputs, and resiliency and security standards.
- We expect practical experience with system design, automated testing, debugging, and maintaining operational stability in production software.
- We require experience implementing observability, logging, metrics, alerts, service level objectives, incident response, and root-cause analysis for production services.
- We seek working knowledge of software development and technical processes, with depth in one or more areas such as cloud platforms, AI, machine learning platforms, distributed systems, or infrastructure engineering.
- We need the ability to break down technical requirements into actionable engineering tasks, manage dependencies, and deliver against milestones alongside product and application teams.
- We require strong written and verbal communication skills to explain technical decisions, trade-offs, issues, and risks to engineering teams and stakeholders.
- Preferred experience includes supporting AI/ML or generative AI platform capabilities such as model hosting, inference services, model gateways, managed AI services, or developer-facing AI/ML infrastructure.
- Preferred experience includes managing AI infrastructure on cloud platforms, including deployment, scaling, monitoring, and optimization of machine learning workloads.
- Preferred experience includes building reusable golden path assets such as templates, reference implementations, SDKs, automated tests, onboarding guides, and deployment patterns.
- Preferred experience includes developing generative AI applications or AI agents, and/or implementing AI-assisted operations with appropriate guardrails.
- We design and implement solutions that improve the reliability and scalability of AI/ML platforms and applications for rapidly growing demand.
- We develop secure, stable, high-quality production code and contribute to code reviews, debugging, testing, and defect remediation across AI Foundation Services components.
- We build and improve reusable platform services, APIs, SDKs, and libraries that standardize how application teams use model hosting, inference, and managed AI/ML services.
- We partner with application teams to implement AI Foundation Services capabilities that enable generative AI and other AI use cases, supporting delivery from technical design through build, launch, and early operational support.
- We own and evolve non-functional requirements and build tooling for observability, resilience, security controls, infrastructure management, and cost optimization.
- We establish and enforce standards and reference architectures for reliability, observability, automation, and operational readiness across services.
- We work with product and platform engineering teams to define and meet service reliability targets, including performance, availability, and recoverability.
- We take part in on-call rotations, troubleshoot and resolve complex production issues, identify systemic gaps, and drive lasting remediation.
- We mentor and guide engineers while raising the bar for engineering quality, documentation, and operational discipline.
- AI
- AI Agents
- Cloud
- Support
- Machine Learning
- Marketing
- Python
- Security
More:
hackajob is partnering with JPMorganChase to connect exceptional professionals with this opportunity. We are a Lead Software Engineer role within our AI/ML Data Platforms organization, on the Reliability Engineering team, where we focus on building trusted, market-leading technology products that improve the reliability and scalability of AI platforms. Our company has a history spanning over 200 years and serves consumers, small businesses, and major corporate, institutional, and government clients under the J.P. Morgan and Chase brands. We offer a competitive total rewards package, including base salary tied to role, experience, skill set, and location, with possible commission or discretionary incentive compensation for eligible roles. Our benefits and programs may include comprehensive health care, wellness centers, retirement savings, backup childcare, tuition reimbursement, mental health support, and financial coaching. We value diversity and inclusion, provide reasonable accommodations where needed, and operate with a global workforce across Corporate Functions that supports finance, risk, human resources, marketing, and other essential business areas.
last updated 40 week of 2026
$100k - $150k
...Platform Reliability Engineer – Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States... ...Demonstrated experience leading incident response and conducting...SuggestedFull timeH1bLocal areaImmediate startRemote workVisa sponsorship$61k - $101k
...formal training or certification in software engineering concepts, along with 5+ years... ...demonstrated experience leading the effective use of enterprise-approved AI-assisted software development tools... ...high-performance LLM inference platform using vLLM and GPU/CPU serving...SuggestedFull timeFor contractors$20k
...enabling human life on Mars.SITE RELIABILITY ENGINEER (STARSHIELD) At SpaceX we’re... ...within minutes, and the software that brings it all together.... ...Operations, and GPU platforms. You will develop automation... ...and productize solutions for AI clusters (100k+ GPU scale)Develop...SuggestedPermanent employmentTemporary workImmediate startWeekend work$175k - $250k
...Senior Cloud Infrastructure Engineer Location: San Francisco... ...with generative AI. They are the team behind... ...complete flexibility. Their platform allows users to connect... ..., you will take the lead on designing, deploying... ..., performance, and reliability across environments. What...SuggestedFull timeRemote workRelocationRelocation package- ...Site Reliability Engineer III (AI Platform) Location: Mount Laurel, NJ (Onsite) Duration: Contract Experience: 4+ years About the Role... .... ~ Strong understanding of distributed systems, software design, algorithms, and system performance. ~ Experience...SuggestedContract work
- ...Database, Infra (Container platforms) and Network )... ...Expected to actively lead and triage proactively... ...alerts for “Senior Site Reliability Engineer” roles. Bellevue, WA... ...Diego, CA) Senior Software Engineer - Optical Network... ...Reliability Engineer, AI/ML Platforms Senior...Contract workRemote work
$166k - $220k
...powered by Lattice OS, an AI‑powered operating... ...combinations of hardware and software tailored to different... ...systems integration engineers internalize the... ...are looking for a Site Reliability Engineer (SRE) to join... ...They are comfortable leading large, focused projects...Full timeWork experience placement- ...seeking a high-caliber Site Reliability Engineer (SRE) to join our Forward... ...that our complex, data-driven AI platforms remain resilient, scalable,... .... This role is a hybrid of software engineering and systems... ...responder in on-call rotations, leading the technical resolution of...Local area
- ...intelligent, cloud identity platform lets people shop, work... ..., performance, and reliability—across our... ...maintain production-quality software—with tests, documentation... ...automates reliability engineering work: Deployment tooling... ...building with AI coding assistants or agentic...Full timeCasual workLocal areaWorldwideFlexible hoursShift work
- ...client seeks a Sr. Site Reliability Engineer III to design,... ...Kubernetes or VMWare platforms, CI/CD enablement, observability... ...friction across the software development lifecycle.... ...level objectives. Lead or participate in incident... ...intelligence (AI) tools as part of its...Hourly payPermanent employmentFull timeLocal areaImmediate start
$95k - $171k
...passionate about cutting-edge AI infrastructure? Do... ...of the most exciting platforms in cloud computing?... ..., and ensuring reliability for AI workloads within... ...As an Site Reliability Engineer II, you will be responsible... ...protects life online. Leading companies worldwide...Permanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours$168k - $200k
...Datavant is the data collaboration platform trusted for healthcare. Guided by our mission... ...their medical records to powering the AI revolution in healthcare, Datavanters... ...For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You...Remote work$121.4k - $218.6k
...Join our highly skilled Site Reliability Engineering team! Our team designs,... ...Investigating and analyzing platform components to identify opportunities... ...without a glitch. AI : Enabling our customers to... ...). Akamai provides industry-leading benefits including healthcare...Work experience placementWork at office$81.1k - $187k
...We are looking for a Site Reliability Engineer 3 to support mission-critical... ...capacity needs. Collaborates with software development teams to develop... ...life-saving care. And with AI embedded across our products... ...your potential at a company leading the way in AI and cloud...Temporary workImmediate startFlexible hoursShift work$75.7k - $136.3k
...our highly skilled Site Reliability Engineering team! Our team... ...Collaborating with software engineering, infrastructure, and platform teams to investigate complex... ...moments without a glitch. AI : Enabling our... ...Akamai provides industry-leading benefits including healthcare...Work experience placementWork at office$107k - $220k
...The Site Reliability Engineer (SRE) will ensure the reliability, performance, and scalability of the WDP System. This person will define and track... ...due to a disability. Please send your request to ****@*****.***.ai. Avalore is an Equal Opportunity Employer - We not...Full timeContract workTemporary workWork at officeVisa sponsorshipWork visa$115.5k - $184.8k
...Site Reliability Engineer II Join Axon and be a Force for Good. At Axon... ...ecosystem of devices and cloud software. Like our products, we work... ...court that's us. Axon's Platform team is the engine behind... ...~ Familiarity with agentic AI tooling or LLM-powered developer...Work at officeRemote work$136.2k - $214.01k
...cybersecurity. We protect how people, data, and AI agents connect across email, cloud, and... ...impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep... ...stakeholders, network, hardware and software that relate to scaling and...Full timeFlexible hours$169.3k - $304.7k
...efficient, scalable, and reliable routing software and infrastructure... ...stability of our global platform. As a Principal Site Reliability Engineer - Network, you will be... ...without a glitch. AI: Enabling our customers... ...provides industry-leading benefits including healthcare...Work experience placementWork at office- ...Manager, Site Reliability Engineering Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires...
$140k - $210k
...26) Day to Day As an Engineering Manager in Site Reliability Engineering at Indeed, you... ...grow a team that applies software engineering principles to... ...infrastructure and observability tools/platforms, such AWS, Kubernetes,... ...for that opening. AI Notice Indeed is...Work experience placementLocal area$84.9k - $209.5k
...professionals can reliably access the applications... ...Site Reliability Engineer to strengthen the... ...of our Citrix platform. The platform delivers... ...expertise with software engineering practices... ...and capacity. Lead complex... ...saving care. And with AI embedded across our...Temporary workImmediate startFlexible hoursShift work- ...Overview Sedaro is hiring a Lead Software Engineer to build powerful tools for... ...the results. ~ Team: Platform Services ~ Location: 80... ...the Space Force and Shield AI to deliver next-gen systems... ...frontend performance and reliability issues as our data scale and...Permanent employmentFlexible hours3 days per week
- ...technology products. As a Sr. Lead Software Engineer at JPMorgan Chase within... ...– Infrastructure Platforms Group, you are an integral... ...tools to ensure infrastructure reliability Develop and maintain network... ...and governance of approved AI-assisted engineering practices...
$61k - $101k
...training or certification in software engineering, along with 5+ years of... ...development, testing, and operational reliability. We need demonstrated... ...using enterprise-approved AI-assisted software development... ...decision-makers to adopt leading-edge technologies where appropriate...Full timeFor contractors$207k - $284.9k
...Secure Every Identity, from AI to Human Identity is the... ...talk. Senior Manager, Site Reliability Engineering Secure Every Identity, from... ...Okta's federal environments, lead a team of engineers working... ...Strong background in SRE or platform engineering fundamentals: Kubernetes...Permanent employmentLocal areaWorldwideFlexible hoursDay shift$174k - $239k
...Every Identity, from AI to Human Identity is... ...infrastructure to enterprise platforms, we partner across... ...to drive scale, reliability, and innovation through... ...Staff Site Reliability Engineer Opportunity Okta Federal... ...leveraging secure software development practices....Work experience placementLocal areaWorldwideFlexible hours$174k - $238k
...Every Identity, from AI to Human Identity is... ...experienced Staff Site Reliability Engineer to join Okta's Federal... ...invest in platform engineering, observability... ...partnering closely with software engineers, architects,... ...customer-facing systems. Lead incident response...Local areaWorldwideFlexible hours$204k - $306k
...Secure Every Identity, from AI to Human Identity is the... ...let's talk. Manager, Site Reliability Engineering San Francisco, California... ...the Manager of Infrastructure Platform and Shared Services, you will... ...robust self-healing patterns. Lead, mentor, and grow a high-...Permanent employmentWork at officeLocal areaWorldwideFlexible hours2 days per week$194k - $267k
...Secure Every Identity, from AI to Human Identity is the key to unlocking the potential... ...and tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and...Permanent employmentWork at officeLocal areaWorldwideFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Software Engineer, AI Platform Reliability. Be the first to apply!
- lead operating engineer Washington DC
- lead engineer Washington DC
- client platform engineer Washington DC
- data platform engineer Washington DC
- senior platform engineer Washington DC
- platform engineer Washington DC
- platform developer Washington DC
- software sales executive Washington DC
- software technology Washington DC
- software intern Washington DC



