Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Software Engineer, AI Platform Reliability

$61k - $101k
Full-time

J.P. Morgan

Salary: $61,000 - 101,000 per year Requirements:

  • We require formal training or certification in software engineering concepts, plus 5+ years of applied experience.
  • We need strong hands-on Python development experience, with a background in delivering production-grade services.
  • We value demonstrated experience leading the effective use of approved AI-assisted software development tools, with the ability to set expectations for verifying correctness, performance, and security.
  • We look for a strong understanding of responsible AI practices in engineering workflows, including data sensitivity, secure handling of inputs and outputs, and resiliency and security standards.
  • We expect practical experience with system design, automated testing, debugging, and maintaining operational stability in production software.
  • We require experience implementing observability, logging, metrics, alerts, service level objectives, incident response, and root-cause analysis for production services.
  • We seek working knowledge of software development and technical processes, with depth in one or more areas such as cloud platforms, AI, machine learning platforms, distributed systems, or infrastructure engineering.
  • We need the ability to break down technical requirements into actionable engineering tasks, manage dependencies, and deliver against milestones alongside product and application teams.
  • We require strong written and verbal communication skills to explain technical decisions, trade-offs, issues, and risks to engineering teams and stakeholders.
  • Preferred experience includes supporting AI/ML or generative AI platform capabilities such as model hosting, inference services, model gateways, managed AI services, or developer-facing AI/ML infrastructure.
  • Preferred experience includes managing AI infrastructure on cloud platforms, including deployment, scaling, monitoring, and optimization of machine learning workloads.
  • Preferred experience includes building reusable golden path assets such as templates, reference implementations, SDKs, automated tests, onboarding guides, and deployment patterns.
  • Preferred experience includes developing generative AI applications or AI agents, and/or implementing AI-assisted operations with appropriate guardrails.
Responsibilities:
  • We design and implement solutions that improve the reliability and scalability of AI/ML platforms and applications for rapidly growing demand.
  • We develop secure, stable, high-quality production code and contribute to code reviews, debugging, testing, and defect remediation across AI Foundation Services components.
  • We build and improve reusable platform services, APIs, SDKs, and libraries that standardize how application teams use model hosting, inference, and managed AI/ML services.
  • We partner with application teams to implement AI Foundation Services capabilities that enable generative AI and other AI use cases, supporting delivery from technical design through build, launch, and early operational support.
  • We own and evolve non-functional requirements and build tooling for observability, resilience, security controls, infrastructure management, and cost optimization.
  • We establish and enforce standards and reference architectures for reliability, observability, automation, and operational readiness across services.
  • We work with product and platform engineering teams to define and meet service reliability targets, including performance, availability, and recoverability.
  • We take part in on-call rotations, troubleshoot and resolve complex production issues, identify systemic gaps, and drive lasting remediation.
  • We mentor and guide engineers while raising the bar for engineering quality, documentation, and operational discipline.
Technologies:
  • AI
  • AI Agents
  • Cloud
  • Support
  • Machine Learning
  • Marketing
  • Python
  • Security

More:

hackajob is partnering with JPMorganChase to connect exceptional professionals with this opportunity. We are a Lead Software Engineer role within our AI/ML Data Platforms organization, on the Reliability Engineering team, where we focus on building trusted, market-leading technology products that improve the reliability and scalability of AI platforms. Our company has a history spanning over 200 years and serves consumers, small businesses, and major corporate, institutional, and government clients under the J.P. Morgan and Chase brands. We offer a competitive total rewards package, including base salary tied to role, experience, skill set, and location, with possible commission or discretionary incentive compensation for eligible roles. Our benefits and programs may include comprehensive health care, wellness centers, retirement savings, backup childcare, tuition reimbursement, mental health support, and financial coaching. We value diversity and inclusion, provide reasonable accommodations where needed, and operate with a global workforce across Corporate Functions that supports finance, risk, human resources, marketing, and other essential business areas.

last updated 40 week of 2026

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Lead Software Engineer, AI Platform Reliability in Washington DC vacancy
  • $100k - $150k

     ...Platform Reliability Engineer – Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States...  ...Demonstrated experience leading incident response and conducting... 
    Suggested
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    Washington DC
    more than 2 months ago
  • $61k - $101k

     ...formal training or certification in software engineering concepts, along with 5+ years...  ...demonstrated experience leading the effective use of enterprise-approved AI-assisted software development tools...  ...high-performance LLM inference platform using vLLM and GPU/CPU serving... 
    Suggested
    Full time
    For contractors

    J.P. Morgan

    Washington DC
    9 days ago
  • $20k

     ...enabling human life on Mars.SITE RELIABILITY ENGINEER (STARSHIELD) At SpaceX we’re...  ...within minutes, and the software that brings it all together....  ...Operations, and GPU platforms. You will develop automation...  ...and productize solutions for AI clusters (100k+ GPU scale)Develop... 
    Suggested
    Permanent employment
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Washington DC
    2 days ago
  • $175k - $250k

     ...Senior Cloud Infrastructure Engineer Location: San Francisco...  ...with generative AI. They are the team behind...  ...complete flexibility. Their platform allows users to connect...  ..., you will take the lead on designing, deploying...  ..., performance, and reliability across environments. What... 
    Suggested
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    Washington DC
    3 days ago
  •  ...Site Reliability Engineer III (AI Platform) Location: Mount Laurel, NJ (Onsite) Duration: Contract Experience: 4+ years About the Role...  .... ~ Strong understanding of distributed systems, software design, algorithms, and system performance. ~ Experience... 
    Suggested
    Contract work

    GCS Recruitment

    Laurel, MD
    2 days ago
  •  ...Database, Infra (Container platforms) and Network )...  ...Expected to actively lead and triage proactively...  ...alerts for “Senior Site Reliability Engineer” roles. Bellevue, WA...  ...Diego, CA) Senior Software Engineer - Optical Network...  ...Reliability Engineer, AI/ML Platforms Senior... 
    Contract work
    Remote work

    Signature IT World Inc

    Washington DC
    3 days ago
  • $166k - $220k

     ...powered by Lattice OS, an AI‑powered operating...  ...combinations of hardware and software tailored to different...  ...systems integration engineers internalize the...  ...are looking for a Site Reliability Engineer (SRE) to join...  ...They are comfortable leading large, focused projects... 
    Full time
    Work experience placement

    Mosaic

    Washington DC
    2 days ago
  •  ...seeking a high-caliber Site Reliability Engineer (SRE) to join our Forward...  ...that our complex, data-driven AI platforms remain resilient, scalable,...  .... This role is a hybrid of software engineering and systems...  ...responder in on-call rotations, leading the technical resolution of... 
    Local area

    Tiger Analytics Inc.

    Washington DC
    a month ago
  •  ...intelligent, cloud identity platform lets people shop, work...  ..., performance, and reliability—across our...  ...maintain production-quality software—with tests, documentation...  ...automates reliability engineering work: Deployment tooling...  ...building with AI coding assistants or agentic... 
    Full time
    Casual work
    Local area
    Worldwide
    Flexible hours
    Shift work

    Ping Identity

    Washington DC
    4 days ago
  •  ...client seeks a Sr. Site Reliability Engineer III to design,...  ...Kubernetes or VMWare platforms, CI/CD enablement, observability...  ...friction across the software development lifecycle....  ...level objectives. Lead or participate in incident...  ...intelligence (AI) tools as part of its... 
    Hourly pay
    Permanent employment
    Full time
    Local area
    Immediate start

    Eliassen Group

    Washington DC
    a month ago
  • $95k - $171k

     ...passionate about cutting-edge AI infrastructure? Do...  ...of the most exciting platforms in cloud computing?...  ..., and ensuring reliability for AI workloads within...  ...As an Site Reliability Engineer II, you will be responsible...  ...protects life online. Leading companies worldwide... 
    Permanent employment
    Work experience placement
    Work at office
    Remote work
    Work from home
    Worldwide
    Flexible hours

    Akamai

    Washington DC
    2 days ago
  • $168k - $200k

     ...Datavant is the data collaboration platform trusted for healthcare. Guided by our mission...  ...their medical records to powering the AI revolution in healthcare, Datavanters...  ...For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You... 
    Remote work

    Datavant

    Washington DC
    3 days ago
  • $121.4k - $218.6k

     ...Join our highly skilled Site Reliability Engineering team! Our team designs,...  ...Investigating and analyzing platform components to identify opportunities...  ...without a glitch. AI : Enabling our customers to...  ...). Akamai provides industry-leading benefits including healthcare... 
    Work experience placement
    Work at office

    Akamai

    Washington DC
    2 days ago
  • $81.1k - $187k

     ...We are looking for a Site Reliability Engineer 3 to support mission-critical...  ...capacity needs. Collaborates with software development teams to develop...  ...life-saving care. And with AI embedded across our products...  ...your potential at a company leading the way in AI and cloud... 
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    Washington DC
    3 days ago
  • $75.7k - $136.3k

     ...our highly skilled Site Reliability Engineering team! Our team...  ...Collaborating with software engineering, infrastructure, and platform teams to investigate complex...  ...moments without a glitch. AI : Enabling our...  ...Akamai provides industry-leading benefits including healthcare... 
    Work experience placement
    Work at office

    Akamai

    Washington DC
    2 days ago
  • $107k - $220k

     ...The Site Reliability Engineer (SRE) will ensure the reliability, performance, and scalability of the WDP System. This person will define and track...  ...due to a disability. Please send your request to ****@*****.***.ai.  Avalore is an Equal Opportunity Employer - We not... 
    Full time
    Contract work
    Temporary work
    Work at office
    Visa sponsorship
    Work visa

    Avalore, LLC

    Arlington, VA
    1 day ago
  • $115.5k - $184.8k

     ...Site Reliability Engineer II Join Axon and be a Force for Good. At Axon...  ...ecosystem of devices and cloud software. Like our products, we work...  ...court that's us. Axon's Platform team is the engine behind...  ...~ Familiarity with agentic AI tooling or LLM-powered developer... 
    Work at office
    Remote work

    Axon

    Washington DC
    4 days ago
  • $136.2k - $214.01k

     ...cybersecurity. We protect how people, data, and AI agents connect across email, cloud, and...  ...impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep...  ...stakeholders, network, hardware and software that relate to scaling and... 
    Full time
    Flexible hours

    Proofpoint

    Laurel, MD
    4 days ago
  • $169.3k - $304.7k

     ...efficient, scalable, and reliable routing software and infrastructure...  ...stability of our global platform. As a Principal Site Reliability Engineer - Network, you will be...  ...without a glitch. AI: Enabling our customers...  ...provides industry-leading benefits including healthcare... 
    Work experience placement
    Work at office

    Akamai

    Washington DC
    3 days ago
  •  ...Manager, Site Reliability Engineering Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires... 

    Okta, Inc.

    Washington DC
    2 days ago
  • $140k - $210k

     ...26) Day to Day As an Engineering Manager in Site Reliability Engineering at Indeed, you...  ...grow a team that applies software engineering principles to...  ...infrastructure and observability tools/platforms, such AWS, Kubernetes,...  ...for that opening. AI Notice Indeed is... 
    Work experience placement
    Local area

    Indeed

    Washington DC
    1 day ago
  • $84.9k - $209.5k

     ...professionals can reliably access the applications...  ...Site Reliability Engineer to strengthen the...  ...of our Citrix platform. The platform delivers...  ...expertise with software engineering practices...  ...and capacity. Lead complex...  ...saving care. And with AI embedded across our... 
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    Washington DC
    3 days ago
  •  ...Overview  Sedaro is hiring a Lead Software Engineer to build powerful tools for...  ...the results.  ~ Team: Platform Services  ~ Location: 80...  ...the Space Force and Shield AI to deliver next-gen systems...  ...frontend performance and reliability issues as our data scale and... 
    Permanent employment
    Flexible hours
    3 days per week

    Sedaro

    Arlington, VA
    a month ago
  •  ...technology products.  As a Sr. Lead Software Engineer at JPMorgan Chase within...  ...– Infrastructure Platforms Group, you are an integral...  ...tools to ensure infrastructure reliability Develop and maintain network...  ...and governance of approved AI-assisted engineering practices... 

    JPMorgan Chase & Co.

    Washington DC
    9 days ago
  • $61k - $101k

     ...training or certification in software engineering, along with 5+ years of...  ...development, testing, and operational reliability. We need demonstrated...  ...using enterprise-approved AI-assisted software development...  ...decision-makers to adopt leading-edge technologies where appropriate... 
    Full time
    For contractors

    J.P. Morgan

    Washington DC
    28 days ago
  • $207k - $284.9k

     ...Secure Every Identity, from AI to Human Identity is the...  ...talk. Senior Manager, Site Reliability Engineering Secure Every Identity, from...  ...Okta's federal environments, lead a team of engineers working...  ...Strong background in SRE or platform engineering fundamentals: Kubernetes... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours
    Day shift

    Okta

    Washington DC
    2 days ago
  • $174k - $239k

     ...Every Identity, from AI to Human Identity is...  ...infrastructure to enterprise platforms, we partner across...  ...to drive scale, reliability, and innovation through...  ...Staff Site Reliability Engineer Opportunity Okta Federal...  ...leveraging secure software development practices.... 
    Work experience placement
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    25 days ago
  • $174k - $238k

     ...Every Identity, from AI to Human Identity is...  ...experienced Staff Site Reliability Engineer to join Okta's Federal...  ...invest in platform engineering, observability...  ...partnering closely with software engineers, architects,...  ...customer-facing systems. Lead incident response... 
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    2 days ago
  • $204k - $306k

     ...Secure Every Identity, from AI to Human Identity is the...  ...let's talk. Manager, Site Reliability Engineering San Francisco, California...  ...the Manager of Infrastructure Platform and Shared Services, you will...  ...robust self-healing patterns. Lead, mentor, and grow a high-... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours
    2 days per week

    Okta

    Washington DC
    2 days ago
  • $194k - $267k

     ...Secure Every Identity, from AI to Human Identity is the key to unlocking the potential...  ...and tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Software Engineer, AI Platform Reliability. Be the first to apply!