Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Software Engineer, AI Platform Reliability

$61k - $101k
Full-time

J.P. Morgan

Salary: $61,000 - 101,000 per year Requirements:

  • We require formal training or certification in software engineering concepts, plus 5+ years of applied experience.
  • We need strong hands-on Python development experience, with a background in delivering production-grade services.
  • We value demonstrated experience leading the effective use of approved AI-assisted software development tools, with the ability to set expectations for verifying correctness, performance, and security.
  • We look for a strong understanding of responsible AI practices in engineering workflows, including data sensitivity, secure handling of inputs and outputs, and resiliency and security standards.
  • We expect practical experience with system design, automated testing, debugging, and maintaining operational stability in production software.
  • We require experience implementing observability, logging, metrics, alerts, service level objectives, incident response, and root-cause analysis for production services.
  • We seek working knowledge of software development and technical processes, with depth in one or more areas such as cloud platforms, AI, machine learning platforms, distributed systems, or infrastructure engineering.
  • We need the ability to break down technical requirements into actionable engineering tasks, manage dependencies, and deliver against milestones alongside product and application teams.
  • We require strong written and verbal communication skills to explain technical decisions, trade-offs, issues, and risks to engineering teams and stakeholders.
  • Preferred experience includes supporting AI/ML or generative AI platform capabilities such as model hosting, inference services, model gateways, managed AI services, or developer-facing AI/ML infrastructure.
  • Preferred experience includes managing AI infrastructure on cloud platforms, including deployment, scaling, monitoring, and optimization of machine learning workloads.
  • Preferred experience includes building reusable golden path assets such as templates, reference implementations, SDKs, automated tests, onboarding guides, and deployment patterns.
  • Preferred experience includes developing generative AI applications or AI agents, and/or implementing AI-assisted operations with appropriate guardrails.
Responsibilities:
  • We design and implement solutions that improve the reliability and scalability of AI/ML platforms and applications for rapidly growing demand.
  • We develop secure, stable, high-quality production code and contribute to code reviews, debugging, testing, and defect remediation across AI Foundation Services components.
  • We build and improve reusable platform services, APIs, SDKs, and libraries that standardize how application teams use model hosting, inference, and managed AI/ML services.
  • We partner with application teams to implement AI Foundation Services capabilities that enable generative AI and other AI use cases, supporting delivery from technical design through build, launch, and early operational support.
  • We own and evolve non-functional requirements and build tooling for observability, resilience, security controls, infrastructure management, and cost optimization.
  • We establish and enforce standards and reference architectures for reliability, observability, automation, and operational readiness across services.
  • We work with product and platform engineering teams to define and meet service reliability targets, including performance, availability, and recoverability.
  • We take part in on-call rotations, troubleshoot and resolve complex production issues, identify systemic gaps, and drive lasting remediation.
  • We mentor and guide engineers while raising the bar for engineering quality, documentation, and operational discipline.
Technologies:
  • AI
  • AI Agents
  • Cloud
  • Support
  • Machine Learning
  • Marketing
  • Python
  • Security

More:

hackajob is partnering with JPMorganChase to connect exceptional professionals with this opportunity. We are a Lead Software Engineer role within our AI/ML Data Platforms organization, on the Reliability Engineering team, where we focus on building trusted, market-leading technology products that improve the reliability and scalability of AI platforms. Our company has a history spanning over 200 years and serves consumers, small businesses, and major corporate, institutional, and government clients under the J.P. Morgan and Chase brands. We offer a competitive total rewards package, including base salary tied to role, experience, skill set, and location, with possible commission or discretionary incentive compensation for eligible roles. Our benefits and programs may include comprehensive health care, wellness centers, retirement savings, backup childcare, tuition reimbursement, mental health support, and financial coaching. We value diversity and inclusion, provide reasonable accommodations where needed, and operate with a global workforce across Corporate Functions that supports finance, risk, human resources, marketing, and other essential business areas.

last updated 36 week of 2026

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Lead Software Engineer, AI Platform Reliability in Washington DC vacancy
  • $117.2k - $313.7k

     ...SalesforceSalesforce is the #1 AI CRM, where humans...  ...at the company leading workforce transformation...  .../ Lead / Principal Software Engineer - Foundations Team...  ...deeply about performance, reliability, and engineering...  ...backend systems, cloud platforms, and infrastructure.Languages... 
    Suggested
    Full time
    Work experience placement
    Remote work

    Salesforce

    Washington DC
    3 days ago
  • $145k - $250k

     ...decisions, and context. As AI takes on more of that...  ...to meet you. Site Reliability Engineer / DevSecOps Engineer...  ...pipeline for an AI/ML platform deployed into tightly...  ...release artifacts with a software bill of materials...  ...and OpenTelemetry Lead incident response and... 
    Suggested
    Full time
    Contract work
    Remote work

    OpenTeams

    Washington DC
    6 days ago
  • $174k - $226k

     ...roleWe're hiring a hands-on engineering leader to own both the...  ...the delivery of our platform engineering team. This...  ...that space.We build in an AI-forward environment....  ...culture that is organized, reliable, and focused on...  ...people manager or tech lead with direct reports — including... 
    Suggested
    Work at office
    Remote work
    Flexible hours

    Koalafi

    Arlington, VA
    4 days ago
  • $132.23k - $176.31k

     ...the trusted network for the AI‑powered world, connecting...  ...skilled and proactive Senior Lead Site Reliability Engineer (SRE) to join our team,...  ...native mindset, understands the software development lifecycle (from...  ...for the Lumen Connect platform. Collaboration & Communication... 
    Suggested
    Full time
    Temporary work
    Remote work

    Lumen

    Largo, MD
    1 day ago
  • $166k - $220k

     ...powered by Lattice OS, an AI-powered operating...  ...combinations of hardware and software tailored to different...  ...systems integration engineers internalize the...  ...are looking for a Site Reliability Engineer (SRE) to join...  ...They are comfortable leading large, focused projects... 
    Suggested
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Washington DC
    2 days ago
  •  ...in others.About the Role:As a Site Reliability Engineer, you’ll join the global Platform SRE team responsible for building,...  ...and postmortem practices.Lead and contribute to scaling initiatives...  ...., a leading developer of API and AI connectivity technologies, is building... 
    Temporary work

    Kong

    Washington DC
    9 hours ago
  • $207k - $284.9k

     ...Every Identity, from AI to HumanIdentity is the...  ....Senior Manager, Site Reliability EngineeringSecure Every...  ...Federal Operations Engineering GroupOkta's Federal Operations...  ...federal environments, lead a team of engineers...  ...background in SRE or platform engineering... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours
    Day shift

    Okta

    Washington DC
    2 days ago
  • $182k - $250.8k

     ...Every Identity, from AI to HumanIdentity is the...  ...is the backbone of our platform's reliability and operational...  ...forward-thinking group of engineers and leaders who believe...  ...Reliability Engineer, you'll lead this team with a focus...  ..., resilience, and software engineering rigor into... 
    Permanent employment
    Local area
    Remote work
    Worldwide
    Flexible hours
    Weekend work
    Weekday work

    Okta

    Washington DC
    9 hours ago
  • $174k - $239k

     ...Every Identity, from AI to HumanIdentity is the...  ...infrastructure to enterprise platforms, we partner across...  ...to drive scale, reliability, and innovation through...  ...Staff Site Reliability Engineer OpportunityOkta Federal...  ...while leveraging secure software development practices.... 
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    4 days ago
  • $174k - $238k

     ...Every Identity, from AI to HumanIdentity is the...  ...experienced Staff Site Reliability Engineer to join Okta's Federal...  ...invest in platform engineering, observability...  ...partnering closely with software engineers, architects,...  ...customer-facing systems.Lead incident response efforts... 
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    3 days ago
  • $113.3k - $205.52k

     ...do at Jamf: As a Senior Site Reliability Engineer, you'll help us balance development...  ...their services are measured, lead the work to improve them, and...  ..., and network layers, using AI to correlate logs, metrics,...  ...of 5 years experience in software engineering, SRE or production... 
    Work at office
    Remote work
    Worldwide
    Flexible hours

    GrabJobs

    Alexandria, VA
    1 day ago
  •  ...light a fire within you. Lead Software Engineer The Lead Software...  ...services within the NICE CXone platform, building scalable cloud systems...  ...practices, including AI-assisted engineering tools...  ...services, balancing performance, reliability, and maintainability.... 
    Full time
    Worldwide

    Nice

    Washington DC
    9 hours ago
  •  ...Overview   Sedaro is hiring a Lead Software Engineer to build powerful tools for...  ...the results.   ~ Team: Platform Services   ~ Location:...  ...Space Force and Shield AI to deliver next-gen...  ...address frontend performance and reliability issues as our data scale... 
    Permanent employment
    Full time
    Flexible hours
    3 days per week

    Sedaro

    Arlington, VA
    9 hours ago
  • $61k - $101k

     ...formal training or certification in software engineering concepts, along with 5+ years...  ...demonstrated experience leading the effective use of enterprise-approved AI-assisted software development tools...  ...high-performance LLM inference platform using vLLM and GPU/CPU serving... 
    Full time
    For contractors

    J.P. Morgan

    Washington DC
    3 days ago
  •  ...-notch technology products. As a Senior Lead Software Engineer at JPMorgan Chase within the Corporate Sector, Infrastructure Platforms team, you are an integral part of an agile...  ...secure, scalable cloud platforms optimized for AI/ML workloads. Partner with AI teams to... 
    For contractors

    JPMorgan Chase & Co.

    Washington DC
    7 days ago
  •  ...Role Wine & More is seeking a Lead Platform Engineer to join our Technology team...  ...strong focus on security, reliability, and cost efficiency—...  ...Node.js, while leading our AI initiative by identifying practical...  .... Work in platform software development using Go, Node.... 
    Work at office

    Total Wine & More

    Bethesda, MD
    5 days ago
  •  ...client seeks a Sr. Site Reliability Engineer III to design,...  ...Kubernetes or VMWare platforms, CI/CD enablement, observability...  ...friction across the software development lifecycle....  ...level objectives. Lead or participate in incident...  ...intelligence (AI) tools as part of its... 
    Hourly pay
    Permanent employment
    Full time
    Local area
    Immediate start

    Eliassen Group

    Washington DC
    13 days ago
  • $165k - $225.6k

     ...Every Identity, from AI to Human Identity is...  ...infrastructure to enterprise platforms, we partner across...  ...to drive scale, reliability, and innovation through...  ...Senior Site Reliability Engineer Opportunity Reporting...  ...: Partner with software engineering teams to champion... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    1 day ago
  •  ...Position: Senior Site Reliability Engineer (SRE) Location:...  ...-based production platforms. This role requires a...  ...cloud architecture, software engineering, platform...  ..., observability, and AI-driven engineering practices...  .... Experience leading technical execution... 
    Full time

    Siri InfoSolutions Inc

    Washington DC
    1 day ago
  •  ...Infrastructure Experts to evaluate AI-powered workflows across software development, cloud infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,...  ...deployment workflows for accuracy and reliability. Work with AWS, Azure, GCP, Kubernetes... 
    For contractors
    Remote work

    YO AI Labs

    Washington DC
    12 days ago
  • $128k - $252.5k

     ...analytics, Generative AI, transformative...  ...creatives, designers, engineers, and architects. Our...  ...Work You’ll Do As a Lead Agentic Software Engineer II, you will...  ...using Python, SQL, cloud platforms, and modern data tools...  ...engineering, DevOps, site reliability engineering, quality... 
    Local area

    Deloitte

    Rosslyn, VA
    9 hours ago
  • $160k - $210k

     ...our Deep Learning Advertising Platform. Since 2015, we have...  ...will be at the forefront of AI-driven advertising solutions...  ...are looking for a Senior Site Reliability Engineer to strengthen our AWS infrastructure...  ...experience in operations, software engineering, or as an SRE.... 
    Work at office
    Immediate start
    Remote work
    Work from home

    Cognitiv

    Washington DC
    5 days ago
  • $147k - $202.4k

     ...Every Identity, from AI to Human Identity is...  ...highly skilled Senior Site Reliability Engineer to join our team. We...  ...role is a blend of software engineering and systems...  ...Responsibilities Platform & Reliability: Design...  ...for critical incidents, leading root cause analysis... 
    Work at office
    Local area
    Worldwide
    Flexible hours
    Shift work

    Okta

    Washington DC
    15 days ago
  •  ...respond to critical incidents, lead root cause analysis, and...  ...experience with the Snowflake platform and large-scale data systems....  ...with a proactive approach. AI/ML experience or a strong interest in applying AI/ML to reliability, security, or operational efficiency... 
    Full time
    Work at office
    Flexible hours

    Okta

    Washington DC
    14 days ago
  •  ...seeking a high-caliber Site Reliability Engineer (SRE) to join our Forward...  ...that our complex, data-driven AI platforms remain resilient, scalable,...  .... This role is a hybrid of software engineering and systems...  ...responder in on-call rotations, leading the technical resolution of... 
    Local area

    Tiger Analytics Inc.

    Washington DC
    13 days ago
  •  ...Database, Infra (Container platforms) and Network )...  ...Expected to actively lead and triage proactively...  ...alerts for “Senior Site Reliability Engineer” roles. Bellevue, WA $...  ...San Diego, CA) Senior Software Engineer - Optical Network...  ...Reliability Engineer, AI/ML Platforms Senior... 
    Contract work
    Remote work

    Signature IT World Inc

    Washington DC
    3 days ago
  • $175k - $250k

    Senior Cloud Infrastructure Engineer Location: San Francisco...  ...with generative AI. They are the team behind...  ...flexibility. Their platform allows users to connect...  ...role, you will take the lead on designing, deploying...  ...scalability, performance, and reliability across environments.... 
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    Washington DC
    3 days ago
  •  ...the importance and value of site reliability and develops new methods and organizational...  ...offered or in DevOps, Production Software Engineering. Tools and Resource Training...  ...business processes. The Appian AI-Powered Process Platform includes everything you need to... 
    Full time

    Appian Corporation

    McLean, VA
    9 days ago
  • $117.2k - $260.1k

     ...SalesforceSalesforce is the #1 AI CRM, where humans with...  ...your career at the company leading workforce transformation in...  ...distributed systems and enterprise reliability, availability and scale.The...  ...systems. The Senior/Lead Software Engineer will be responsible both for... 
    Full time

    Salesforce

    Washington DC
    9 hours ago
  • $194k - $267k

     ...Secure Every Identity, from AI to Human Identity is the key to unlocking the potential...  ...and tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    12 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Software Engineer, AI Platform Reliability. Be the first to apply!