Lead Site Reliability Engineer - Operations Excellence for AI Platforms
JP Morgan Chase
Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer focussed on Operations Excellence, you will have the opportunity to shape how we respond to, learn from, and prevent incidents across a complex, high-stakes technology environment. This is a role where your impact is visible, your voice carries weight, and your work directly influences the stability of services relied upon by millions.As a Lead Software Engineer - Site Reliability Engineer, Operations Excellence at JPMorganChase within the AI/ML & Data Platforms area, you will serve as a technical leader at the intersection of software engineering and reliability engineering — owning the incident management lifecycle, defining reliability standards, and partnering across Engineering, Product, Infrastructure, and Security to deliver durable operational improvements. You will bring structure to complexity, clarity to high-urgency situations, and a continuous improvement mindset to everything from alert quality to executive reporting. Your work will directly strengthen the firm's ability to detect, respond to, and prevent production issues at scale.Job ResponsibilitiesOwn and continuously improve the incident management lifecycle, including triage, escalation, stakeholder communications, and recovery, ensuring consistent execution and measurable improvement over time.Serve as Incident Commander for major incidents, coaching responders to follow defined processes and site reliability engineering best practices while maintaining clear and timely stakeholder communication.Drive operational readiness through drills, game days, and failure-mode exercises that improve response effectiveness, surface gaps, and build team resilience before incidents occur.Define, implement, and enforce reliability standards across services, including service level indicators and objectives, error budgets, monitoring coverage, and alert quality.Strengthen change and release management practices by establishing readiness checks, progressive delivery standards, rollback procedures, runbook quality, and production hygiene expectations.Partner with Engineering, Product, Infrastructure, and Security teams to prioritize reliability work and deliver solutions that address user and operational pain points with lasting impact.Lead problem management and root cause analysis by debugging complex production issues, identifying systemic root causes, and driving durable remediation and preventative actions.Own reliability reporting and operational governance by tracking key performance indicators — including availability versus service level objectives, mean time to detect and recover, incident trends, alert noise, and change failure rate — and producing executive-ready summaries.Facilitate operational forums including operations reviews, incident review boards, and reliability councils to drive accountability, share learnings, and align stakeholders on priorities.Drives team adoption of enterprise-authorized AI-assisted engineering practices within the work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across the team.Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.Required qualifications, capabilities, and skillsFormal training or certification on software engineering concepts and advanced applied experience.Demonstrated experience improving operational processes, including incident management, root cause analysis, problem management, and operational governance in complex production environments.Strong knowledge of reliability engineering concepts including service level objectives and indicators, error budgets, capacity planning, resilience patterns, and observability.Proven ability to influence across teams, drive cross-functional initiatives, and communicate clearly with stakeholders — particularly in high-urgency, time-sensitive situations.Strong analytical and reporting skills with the ability to define metrics, generate actionable insights, and drive decisions based on operational data.Hands-on experience with cloud platforms such as Amazon Web Services, Google Cloud Platform, or Microsoft Azure, and infrastructure-as-code tooling including Terraform.Systematic problem-solving and troubleshooting skills applied to complex, distributed systems in production environments.Demonstrated ability to operate with a high degree of ownership, self-direction, and urgency in ambiguous or fast-moving situations.Hands-on experience using enterprise-authorized AI-assisted software development tools within the work environment (e.g., for coding, test creation, troubleshooting, or documentation) with demonstrated ability to critically evaluate, validate, and refine AI-generated outputs for correctness, performance, and security.Understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; ability to guide peers on safe and effective usage within team practices.Preferred qualifications, capabilities, and skillsExperience leading reliability programs across multiple services or teams at a platform or domain level.Experience building automation for operations, including runbook automation, event correlation, alert tuning, and self-healing patterns.Exposure to agentic AI or AI-driven operations concepts, including workflow orchestration, decision support, and safe automation patterns.Familiarity with operational governance frameworks and experience facilitating reliability forums or review boards at scale. JPMorganChase, one of the oldest financial institutions, offers innovative financial solutions to millions of consumers, small businesses and many of the world’s most prominent corporate, institutional and government clients under the J.P. Morgan and Chase brands. Our history spans over 200 years and today we are a leader in investment banking, consumer and small business banking, commercial banking, financial transaction processing and asset management. We offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may receive commission-based pay and/or discretionary incentive compensation, paid in the form of cash and/or forfeitable equity, awarded in recognition of individual achievements and contributions. We also offer a range of benefits and programs to meet employee needs, based on eligibility. These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare, tuition reimbursement, mental health support, financial coaching and more. Additional details about total compensation and benefits will be provided during the hiring process. We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants’ and employees’ religious practices and beliefs, as well as mental health or physical disability needs. Visit our FAQs for more information about requesting an accommodation.JPMorgan Chase & Co. is an Equal Opportunity Employer, including Disability/VeteransOur professionals in our Corporate Functions cover a diverse range of areas from finance and risk to human resources and marketing. Our corporate teams are an essential part of our company, ensuring that we’re setting our businesses, clients, customers and employees up for success.Full timePosting Date: 2026-08-18
$113.1k - $232.3k
Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead... ..., performance, and operational integrity of high-visibility products and platforms and the environments they... ...tradition of delivering with excellence. The successful candidate...PlatformOperationsWork at officeLocal areaVisa sponsorshipFlexible hours3 days per week- ...realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the AI Machine Learning and Data platform team, you hold a leadership role... ...root cause analysis. Drive operational readiness and strengthen change and...PlatformOperations
$150k - $200k
...Runpod is the AI Developer Cloud... ...AI on one platform. The platform... ...announcement: . The Reliability team owns the... ..., and operational excellence of Runpod’s global... ...standards across engineering Designing... ...systems. As a Site Reliability Engineer... ...services Lead incident response...PlatformOperationsRemote workVisa sponsorshipWork visaFlexible hours- ...leaders across engineering and technology... ...objective reliability goals for services... ...Senior Azure Site Reliability Engineer... ...Azure platform. The role focuses... ..., and operational excellence. The ideal candidate... ...they support Lead complex... ...Experience with Azure AI Foundry,...PlatformOperationsWork at officeShift workDay shift
$142.32k - $213.48k
...CitiAbout Citi:Citi, the leading global bank, has... ...daily life, our Operations & Technology teams are... ...architecture and ensuring our platforms provide a first-... ...to deliver excellence through secure, reliable, and efficient services... ...together.The Role:The AI Partnerships Lead (...PlatformOperationsFull timeWork at office- ...tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the AIML Platform team, you hold a... ...Uses enterprise-authorized AI capabilities within the work... ...outputs and handling operational data according to sensitivity...PlatformOperations
- ...and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Chief Data & Analytics Office (CDAO) AI/ML & Data Platforms team,, you will solve complex and broad... ...your knowledge of end-to-end operations, availability, reliability, and scalability...PlatformOperationsWork at office
- ...environment. As a Lead Technical... ...Observability Platforms, you will drive... ...development of operational plans and risk... ...solutionsPartner with engineering and product... ...enterprise-authorized AI capabilities... ...experience ( Excel and Sharepoint)... ...care coverage, on-site health and...PlatformOperationsShift work
$155k - $175k
...and experienced Site Reliability Manager to join... ...services. You will lead a team of engineers focusing on... ...DevSecOps, and Platform Lifecycle Management... ...experience supporting operations and maintenance... ...Cloudwatch). Excellent communication... ...transfer. Statement on AI and Hiring...PlatformOperationsWork experience placementH1bWork at officeLocal area$114k - $148k
...Site Reliability Engineer Location: Remote, United States Employment... ...focus on ensuring the platform and services... ...implement and maintain operations. A passion for technology... ...operational data, embeds AI for better decisions... ...we provide are: Excellent Medical Plan. Dental...PlatformOperationsFull timeTemporary workWork experience placementRemote work- ...Zscaler Zero Trust Exchange platform combined with advanced AI combats billions of... ...We are looking for a Site Reliability Engineer-SkillBridge Intern (San... ...your bias for action. You operate with integrity because... ...drive for technical excellence with the need to deliver...PlatformOperationsInternshipWork at officeLocal areaRemote workWorldwide
$253k - $336k
...by Lattice OS, an AI-powered operating system that turns... ...TEAM: CorpTech Platform is the internal engineering force multiplier... ..., QA and release excellence, ERP engineering,... ...The Director of Site Reliability Engineering owns... ...operations. This role leads the SRE...PlatformOperationsFull timeWork experience placement- ...improving how Markets Operations runs at scale.... ..., product, engineering, and data to... ...capabilities designed for reliability, control, and... ...Learning Lead at... ...and generative AI solutions that... ...AI products and platforms across multiple... ...care coverage, on-site health and wellness...PlatformOperations
- Elevate your engineering prowess to unprecedented levels by... ...the top echelon in site reliability.As a Senior Lead Site Reliability Engineer... ...Analytics Office (CDAO) AI/ML & Data Platforms team, you work with your... ...reliability design and operational decisioning (e.g.,...PlatformOperationsWork at office
- ...-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within... ...knowledge of end-to-end operations, availability, reliability,... ...scalability of your application or platform. Job responsibilities... ...Uses enterprise-authorized AI capabilities within the...PlatformOperationsLocal area
$176.72k - $265.08k
...ProfessionalCompany: CitiThe AI Transformation... ...for leading and... ...that builds and operates Master and Reference... ...70 engineers, engineering managers... ...Reference Data platforms and services —... ...velocity, quality, reliability, AI adoption,... ...operational excellence.Experience in...PlatformOperationsFull timeImmediate startShift work$110k - $140k
...Group, Inc. (AIG) is a leading global insurance... ...analytics capabilities, excellent communication skills,... ...into actionable data engineering specifications by collaborating... ...Owner for Data Platform & Tool Development: Develop... ...join us — across our operations, we are thinking in...PlatformOperationsFull timeWork at office$200k - $275k
## Strategy, Ops & Excellence Lead - Statistical SciencesApplyremote type... ...& Process Innovation, Operational Excellence, and Communications... ...execution.* Evaluation of vendor platforms and solutions against key... ...Current, strong knowledge of AI tools to drive efficiency* Good...PlatformOperationsTemporary workRemote work$1,000 per month
...Improvement & AI Intelligence Automation Lead Location Employment... ...award-winning platform that enables... ...parent rooms, on-site gyms,... ...improvement, scalable operating models, and intelligent... ...of operational excellence, process... ...architects, and engineering teams to intake...PlatformOperationsPermanent employmentFull timeWork at officeLocal areaRemote workWorldwideFlexible hours3 days per week$176.72k - $265.08k
...CitiOverview of Citi:Citi, the leading global bank, has... ..., our Enterprise Operations & Technology teams... ...and ensuring our platforms provide a first-class... ...experiences to deliver excellence through secure, reliable, and efficient... ...skillsFamiliarity with AI tools and use AI tools...PlatformOperationsFull timeAfternoon shift- ...Group Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure... ...management, and daily operations. Bitdeer also offers... ...focus on judgment calls the platform can't yet make, and every... ...rack and stack, labeling (on-site roles). Execute structured...PlatformOperationsShift workNight shift
- ...Administrator Lead provides... ...projects and operational objectives,... ...Manager, Systems Engineering. The... ..., balancing reliability, security, cost... ...or failover sites to test, validate... ...Excellent written and... ...and related platforms Expert knowledge... ...Experience with AI‑enabled technologies...PlatformOperationsVisa sponsorship
- ...environment.As a Lead Technical Program... ...Sector- Infrastructure Platforms Foundational... ...the development of operational plans and risk management... ...& Due Diligence, AI Amplification &... ...teams, including engineering, product, and business... ...care coverage, on-site health and...PlatformOperations
- ...core workflows, improve operating leverage, and surface... ...practical lessons about how AI changes enterprise... ...Storytelling & AI Innovation Lead partners directly with... ...Finance and Enterprise Platforms & Technology leaders;... ...superficial demo. You excel at translating between...PlatformOperationsWork at officeImmediate startRemote work
- ...group delivers secure, reliable technology... ...performance of enterprise platforms.As a Principal Site Reliability Engineer (SRE), you will drive operational excellence across mission-... ...systems. You will lead reliability initiatives... ..., automation, and AI-powered technologies...PlatformRemote workFlexible hours
- ...solutions.As a Senior Lead Architect -... ..., secure, and AI-enabled... ...and technical operations and processes.Serves... ...and partner with engineering, product organizations... ...- platform choices, preferred... ...architecture domains. Excellent communication... ...coverage, on-site health and wellness...PlatformOperations
- ...owners, business, operations, and software developers... ...closely with engineering to deliver reliable, scalable platforms. Expect to work... ...opportunities to apply AI/LLM-based... ...easy onboarding.Your excellent communication skills... ...care coverage, on-site health and wellness...PlatformOperations
- ...deliver market-leading technology... ...software engineering career to the... ...engineering excellence across the team... ...end-to-end platform capabilities... ...-authorized AI-assisted... ...speed, and operational outcomes (e.... ...consistency, reliability, and adherence... ...coverage, on-site health and wellness...PlatformOperations
- ...Data Centers LLC is seeking a Sr. Platform Owner, Halo, Global, to lead the Halo AI platform's health, scalability,... ...from architecture to production operations. This remote US-based role... ...cloud infrastructure to balance reliability with rapid innovation. #J-18808...PlatformOperationsRemote work
- ...you're a Senior Lead Software Engineer who takes ownership... ...standards, and reliability posture across... ...within the Corporate AI/ML Data Platforms - Machine Learning Center of Excellence, you will design... ...the firm to operate at the forefront... ...coverage, on-site health and wellness...PlatformOperations
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer - Operations Excellence for AI Platforms. Be the first to apply!
- lead web developer Jersey City, NJ
- lead operating engineer Jersey City, NJ
- lead algorithm engineer Jersey City, NJ
- lead infrastructure engineer Jersey City, NJ
- lead network engineer Jersey City, NJ
- lead engineer Jersey City, NJ
- site reliability engineer Jersey City, NJ
- site reliability engineer sre Jersey City, NJ
- website content developer Jersey City, NJ
- site leader Jersey City, NJ


