Lead Site Reliability Engineer
JPMorgan Chase & Co.
Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Commercial & Investment Bank, Production Management team, you hold a leadership role in your team, demonstrate strong knowledge across multiple technical domains, and advise others on the technical and business issues facing them. Take lead and conduct resiliency design reviews, break up complex problems into digestible work for other engineers, act as a technical lead for medium to large-sized products, and provide advice and mentoring to other engineers. Job Responsibilities * Lead the Production Management team supporting Sales Execution platforms across Rates, Credit, FX, SPG, and Repo, setting direction, priorities, and performance expectations. * Own stability, availability, resiliency, and end-to-end operational performance of business-critical Sales platforms, with clear accountability for outcomes. * Act as a senior escalation point during critical incidents, driving rapid triage, decisive coordination, and recovery actions aligned to business impact. * Build deep understanding of Sales Execution workflows (RFQ, pricing, execution, booking, market data, and trade lifecycle) to anticipate risks and improve support effectiveness. * Partner closely with Sales, Trading, Product, Application Development, Operations, and Infrastructure to improve platform stability, user experience, and operational efficiency. * Serve as a trusted advisor to Front Office Sales stakeholders, providing concise, business-focused communications during incidents and key initiatives. * Drive operational consistency and service maturity through standardization, governance participation, service reviews, and disciplined support model integration for new capabilities. * Lead reliability improvements by applying systems thinking and root-cause practices, expanding SRE adoption (observability, monitoring, automation, operational analytics), and improving supportability with engineering teams. * Run strong incident, problem, and change management—major incident response, RCA and remediation to eliminate recurrence, and ensuring changes meet readiness standards (testing, monitoring, and resiliency/DR validation). * Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements. * Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls. Required qualifications, capabilities, and skills * Formal training or certification on site reliability engineering concepts and 5+ years applied experience * Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity. * Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations. * Leadership experience across Production Support, Production Management, SRE, and Technology Operations teams, delivering stable and resilient production services. * Proven background supporting Front Office Sales users and Sales Execution platforms in high-availability, time-sensitive environments. * Strong knowledge of Rates, Credit, FX, Repo, and/or SPG workflows, with solid understanding of RFQs, pricing, execution, booking, market data, and trade lifecycle processes. * Strong systems thinking and problem-solving capability to assess complex, cross-domain production issues and drive end-to-end resolution. * Demonstrated major incident leadership, coordinating effectively across teams to restore service rapidly and drive root-cause remediation. * Track record of partnering with Application Development, Product, Sales, and business stakeholders to improve reliability, service quality, and operational maturity. * Strong observability and service management expertise (Dynatrace, Splunk, Geneos, Grafana; ITIL Incident/Problem/Change/Availability), with excellent verbal/written communication and people leadership. Preferred qualifications, capabilities, and skills * Extensive experience supporting electronic Sales Execution platforms within Capital Markets environments, ensuring high availability and business-critical performance. * Deep knowledge of Rates, Credit, FX, Repo, and/or SPG workflows, enabling effective support aligned to trading and sales execution needs. * Proven ability to support a broad portfolio of Sales technology applications, managing operational risk and prioritization across multiple platforms. * Demonstrated experience integrating support teams and standardizing operating models across multiple application groups to drive consistency and service maturity. * Strong track record partnering directly with Front Office Sales teams to support client-facing electronic execution services and deliver business-focused outcomes. * Hands-on technology experience across cloud (AWS/Azure/GCP), automation and scripting (Python, Shell, PowerShell, Ansible, Terraform), and modern distributed platforms (containers, microservices, Kubernetes/OpenShift). JPMorganChase, one of the oldest financial institutions, offers innovative financial solutions to millions of consumers, small businesses and many of the world’s most prominent corporate, institutional and government clients under the J.P. Morgan and Chase brands. Our history spans over 200 years and today we are a leader in investment banking, consumer and small business banking, commercial banking, financial transaction processing and asset management. We offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may receive commission-based pay and/or discretionary incentive compensation, paid in the form of cash and/or forfeitable equity, awarded in recognition of individual achievements and contributions. We also offer a range of benefits and programs to meet employee needs, based on eligibility. These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare, tuition reimbursement, mental health support, financial coaching and more. Additional details about total compensation and benefits will be provided during the hiring process. We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants’ and employees’ religious practices and beliefs, as well as mental health or physical disability needs. Visit our FAQs [ for more information about requesting an accommodation. JPMorgan Chase & Co. is an Equal Opportunity Employer, including Disability/Veterans J.P. Morgan’s Commercial & Investment Bank is a global leader across banking, markets, securities services and payments. Corporations, governments and institutions throughout the world entrust us with their business in more than 100 countries. The Commercial & Investment Bank provides strategic advice, raises capital, manages risk and extends liquidity in markets around the world.
$111k - $218k
...The Site Reliability Engineering team designs and builds the global infrastructure on which we deploy our services, focusing on the above mentioned flagship MongoDB Atlas platform. As our customers grow and globalize, our services must satisfy demands for low-latency...SuggestedFull timeLocal areaWorldwideFlexible hours$158.5k - $172k
...velocity energy of a powerhouse startup. As a leading U.S. ordering and delivery marketplace,... .... About The Opportunity As a Senior Engineer on the Runtime Automation team, you will... ...high-impact position driving continuous reliability, deep system optimization, and...SuggestedFull timeWork at office3 days per week- ...researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: At... ...new cloud infrastructure company, we seek to improve our reliability dramatically while scaling the size of our platform and customer...Suggested
- ...A dynamic fintech company in New York is seeking a Product & Platform Monitoring Manager to ensure the reliability of its fintech products. The role focuses on end-to-end monitoring of customer journeys, incident management, and collaboration with various teams. Candidates...Suggested
- ...Karsun Solutions, LLC is seeking a Site Reliability Manager to lead a multi-disciplinary team responsible for reliability, security, and platform lifecycle across AWS-based services. The role emphasizes collaboration, observability, and continuous improvement in a client...Suggested
- ...Karsun Solutions in the DMV area is seeking a Site Reliability Manager to ensure reliability, scalability, and performance of our systems. You will lead a team focusing on Application Reliability, DevSecOps, and Platform Lifecycle Management. The ideal candidate has 1...
- ...Komodor, a remote-first company, is seeking a Solutions Engineer to connect customer business initiatives to the Komodor platform, understanding developers, DevOps and Incident response teams working with Kubernetes. You will identify customer pain points and communicate...Remote work
- ...optimize production infrastructure across CI/CD, cloud deployments, and security. You will collaborate with our internal product and engineering teams to keep services scalable, secure, and highly available. The role emphasizes GitHub Actions, Terraform, Vercel, AWS core...Remote work
- ...A financial technology company based in New York is seeking a Product & Platform Monitoring Manager to ensure the reliability and health of their fintech products. The role involves monitoring customer journeys, APIs, and incident management, requiring 5+ years of experience...
- ...and Antler, we empower CISOs to proactively manage human risk—the leading cause of cybersecurity breaches—and build safer, more resilient organizations. The Role: As a Senior Site Reliability Engineer (SRE) at Dune Security, you will play a critical role in ensuring our...Full timeWork at office
$70k - $90k
...TD SYNNEX is seeking a Business Systems Engineer for its Global eCommerce and ERP team. This position involves translating business requirements into technical specifications for their systems. The ideal candidate has 1-3 years of experience, is knowledgeable in Java,...Remote work- ...The Consulting Solutions is seeking an experienced Senior / Staff Engineer for our SRE, InfraSec team in Seattle. The role involves leading the security of cloud-based infrastructure, mentoring a team of SREs, and collaborating with other engineering teams to ensure high...Remote work
- Fun is seeking a Business Development professional to own enterprise deals end-to-end and accelerate commercial growth at the frontier of on-chain payments. This role is primarily in-person at our Midtown, NYC headquarters with a Monday–Thursday collaboration rhythm and...Work from home
$93k - $160k
...Palantir Technologies is seeking a Site Reliability Operations Analyst in New York, NY. In this role, you will streamline workflows and reduce friction in deployments. Your responsibilities include supporting deployments, removing roadblocks, and managing multiple challenges...$130k - $170k
...NBC Universal is looking for a Staff Software Engineer (SRE Lead) in New York, NY. This role involves overseeing day-to-day operations of SAP BTP CPI applications, managing incidents, leading offshore support teams, and ensuring high system performance. Candidates should...Remote work$160k - $180k
...Socure is seeking a Site Reliability Engineer in New York to enhance our identity trust infrastructure. In this role, you will take full ownership of AWS and Kubernetes platforms, ensuring high reliability and operability. The ideal candidate will possess extensive experience...$153k - $210k
...Senior Software Engineer, Site Reliability Engineering Reno, NV; San Ramon, CA; NYC - Hybrid Are you passionate about building resilient, highly... ...issues to resolution with very infrequent after‑hours support. Lead blameless postmortems and implement long‑term improvements...$123k - $165k
...Site Reliability Engineer II Our engineering fleet is a horizontal set of teams providing engineering services across the organization. Our specific team provides reliability engineering and operational support to backend service development teams. Technology is...$176.75k - $209.1k
...Site Reliability Engineer At Peloton, we view Platform as a Product. A phenomenal platform unlocks speed of development and learning. It allows us to scale easily, enabling our engineers to maximize attention on new features and capabilities. A key to crafting a phenomenal...Temporary work$150k - $175k
...Site Reliability Engineer At ASAPP, our mission is simple: deliver the best AI-powered customer experience—faster than anyone else. To achieve that, we're guided by principles that shape how we think, build, and execute. We value customer obsession, purposeful speed...Remote work- ...self-healing, deployment/rollback automation). Establish reliability standards: SLOs/SLIs, error budgets, production readiness reviews... ..., and release risk controls. Performance and reliability engineering: capacity planning, load/performance analysis, resilience...
$100k - $250k
...financial markets. Role Roadmap As a member of Kalshi's engineering team, you'll help build the next-generation financial... ..., and evolve. What You'll Do Improve observability, reliability, and service availability by defining and measuring key metrics...Local area- ...Applications Deployment Responsible for reliability and support of Container Platform on-... ...Perform blameless RCA, partner with engineering and operation teams across the... ...Additional Skills : Automation Process Engineer,Site Reliability Engineer,Full Stack DeveloperThis...
- ...exceptional professionals for this role. JOB DESCRIPTION As a Site Reliability Engineering at JPMorgan Chase within the Enterprise technology,... ...situations with composure and tact. J ob responsibilitie s Lead SRE practices that balance delivery speed, efficiency, and...
- ...strategy sessions with other Optum Teams Require 2+ years of experience with Terraform Require 2+ years of experience with DevOps Solution Architect, DevOps, or System Engineer certification in one or more public cloud providers Terraform certification #J-18808-Ljbffr...
$175k - $230k
...those residing in senior living facilities. Falls are the leading cause of injury-related death among adults over 65. And yet... ...be a 24x7, highly available platform for elder care. As a Site Reliability Engineer, you'll partner with engineering teams across the organization...ApprenticeshipWork at officeLocal areaRemote work2 days per week- ...coasts. If you're driven by impact, pace, and raising the bar. This is the place. The Role As a Staff Site Reliability Engineer you'll play a lead role on the founding SRE team at our new NYC engineering hub. You'll own multi-team reliability and...Work at office
$120k - $165k
...and shape the future of our communities. This is a Lead Software Production Management & Reliability Engineering position at Director level which is part of the... ...Overview The Wealth Management Production Management Site Reliability Engineer position is a highly visible/...Temporary workWork at office$105.79k - $141.05k
...shape the future of AI‑ready connectivity, join us today. The Role We are seeking a highly skilled and proactive Lead Site Reliability Engineer (SRE) to join our team, focusing on production support and performance optimization across our portal ecosystem. This role...Full timeTemporary workRemote work$132.23k - $176.31k
...the future of AI‑ready connectivity, join us today. The Role We are seeking a highly skilled and proactive Senior Lead Site Reliability Engineer (SRE) to join our team, focusing on production support and performance optimization across our portal ecosystem. This...Full timeTemporary workRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- lead web developer New York, NY
- lead algorithm engineer New York, NY
- lead network engineer New York, NY
- lead infrastructure engineer New York, NY
- lead engineer New York, NY
- lead operating engineer New York, NY
- lead system engineer New York, NY
- site reliability engineer New York, NY
- site reliability engineer remote New York, NY
- site reliability engineer sre New York, NY


