Lead Site Reliability Engineer
hackajob
Lead Site Reliability Engineer
Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Commercial & Investment Bank, Production Management team, you hold a leadership role in your team, demonstrate strong knowledge across multiple technical domains, and advise others on the technical and business issues facing them. Take lead and conduct resiliency design reviews, break up complex problems into digestible work for other engineers, act as a technical lead for medium to large-sized products, and provide advice and mentoring to other engineers.
Job Responsibilities
- Lead the Production Management team supporting Sales Execution platforms across Rates, Credit, FX, SPG, and Repo, setting direction, priorities, and performance expectations.
- Own stability, availability, resiliency, and end-to-end operational performance of business-critical Sales platforms, with clear accountability for outcomes.
- Act as a senior escalation point during critical incidents, driving rapid triage, decisive coordination, and recovery actions aligned to business impact.
- Build deep understanding of Sales Execution workflows (RFQ, pricing, execution, booking, market data, and trade lifecycle) to anticipate risks and improve support effectiveness.
- Partner closely with Sales, Trading, Product, Application Development, Operations, and Infrastructure to improve platform stability, user experience, and operational efficiency.
- Serve as a trusted advisor to Front Office Sales stakeholders, providing concise, business-focused communications during incidents and key initiatives.
- Drive operational consistency and service maturity through standardization, governance participation, service reviews, and disciplined support model integration for new capabilities.
- Lead reliability improvements by applying systems thinking and root-cause practices, expanding SRE adoption (observability, monitoring, automation, operational analytics), and improving supportability with engineering teams.
- Run strong incident, problem, and change management—major incident response, RCA and remediation to eliminate recurrence, and ensuring changes meet readiness standards (testing, monitoring, and resiliency/DR validation).
- Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
- Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls.
Required qualifications, capabilities, and skills
- Formal training or certification on site reliability engineering concepts and 5+ years applied experience
- Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.
- Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations.
- Leadership experience across Production Support, Production Management, SRE, and Technology Operations teams, delivering stable and resilient production services.
- Proven background supporting Front Office Sales users and Sales Execution platforms in high-availability, time-sensitive environments.
- Strong knowledge of Rates, Credit, FX, Repo, and/or SPG workflows, with solid understanding of RFQs, pricing, execution, booking, market data, and trade lifecycle processes.
- Strong systems thinking and problem-solving capability to assess complex, cross-domain production issues and drive end-to-end resolution.
- Demonstrated major incident leadership, coordinating effectively across teams to restore service rapidly and drive root-cause remediation.
- Track record of partnering with Application Development, Product, Sales, and business stakeholders to improve reliability, service quality, and operational maturity.
- Strong observability and service management expertise (Dynatrace, Splunk, Geneos, Grafana; ITIL Incident/Problem/Change/Availability), with excellent verbal/written communication and people leadership.
Preferred qualifications, capabilities, and skills
- Extensive experience supporting electronic Sales Execution platforms within Capital Markets environments, ensuring high availability and business-critical performance.
- Deep knowledge of Rates, Credit, FX, Repo, and/or SPG workflows, enabling effective support aligned to trading and sales execution needs.
- Proven ability to support a broad portfolio of Sales technology applications, managing operational risk and prioritization across multiple platforms.
- Demonstrated experience integrating support teams and standardizing operating models across multiple application groups to drive consistency and service maturity.
- Strong track record partnering directly with Front Office Sales teams to support client-facing electronic execution services and deliver business-focused outcomes.
- Hands-on technology experience across cloud (AWS/Azure/GCP), automation and scripting (Python, Shell, PowerShell, Ansible, Terraform), and modern distributed platforms (containers, microservices, Kubernetes/OpenShift).
$113.1k - $232.3k
Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied AI Site Reliability Engineer II, you will actively engage in your engineering craft, taking a hands-on approach to the reliability, performance, and operational integrity...SuggestedWork at officeLocal areaVisa sponsorshipFlexible hours3 days per week- ...EIT) organization is expanding, and we are seeking a Senior Site Reliability Engineer to help drive a major architectural modernization. In this... ...SLOs and error budgets. • Modernization & Migration: Lead the technical execution of re-architecting and redeploying...SuggestedPermanent employmentFull timeH1bLocal areaRemote workShift work
$153k - $210k
...Senior Software Engineer, Site Reliability Engineering Reno, NV; San Ramon, CA; NYC - Hybrid Are you passionate about building resilient... ...to resolution with very infrequent after-hours support. Lead blameless postmortems and implement long-term improvements...SuggestedFull time$207k - $300k
...designs from peers, providing feedback to ensure best practices in reliability, security, and efficiency.Triage and resolve complex system... ...execution of software development initiatives. Mentor other engineers and contribute to the engineering community through documentation...SuggestedFull timeWork at office$200k - $250k
Hudson River Trading (HRT) is seeking a Senior Site Reliability Engineer focused on storage to join our growing Enterprise SRE team. This team is responsible for developing and maintaining productivity service infrastructure for the entire firm, both on-prem and in the...SuggestedWork at officeLocal areaImmediate start$139k - $257.55k
...Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning,... ...productivity and personalized customer experiences. Adobe’s industry-leading offerings including Adobe Acrobat Studio, Adobe Express,...Full timeTemporary workLocal areaRemote workWorldwide$167.7k - $245.2k
...assurance insights within Cisco’s leading Networking, Security,... ...effective.We’re looking for talented engineers with a software or operations... ...teams to ensure the reliability, performance and security of... ...Please see the Cisco careers site to discover more benefits and...Full timeTemporary workWork at officeLocal areaFlexible hours1 day per week$141k - $216.6k
...building a safer, more connected world.Position OverviewAs a Site Reliability Engineer, you'll own the reliability, observability, and operational... ...of ambiguous technical problems with minimal direction.Leading operational improvements rather than simply executing assigned...Work experience placementWork at office$158.5k - $172k
...velocity energy of a powerhouse startup.As a leading U.S. ordering and delivery marketplace,... ....About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will... ...high-impact position driving continuous reliability, deep system optimization, and automation...Full timeTemporary workWork at officeFlexible hours3 days per week$45 - $85 per hour
DescriptionThe Site Reliability Engineering groups goal is to ensure Customers can always use the service reliably.We're looking for engineers to... ...law.About TEKsystems and TEKsystems Global Services We’re a leading provider of business and technology services. We...Contract workTemporary work- ...the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial & Investment Bank... ...tooling to reduce manual steps and speed repeatable triage.Lead L1/L2 production support using SRE practices: quickly...Shift work
$182.8k - $247.3k
...mission to develop education for our half a billion (and growing!) learners around the world.About the role...As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed systems...Work experience placement$150k - $250k
What We DoAt Goldman Sachs, our Engineers don't just make things - we make things possible... ...Global Banking & Markets business, the Site Reliability Engineering (SRE) team ensures the... ...into durable engineering improvements.Lead incident response for latency-sensitive...Full timeTemporary workPart time$194k - $267k
...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$194k - $267k
...let's talk.The TeamWe are looking for an experienced Staff Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is... ...technical challenges. You will serve as a key technical lead within the EPG SRE organization, partnering with software engineers...Local areaWorldwideFlexible hours$194k - $267k
...Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk... ...servicesIncident Response: Participate in on-call rotations and lead post-incident reviews to drive systemic improvements and "...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$131k - $164k
Position OverviewWe are seeking a highly skilled Staff Site Reliability Engineer with deep technical expertise across VMware, Linux, and automation... ...they need to drive greater impact and accountability - to lead with purpose. Our employees are passionate, smart, and...Work at officeLocal areaVisa sponsorshipFlexible hours$190k - $260k
Who are we?Cohere is the leading security-first enterprise AI company... ...is a team of researchers, engineers, designers, and more, who are... ...high-performance, scalable and reliable machine learning systems? Do... ...? We are looking for a Site Reliability Engineer to join...Full timeWork experience placementWork at officeLocal areaRemote workHome office- ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Commercial & Investment Bank, Production Management team, you hold a leadership...
$150k - $220k
...and innovators in this way. The Role: As an engineering organization, we pride ourselves on engineering as a creative... ...achieve autonomy, mastery, and purpose. The Manager, Site Reliability Engineering will lead Forge’s SRE team responsible for keeping Forge systems...Local area$182k - $250.8k
...Okta is the backbone of our platform's reliability and operational excellence. We are a forward-thinking group of engineers and leaders who believe that great... ...of users worldwide. As a Manager, Site Reliability Engineer, you'll lead this team with a focus on scalability...Permanent employmentLocal areaRemote workWorldwideFlexible hoursWeekend workWeekday work$150k - $190k
Senior Site Reliability Engineer, VPAt Morgan Stanley, we advise, originate, trade, manage and distribute capital for governments, institutions... ..., and always do so with a standard of excellence. We are a leading global financial services firm that conducts its business through...Temporary workWorldwideFlexible hoursWeekend work$195k - $275k
Morgan Stanley is a leading global financial services firm providing a wide range of investment banking, securities, investment management... ...& Release Management, and the Chief Operating Office.The Reliability Operations (RO) within WMT is responsible for providing swift,...Temporary workWork at officeWorldwideNight shift$130k - $250k
What We DoAt Goldman Sachs, our Engineers don't just make things - we make things possible... ...your journey here.Securities Frontline Site Reliability Engineers (SREs) play a critical role... ...serve to grow. Founded in 1869, we are a leading global investment banking, securities...Full timeTemporary workPart timeImmediate start$150k - $160k
Front-End & AdTech Site Reliability Engineer (SRE)Haymarket Media, Inc. is seeking a Front-End & AdTech Site Reliability Engineer (SRE) to join... ...fast, resilient, and optimized at the edge.Responsibilities:Lead development team initiatives to:Architect and maintain Cloudflare...Work at officeLocal area$105k - $300k
...Site Reliability Engineer At Citadel, a leading investor in the world's financial markets, we aim to win together as one team to earn the long-term trust of our capital partners and each other. Our collaborative approach allows technologists to grow alongside other...$115k - $125k
...Site Reliability Engineer New York City, NY Pico fuels the global capital markets community by providing exceptional market data services and customized managed infrastructure solutions. As financial industry experts at the center of markets and technology, we help...Work experience placementWork at officeWork from homeMonday to FridayFlexible hoursShift workWeekend workAfternoon shiftEarly shift$140k - $215k
...intersection of our Core Platform and Embedded Reliability charters: building the foundational... ..., while embedding directly with product engineering teams and their leadership to drive... ...comprises hundreds of libraries and services.Lead initiatives around reliability,...Full timeWork experience placementWork at officeLocal area2 days per week3 days per week- ...Site Reliability Engineer Our Client, a multinational telecommunications technology company is seeking a Site Reliability Engineer (SRE I) to join our Video Platform Engineering Team. As a Level 1 SRE, you will work closely with senior engineers to respond to incidents...Temporary work
$191k - $226k
...anyone else can. About the role: We are seeking a Senior Site Reliability Engineer to own the reliability, performance, and resilience of the... .... You will run the machine: defining and upholding SLOs, leading incident response, and driving the automation and standards...Remote workWork visaFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- lead algorithm engineer New York, NY
- lead web developer New York, NY
- lead security engineer New York, NY
- lead network engineer New York, NY
- lead infrastructure engineer New York, NY
- lead system engineer New York, NY
- lead operating engineer New York, NY
- lead engineer New York, NY
- site reliability engineer New York, NY
- site reliability engineer remote New York, NY



