Lead Site Reliability Engineer
Chase
Lead Site Reliability Engineer
Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.
As a Lead Site Reliability Engineer at JPMorgan Chase within the Enterprise technology, engineering services and platform team, you hold a leadership role in your team, demonstrate strong knowledge across multiple technical domains, and advise others on the technical and business issues facing them. Take lead and conduct resiliency design reviews, break up complex problems into digestible work for other engineers, act as a technical lead for medium to large-sized products, and provide advice and mentoring to other engineers.
Job responsibilities
- Guides and assists others in the areas of building appropriate level designs and gaining consensus from peers where appropriate
- Collaborates with other software engineers and teams to design and implement deployment approaches using automated continuous integration and continuous delivery pipelines
- Collaborates with other software engineers and teams to design, develop, test, and implement availability, reliability, scalability, and solutions in their applications
- Implements infrastructure, configuration, and network as code for the applications and platforms in your remit
- Collaborates with technical experts, key stakeholders, and team members to resolve complex problems
- Understands service level indicators and utilizes service level objectives to proactively resolve issues before they impact customers
- Supports the adoption of site reliability engineering best practices within your team
- Production 24*7 support for business-critical applications
- Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
- Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls.
Required qualifications, capabilities, and skills
- Formal training or certification on site reliability engineering concepts and 5+ years applied experience
- Proficient in site reliability engineering (SRE) culture and principles, with experience implementing SRE practices within applications and platforms; strong observability background including white/black-box monitoring, SLO-based alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, and similar.
- Proficient in at least one programming language (e.g., Python, Java/Spring Boot,.NET) with strong knowledge of software applications and technical processes within a technical discipline such as cloud, artificial intelligence, Android, or related areas.
- Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.
- Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations.
- Hands-on experience with CI/CD tooling (e.g., Jenkins, GitLab) and infrastructure automation using Terraform to build reliable, repeatable delivery pipelines.
- Strong familiarity with containers and orchestration platforms (Docker, Kubernetes, ECS), including deploying, scaling, and operating containerized services in production.
- Proven ability to troubleshoot and resolve common networking issues (DNS, TCP/IP, routing, TLS, load balancing), applying structured debugging to restore service quickly.
- Collaborative, proactive team contributor: communicates clearly and persuasively with minimal supervision, identifies roadblocks early, learns new technologies quickly, and has experience with event streaming platforms such as Kafka.
Preferred qualifications, capabilities, and skills
- Ability to identify new technologies and relevant solutions to ensure design constraints are met by the software team
- Proven track record of initiating and executing ideas that address complex business challenges
- Deep expertise in networking and systems, including TCP/IP, DNS, load balancing, firewalls, and VPN technologies; strong Linux performance tuning and system-level troubleshooting skills
- Certifications a plus: AWS Certified SysOps Administrator or AWS Professional, Certified Kubernetes Administrator (CKA), Terraform Associate (or equivalent)
- Collaborative leader with a proven track record mentoring junior engineers, driving SRE best-practice adoption across teams, and communicating clearly to both technical and non-technical stakeholders (including presentations)
- Experience in handling critical incident and change management – be part of critical incident taskforce call.
- Familiarity of agile practices – preferably, scrum and Kanban
$165k - $280k
...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s most...SuggestedPermanent employmentTemporary workWorldwideWeekend work$167.7k - $245.2k
...requiring approximately 2 days per week on-site at Cisco offices in either San Francisco... ...AI agents behave as intended, improving reliability and reducing risks. This unified... ...and control.As a Senior Site Reliability Engineer (SRE), you will build, operate, and continuously...SuggestedFull timeTemporary workLocal areaFlexible hours2 days per week$165k - $190k
Obsidian Security is the leading SaaS security platform, trusted by global enterprises... ...DevOps/SRE team at Obsidian ensures that engineering excellence translates into stable,... ...complex challenges around scalability, reliability, observability, and cost efficiencyCollaborate...SuggestedWork from home$210.6k - $305.1k
...powered assurance insights within Cisco’s leading Networking, Security, Collaboration, and... ...: You have led a distributed team of 5+ engineers, can demonstrate strong technical vision... ...insurance. Please see the Cisco careers site to discover more benefits and perks. Employees...SuggestedFull timeTemporary workLocal areaFlexible hours- Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally gifted professionals and position yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Infrastructure Platforms...Suggested
$186.9k - $267.7k
...approximately 2 days per week on-site at Cisco offices in either... ...as intended, improving reliability and reducing risks. This unified... ....As a Staff Site Reliability Engineer (SRE), you will provide technical... ...-term reliability strategy, lead major infrastructure...Full timeTemporary workLocal areaFlexible hours2 days per week$232k - $263k
Obsidian Security is the leading SaaS security platform, trusted by global enterprises... ...growth and IPO readiness.Sr. Staff Site Reliability EngineerAs a Sr. Staff SRE at Obsidian... ...strategic partner to DevOps and Platform Engineering leadership, shaping a unified...Work from home- ...Site Reliability Engineer III There's nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability...Work at office
- ...Site Reliability Engineer There are NO limits to your career: come shape the future and be part of a truly unique global culture at OutSystems... ...here are your key responsibilities and duties: Lead and onboard services and teams to the reliability tenets;...Immediate startRemote workWorldwide
$137.77k - $194.59k
...distributed team of roughly 80 scientists and engineers building and operating Rubin's petascale... ...Your role: \n You will own the reliability and robustness of Rubin Observatory's... ...nature of this position, SLAC is open to on-site, hybrid, and remote work options. \n \...Remote workFlexible hoursNight shift- ...Site Reliability Engineer, Data Platform - USDS Responsibilities Engage in and improve the whole lifecycle of service, from inception and design... ...any immigration-related benefits. About USDS TikTok is the leading destination for short-form mobile video. U.S. Data Security...
$170k - $250k
...Site Reliability Engineer (SRE) Location: San Francisco, CA / Palo Alto, CA Company Stage of Funding: Growth-Stage AI Infrastructure Company ($80M Raised) Office Type: Onsite (4 Days Per Week) Salary: $170,000–$250,000 + Competitive Equity We're representing a rapidly...Work at officeVisa sponsorshipFlexible hours$170k - $230k
...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About Mithril Mithril is an AI infrastructure platform built to make... ...GPU compute more accessible and affordable for the world's leading enterprises, AI startups, and the AI research community, including...Work at officeLocal area1 day per week- ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.As a Lead Site Reliability Engineer at JPMorgan Chase within the Network Product, you hold a leadership role in your team, demonstrate strong knowledge...
$200k - $260k
...enterprise trust, as we bring Work AI to every employee, in every company. About the Role: Glean is seeking a Site Reliability Engineering Lead to foster a culture of engineering excellence, drive technical strategy, and develop a high-performing, collaborative...Work at officeHome officeFlexible hours- ...professionals for this role. JOB DESCRIPTION Elevate your engineering prowess to unprecedented levels by joining a team of... ...professionals and position yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the...
$100k - $200k
OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about...Full time- ...that's more connected, more intelligent, more sustainable for everyone. Role Summary We are seeking an experienced Site Reliability Engineer to help design, build, and operate the infrastructure that underpins the build pipelines that allow our companies to...Full timeContract work
$148k - $235.75k
...on the world.Join our team of innovative engineers who are building an AI Data Center AIOps... ...that turns raw, high-volume telemetry into reliable, job-centric insights and automation for... ...canary checks, post-deploy validation), and lead rollbacks/remediations when needed.Lead...Full time- LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is...Full timeWork at office2 days per week
$170k - $200k
We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance,...Full timeWorldwide$152k - $241.5k
...technology—and amazing people. NVIDIA is leading the way in groundbreaking developments... ...automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-... ..., Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through...Full time$230k - $250k
...network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change... ...how things have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on"...Night shift$101k - $161k
...several prestigious awards, such as Best Engineering Team, Best Company for Diversity,... ...DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s... ...the chance to be drive, develop, and lead projects in any of the following areas:...$140k - $230k
...Platforms and Product /Full-time /HybridZoox is seeking a Site Reliability Engineer to help ensure the availability, performance, and resilience... ...streamline deployment processes, and drive automation initiatives.Lead incident resolution: You will conduct thorough root cause...Full time- Lead Cloud ArchitectCooley is seeking a Lead Cloud Architect to join the Innovation team.About Cooley: Cooley is a global law firm... ...Architect is responsible for owning the technical architecture and engineering standards for a greenfield SOC2-compliant SaaS platform built...Full timeWork at officeLocal areaImmediate startRemote workWork from homeWorldwideWeekend work
$184k - $287.5k
At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability. This demanding position merges software and systems engineering efforts to guarantee flawless service operation...Full time$207k - $301k
...influential relationships with multiple stakeholders across the Site Reliability Engineering and Developer organizations.Serve as an expert on... ...knowledge related to rate limiting or sharding.Develop plans and lead projects on evolving our production systems and their...$65 - $85 per hour
...Site Reliability Engineer Sustainable Talent is partnering with a global leader who's been transforming computer graphics, PC gaming, and accelerated computing for over 25 years. We are looking for a Site Reliability Engineer to support our client's team based out of...Full timeContract workWorldwide$230k - $250k
...Site Reliability Engineer Forward is transforming how the world's most complex networks are managed and secured. Founded in 2013 by four Stanford... ...team always knows what's happening before customers do Lead incident response: on-call rotations, runbooks, post-...Night shift
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- lead operating engineer Palo Alto, CA
- lead engineer Palo Alto, CA
- site reliability engineer Palo Alto, CA
- site reliability engineer sre Palo Alto, CA
- construction site safety Palo Alto, CA
- site leader Palo Alto, CA
- official site Palo Alto, CA
- website content developer Palo Alto, CA
- IT site lead Palo Alto, CA
- site safety Palo Alto, CA

