Lead Site Reliability Engineer
Chase
Lead Site Reliability Engineer
Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.
As a Lead Site Reliability Engineer at JPMorgan Chase within the Enterprise technology, engineering services and platform team, you hold a leadership role in your team, demonstrate strong knowledge across multiple technical domains, and advise others on the technical and business issues facing them. Take lead and conduct resiliency design reviews, break up complex problems into digestible work for other engineers, act as a technical lead for medium to large-sized products, and provide advice and mentoring to other engineers.
Job responsibilities
- Guides and assists others in the areas of building appropriate level designs and gaining consensus from peers where appropriate
- Collaborates with other software engineers and teams to design and implement deployment approaches using automated continuous integration and continuous delivery pipelines
- Collaborates with other software engineers and teams to design, develop, test, and implement availability, reliability, scalability, and solutions in their applications
- Implements infrastructure, configuration, and network as code for the applications and platforms in your remit
- Collaborates with technical experts, key stakeholders, and team members to resolve complex problems
- Understands service level indicators and utilizes service level objectives to proactively resolve issues before they impact customers
- Supports the adoption of site reliability engineering best practices within your team
- Production 24*7 support for business-critical applications
- Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
- Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls.
Required qualifications, capabilities, and skills
- Formal training or certification on site reliability engineering concepts and 5+ years applied experience
- Proficient in site reliability engineering (SRE) culture and principles, with experience implementing SRE practices within applications and platforms; strong observability background including white/black-box monitoring, SLO-based alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, and similar.
- Proficient in at least one programming language (e.g., Python, Java/Spring Boot,.NET) with strong knowledge of software applications and technical processes within a technical discipline such as cloud, artificial intelligence, Android, or related areas.
- Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.
- Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations.
- Hands-on experience with CI/CD tooling (e.g., Jenkins, GitLab) and infrastructure automation using Terraform to build reliable, repeatable delivery pipelines.
- Strong familiarity with containers and orchestration platforms (Docker, Kubernetes, ECS), including deploying, scaling, and operating containerized services in production.
- Proven ability to troubleshoot and resolve common networking issues (DNS, TCP/IP, routing, TLS, load balancing), applying structured debugging to restore service quickly.
- Collaborative, proactive team contributor: communicates clearly and persuasively with minimal supervision, identifies roadblocks early, learns new technologies quickly, and has experience with event streaming platforms such as Kafka.
Preferred qualifications, capabilities, and skills
- Ability to identify new technologies and relevant solutions to ensure design constraints are met by the software team
- Proven track record of initiating and executing ideas that address complex business challenges
- Deep expertise in networking and systems, including TCP/IP, DNS, load balancing, firewalls, and VPN technologies; strong Linux performance tuning and system-level troubleshooting skills
- Certifications a plus: AWS Certified SysOps Administrator or AWS Professional, Certified Kubernetes Administrator (CKA), Terraform Associate (or equivalent)
- Collaborative leader with a proven track record mentoring junior engineers, driving SRE best-practice adoption across teams, and communicating clearly to both technical and non-technical stakeholders (including presentations)
- Experience in handling critical incident and change management – be part of critical incident taskforce call.
- Familiarity of agile practices – preferably, scrum and Kanban
- ...Site Reliability Engineer III There's nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability...SuggestedWork at office
- ...Site Reliability Engineer There are NO limits to your career: come shape the future and be part of a truly unique global culture at OutSystems... ...here are your key responsibilities and duties: Lead and onboard services and teams to the reliability tenets;...SuggestedImmediate startRemote workWorldwide
- ...Senior Lead Site Reliability Engineer Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally gifted professionals and position yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at JPMorgan...Suggested
$200k - $260k
...Site Reliability Engineering Lead Glean is seeking a Site Reliability Engineering Lead to foster a culture of engineering excellence, drive technical strategy, and develop a high-performing, collaborative team. Your role is pivotal in ensuring our services meet stringent...SuggestedWork at officeHome office$217.57k - $260k
...explicitly states otherwise, all roles are on-site five days per week at one of our... ...here. Role Overview The Staff Site Reliability Engineer, Infrastructure role is building a high... ...experience operating at this scale and leading infrastructure through significant...SuggestedFull timeTemporary workWork at officeRemote workFlexible hoursShift work- ...Team: Infra Reliability • SF Bay Area / Remote (US) You'll own the GPU infrastructure Luma's research and product run on - thousands... ...hands-on, close-to-the-metal role for a first-principles Linux engineer. You'll be the final escalation for the hardest GPU, networking...Work experience placementRemote work
- ...world running. Location: 5 on-site days a week in Sunnyvale, CA... ...Our Team's Vision: Our Engineering team is shaping the future of... ...an experienced Senior Site Reliability Engineer (SRE) with a strong... ...and infrastructure updates Lead incident response and resolution...Work experience placementImmediate start
$230k - $250k
...Site Reliability Engineer Forward was founded in 2013 by four Stanford Ph.D.s, building the industry's first network digital twin: a mathematically... ...team always knows what's happening before customers do Lead incident response: on-call rotations, runbooks, post-...Night shift$65 - $85 per hour
...Site Reliability Engineer Sustainable Talent is partnering with a global leader who's been transforming computer graphics, PC gaming, and accelerated computing for over 25 years. We are looking for a Site Reliability Engineer to support our client's team based out of...Full timeContract workWorldwide- • Design, implement, and maintain complex data systems supporting millions of customers with Cloud Native principles and best practices to ensure highly available, secure, performant and scalable database systems • Build and maintain CI/CD pipelines in Jenkins • Build...
$160k - $240k
...another millions of times a day - quickly, reliably, and securely. Any time you swipe your... ...at Fiserv. Job Title Senior Site Reliability Engineer What does a successful Site... ...Participate in on-call rotations and lead incident response activities; run and...- ...Job Title : Senior Site Reliability Engineer Location : Santa Clara, CA Contract ENGAGEMENT SUMMARY The Candidate will... ...cluster dependencies, and shared infrastructure components. Lead or support incident triage for service degradation...Contract work
- ...Site Reliability Engineer (SRE) Share Contractual Sunnyvale, CA PDT - 8450 8-10 Overview: *Must have Apple experience* • At least 8+ years in a Reliability Engineering, DevOps or infrastructure focused role • Advanced experience with programming languages (Python...
- ...Senior Site Reliability Engineer LeanData helps the world's fastest-growing companies automate, simplify, and accelerate revenue. We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly...Full timeWork at officeFlexible hours2 days per week
$132.6k - $214.5k
...As part of this role, you will collaborate closely with our engineering teams to develop innovative solutions that provide clear and... ...team to influence the operability of the product and ensure the reliability and availability of our services. Qualifications DevOps...Full timeWork at officeVisa sponsorshipWork visa- ...Site Reliability Engineer (SRE) Location: Santa Clara Valley (Cupertino), California, Hybrid. Duration: 6+ Months Job Description Deploy, support and monitor new and existing services, platforms, and application stacks. Use scale testing to measure, tune...
- ...Senior Site Reliability Engineer Location: Remote Duration: 12 month contract to start IV Process: 1-3 Round IV process International... ..., or service operations and quality • Participate in, or lead design reviews with peers and stakeholders to decide...Contract workLocal areaRemote work
$128k - $216k
...consumers to one another millions of times a day - quickly, reliably, and securely. Any time you swipe your credit card, pay... ...a global scale, come make a difference at Fiserv. Sr. Site Reliability Engineer About Clover Clover is a pioneer in the fintech space...Worldwide$170k - $200k
...Site Reliability Engineer We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high...Full timeWorldwide$28 per hour
...Position Title: Lead Premium Supervisor Location: Stanford University Athletics Pay Range : $28.00 We Make Applying Easy! Want to apply to this job via text messaging? Text JOB to 75000 and search requisition ID number 1567464 . The advertised...Full timeWork at officeRemote workFlexible hours$61k - $101k
...Requirements: We require formal training or certification in site reliability engineering, along with 5+ years of hands-on experience. We need... ..., communities of practice, guilds, and conferences. We lead reuse-first adoption of AI-assisted reliability workflows across...Full time$255.7k - $300k
...Manager, Software Engineer, Site Reliability Engineering Share Manager, Software Engineer, Site Reliability Engineering ~ link Copy link... ...hybrid schedule as per Google policy. Responsibilities Lead a team of engineers to maintain service uptime while managing...Full timeWork at office- ...Job Description Responsibility • Lead the effort of global expansion of Huobi... ...infrastructure. • Work with engineering teams to make sure new features and changes... ...Constantly improve our system performance and reliability through better tools, process and...Worldwide
- ...infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,... ...and deployment workflows for accuracy and reliability. Work with AWS, Azure, GCP,... ...Azure DevOps Cloud Infrastructure Site Reliability Engineering (SRE) Platform...Remote jobFor contractors
- ...Overview We are seeking a highly motivated Systems Reliability Engineer (SRE) to lead the design and implementation of operational excellence across... ...supporting sensitive and cleared workforces. The Site Reliability Engineer (SRE) - SecOps will embrace our commitment...For contractorsWork at officeFlexible hours
$165k - $190k
Obsidian Security is the leading SaaS security platform, trusted by global enterprises... ...DevOps/SRE team at Obsidian ensures that engineering excellence translates into stable,... ...complex challenges around scalability, reliability, observability, and cost efficiencyCollaborate...Work from home$165k - $280k
...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s most...Permanent employmentTemporary workWorldwideWeekend work$115.5k - $189.75k
...driving by turning on-road signals and incidents into actionable engineering insights. The Release & Triage Tooling sub-team builds AI-... ...end-to-end workflows. Design and implement scalable, reliable internal services used by release and triage teams, ensuring maintainability...Full timeTemporary workWork at officeFlexible hours$200k - $247k
...AI Enablement Lead Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since... ...business processes. The Data Intelligence team is the strategic engine driving the democratization of AI across Waymo G&A. Operating at...Full timeRemote work$22 - $26 per hour
Job Summary: Opens and closes the store in the absence of store management, including all required systems startups, required cash handling, and ensuring the floor and stock room are ready for the business day. Responsible for opening back door of store for deliveries...Hourly payWork experience placementSeasonal workLocal areaFlexible hoursShift workAfternoon shift
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- lead operating engineer Palo Alto, CA
- lead engineer Palo Alto, CA
- site reliability engineer Palo Alto, CA
- junior website developer Palo Alto, CA
- construction site safety Palo Alto, CA
- site services specialist Palo Alto, CA
- website content developer Palo Alto, CA
- on-site clinical research associate (traveling/remote) Palo Alto, CA
- historic site Palo Alto, CA
- official site Palo Alto, CA




