Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineer

Chase

Lead Site Reliability Engineer

Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.

As a Lead Site Reliability Engineer at JPMorgan Chase within the Enterprise technology, engineering services and platform team, you hold a leadership role in your team, demonstrate strong knowledge across multiple technical domains, and advise others on the technical and business issues facing them. Take lead and conduct resiliency design reviews, break up complex problems into digestible work for other engineers, act as a technical lead for medium to large-sized products, and provide advice and mentoring to other engineers.

Job responsibilities

  • Guides and assists others in the areas of building appropriate level designs and gaining consensus from peers where appropriate
  • Collaborates with other software engineers and teams to design and implement deployment approaches using automated continuous integration and continuous delivery pipelines
  • Collaborates with other software engineers and teams to design, develop, test, and implement availability, reliability, scalability, and solutions in their applications
  • Implements infrastructure, configuration, and network as code for the applications and platforms in your remit
  • Collaborates with technical experts, key stakeholders, and team members to resolve complex problems
  • Understands service level indicators and utilizes service level objectives to proactively resolve issues before they impact customers
  • Supports the adoption of site reliability engineering best practices within your team
  • Production 24*7 support for business-critical applications
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
  • Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls.

Required qualifications, capabilities, and skills

  • Formal training or certification on site reliability engineering concepts and 5+ years applied experience
  • Proficient in site reliability engineering (SRE) culture and principles, with experience implementing SRE practices within applications and platforms; strong observability background including white/black-box monitoring, SLO-based alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, and similar.
  • Proficient in at least one programming language (e.g., Python, Java/Spring Boot,.NET) with strong knowledge of software applications and technical processes within a technical discipline such as cloud, artificial intelligence, Android, or related areas.
  • Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.
  • Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations.
  • Hands-on experience with CI/CD tooling (e.g., Jenkins, GitLab) and infrastructure automation using Terraform to build reliable, repeatable delivery pipelines.
  • Strong familiarity with containers and orchestration platforms (Docker, Kubernetes, ECS), including deploying, scaling, and operating containerized services in production.
  • Proven ability to troubleshoot and resolve common networking issues (DNS, TCP/IP, routing, TLS, load balancing), applying structured debugging to restore service quickly.
  • Collaborative, proactive team contributor: communicates clearly and persuasively with minimal supervision, identifies roadblocks early, learns new technologies quickly, and has experience with event streaming platforms such as Kafka.

Preferred qualifications, capabilities, and skills

  • Ability to identify new technologies and relevant solutions to ensure design constraints are met by the software team
  • Proven track record of initiating and executing ideas that address complex business challenges
  • Deep expertise in networking and systems, including TCP/IP, DNS, load balancing, firewalls, and VPN technologies; strong Linux performance tuning and system-level troubleshooting skills
  • Certifications a plus: AWS Certified SysOps Administrator or AWS Professional, Certified Kubernetes Administrator (CKA), Terraform Associate (or equivalent)
  • Collaborative leader with a proven track record mentoring junior engineers, driving SRE best-practice adoption across teams, and communicating clearly to both technical and non-technical stakeholders (including presentations)
  • Experience in handling critical incident and change management – be part of critical incident taskforce call.
  • Familiarity of agile practices – preferably, scrum and Kanban
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineer in Palo Alto, CA vacancy
  • $165k - $280k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s most... 
    Suggested
    Permanent employment
    Temporary work
    Worldwide
    Weekend work

    SpaceX

    Palo Alto, CA
    2 days ago
  • $167.7k - $245.2k

     ...requiring approximately 2 days per week on-site at Cisco offices in either San Francisco...  ...AI agents behave as intended, improving reliability and reducing risks. This unified...  ...and control.As a Senior Site Reliability Engineer (SRE), you will build, operate, and continuously... 
    Suggested
    Full time
    Temporary work
    Local area
    Flexible hours
    2 days per week

    CISCO Systems

    Palo Alto, CA
    1 day ago
  • $165k - $190k

    Obsidian Security is the leading SaaS security platform, trusted by global enterprises...  ...DevOps/SRE team at Obsidian ensures that engineering excellence translates into stable,...  ...complex challenges around scalability, reliability, observability, and cost efficiencyCollaborate... 
    Suggested
    Work from home

    Obsidian Security

    Palo Alto, CA
    4 days ago
  • $210.6k - $305.1k

     ...powered assurance insights within Cisco’s leading Networking, Security, Collaboration, and...  ...:  You have led a distributed team of 5+ engineers, can demonstrate strong technical vision...  ...insurance. Please see the Cisco careers site to discover more benefits and perks. Employees... 
    Suggested
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    Los Altos, CA
    4 days ago
  • Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally gifted professionals and position yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Infrastructure Platforms... 
    Suggested

    JP Morgan Chase

    Palo Alto, CA
    2 days ago
  • $186.9k - $267.7k

     ...approximately 2 days per week on-site at Cisco offices in either...  ...as intended, improving reliability and reducing risks. This unified...  ....As a Staff Site Reliability Engineer (SRE), you will provide technical...  ...-term reliability strategy, lead major infrastructure... 
    Full time
    Temporary work
    Local area
    Flexible hours
    2 days per week

    CISCO Systems

    Palo Alto, CA
    2 days ago
  • $232k - $263k

    Obsidian Security is the leading SaaS security platform, trusted by global enterprises...  ...growth and IPO readiness.Sr. Staff Site Reliability EngineerAs a Sr. Staff SRE at Obsidian...  ...strategic partner to DevOps and Platform Engineering leadership, shaping a unified... 
    Work from home

    Obsidian Security

    Palo Alto, CA
    1 day ago
  •  ...Site Reliability Engineer III There's nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability... 
    Work at office

    Chase

    Palo Alto, CA
    1 day ago
  •  ...Site Reliability Engineer There are NO limits to your career: come shape the future and be part of a truly unique global culture at OutSystems...  ...here are your key responsibilities and duties: Lead and onboard services and teams to the reliability tenets;... 
    Immediate start
    Remote work
    Worldwide

    OutSystems

    Menlo Park, CA
    1 day ago
  • $137.77k - $194.59k

     ...distributed team of roughly 80 scientists and engineers building and operating Rubin's petascale...  ...Your role: \n You will own the reliability and robustness of Rubin Observatory's...  ...nature of this position, SLAC is open to on-site, hybrid, and remote work options. \n \... 
    Remote work
    Flexible hours
    Night shift

    Stanford University

    Menlo Park, CA
    1 day ago
  •  ...Site Reliability Engineer, Data Platform - USDS Responsibilities Engage in and improve the whole lifecycle of service, from inception and design...  ...any immigration-related benefits. About USDS TikTok is the leading destination for short-form mobile video. U.S. Data Security... 

    Tik Tok

    Mountain View, CA
    2 days ago
  • $170k - $250k

     ...Site Reliability Engineer (SRE) Location: San Francisco, CA / Palo Alto, CA Company Stage of Funding: Growth-Stage AI Infrastructure Company ($80M Raised) Office Type: Onsite (4 Days Per Week) Salary: $170,000–$250,000 + Competitive Equity We're representing a rapidly... 
    Work at office
    Visa sponsorship
    Flexible hours

    Recruiting from Scratch

    Palo Alto, CA
    2 days ago
  • $170k - $230k

     ...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About Mithril Mithril is an AI infrastructure platform built to make...  ...GPU compute more accessible and affordable for the world's leading enterprises, AI startups, and the AI research community, including... 
    Work at office
    Local area
    1 day per week

    Mithril

    Palo Alto, CA
    3 days ago
  •  ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.As a Lead Site Reliability Engineer at JPMorgan Chase within the Network Product, you hold a leadership role in your team, demonstrate strong knowledge... 

    JP Morgan Chase

    Palo Alto, CA
    2 days ago
  • $200k - $260k

     ...enterprise trust, as we bring Work AI to every employee, in every company. About the Role: Glean is seeking a Site Reliability Engineering Lead to foster a culture of engineering excellence, drive technical strategy, and develop a high-performing, collaborative... 
    Work at office
    Home office
    Flexible hours

    Glean.info

    Mountain View, CA
    3 days ago
  •  ...professionals for this role. JOB DESCRIPTION Elevate your engineering prowess to unprecedented levels by joining a team of...  ...professionals and position yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the... 

    J.P. Morgan

    Palo Alto, CA
    4 days ago
  • $100k - $200k

    OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about... 
    Full time

    OPPO

    Palo Alto, CA
    3 days ago
  •  ...that's more connected, more intelligent, more sustainable for everyone. Role Summary We are seeking an experienced Site Reliability Engineer to help design, build, and operate the infrastructure that underpins the build pipelines that allow our companies to... 
    Full time
    Contract work

    Rivian and Volkswagen Group Technologies

    Palo Alto, CA
    2 days ago
  • $148k - $235.75k

     ...on the world.Join our team of innovative engineers who are building an AI Data Center AIOps...  ...that turns raw, high-volume telemetry into reliable, job-centric insights and automation for...  ...canary checks, post-deploy validation), and lead rollbacks/remediations when needed.Lead... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is... 
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    2 days ago
  • $170k - $200k

    We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance,... 
    Full time
    Worldwide

    Fortinet

    Sunnyvale, CA
    4 days ago
  • $152k - $241.5k

     ...technology—and amazing people. NVIDIA is leading the way in groundbreaking developments...  ...automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-...  ..., Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $230k - $250k

     ...network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change...  ...how things have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on"... 
    Night shift

    Forward Networks

    Santa Clara, CA
    2 days ago
  • $101k - $161k

     ...several prestigious awards, such as Best Engineering Team, Best Company for Diversity,...  ...DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s...  ...the chance to be drive, develop, and lead projects in any of the following areas:... 

    Arista Networks

    Santa Clara, CA
    10 hours ago
  • $140k - $230k

     ...Platforms and Product /Full-time /HybridZoox is seeking a Site Reliability Engineer to help ensure the availability, performance, and resilience...  ...streamline deployment processes, and drive automation initiatives.Lead incident resolution: You will conduct thorough root cause... 
    Full time

    Zoox

    Foster, CA
    3 days ago
  • Lead Cloud ArchitectCooley is seeking a Lead Cloud Architect to join the Innovation team.About Cooley: Cooley is a global law firm...  ...Architect is responsible for owning the technical architecture and engineering standards for a greenfield SOC2-compliant SaaS platform built... 
    Full time
    Work at office
    Local area
    Immediate start
    Remote work
    Work from home
    Worldwide
    Weekend work

    Cooley

    Palo Alto, CA
    4 days ago
  • $184k - $287.5k

    At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability. This demanding position merges software and systems engineering efforts to guarantee flawless service operation... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $207k - $301k

     ...influential relationships with multiple stakeholders across the Site Reliability Engineering and Developer organizations.Serve as an expert on...  ...knowledge related to rate limiting or sharding.Develop plans and lead projects on evolving our production systems and their... 

    Google

    Sunnyvale, CA
    10 hours ago
  • $65 - $85 per hour

     ...Site Reliability Engineer Sustainable Talent is partnering with a global leader who's been transforming computer graphics, PC gaming, and accelerated computing for over 25 years. We are looking for a Site Reliability Engineer to support our client's team based out of... 
    Full time
    Contract work
    Worldwide

    Sustainable Talent

    Santa Clara, CA
    4 days ago
  • $230k - $250k

     ...Site Reliability Engineer Forward is transforming how the world's most complex networks are managed and secured. Founded in 2013 by four Stanford...  ...team always knows what's happening before customers do Lead incident response: on-call rotations, runbooks, post-... 
    Night shift

    Forward Networks Inc

    Santa Clara, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!