Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineer

Chase

Lead Site Reliability Engineer

Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.

As a Lead Site Reliability Engineer at JPMorgan Chase within the Enterprise technology, engineering services and platform team, you hold a leadership role in your team, demonstrate strong knowledge across multiple technical domains, and advise others on the technical and business issues facing them. Take lead and conduct resiliency design reviews, break up complex problems into digestible work for other engineers, act as a technical lead for medium to large-sized products, and provide advice and mentoring to other engineers.

Job responsibilities

  • Guides and assists others in the areas of building appropriate level designs and gaining consensus from peers where appropriate
  • Collaborates with other software engineers and teams to design and implement deployment approaches using automated continuous integration and continuous delivery pipelines
  • Collaborates with other software engineers and teams to design, develop, test, and implement availability, reliability, scalability, and solutions in their applications
  • Implements infrastructure, configuration, and network as code for the applications and platforms in your remit
  • Collaborates with technical experts, key stakeholders, and team members to resolve complex problems
  • Understands service level indicators and utilizes service level objectives to proactively resolve issues before they impact customers
  • Supports the adoption of site reliability engineering best practices within your team
  • Production 24*7 support for business-critical applications
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
  • Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls.

Required qualifications, capabilities, and skills

  • Formal training or certification on site reliability engineering concepts and 5+ years applied experience
  • Proficient in site reliability engineering (SRE) culture and principles, with experience implementing SRE practices within applications and platforms; strong observability background including white/black-box monitoring, SLO-based alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, and similar.
  • Proficient in at least one programming language (e.g., Python, Java/Spring Boot,.NET) with strong knowledge of software applications and technical processes within a technical discipline such as cloud, artificial intelligence, Android, or related areas.
  • Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.
  • Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations.
  • Hands-on experience with CI/CD tooling (e.g., Jenkins, GitLab) and infrastructure automation using Terraform to build reliable, repeatable delivery pipelines.
  • Strong familiarity with containers and orchestration platforms (Docker, Kubernetes, ECS), including deploying, scaling, and operating containerized services in production.
  • Proven ability to troubleshoot and resolve common networking issues (DNS, TCP/IP, routing, TLS, load balancing), applying structured debugging to restore service quickly.
  • Collaborative, proactive team contributor: communicates clearly and persuasively with minimal supervision, identifies roadblocks early, learns new technologies quickly, and has experience with event streaming platforms such as Kafka.

Preferred qualifications, capabilities, and skills

  • Ability to identify new technologies and relevant solutions to ensure design constraints are met by the software team
  • Proven track record of initiating and executing ideas that address complex business challenges
  • Deep expertise in networking and systems, including TCP/IP, DNS, load balancing, firewalls, and VPN technologies; strong Linux performance tuning and system-level troubleshooting skills
  • Certifications a plus: AWS Certified SysOps Administrator or AWS Professional, Certified Kubernetes Administrator (CKA), Terraform Associate (or equivalent)
  • Collaborative leader with a proven track record mentoring junior engineers, driving SRE best-practice adoption across teams, and communicating clearly to both technical and non-technical stakeholders (including presentations)
  • Experience in handling critical incident and change management – be part of critical incident taskforce call.
  • Familiarity of agile practices – preferably, scrum and Kanban
Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineer in Palo Alto, CA vacancy
  •  ...Site Reliability Engineer III There's nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability... 
    Suggested
    Work at office

    Chase

    Palo Alto, CA
    3 days ago
  •  ...Site Reliability Engineer There are NO limits to your career: come shape the future and be part of a truly unique global culture at OutSystems...  ...here are your key responsibilities and duties: Lead and onboard services and teams to the reliability tenets;... 
    Suggested
    Immediate start
    Remote work
    Worldwide

    OutSystems

    Menlo Park, CA
    3 days ago
  •  ...Senior Lead Site Reliability Engineer Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally gifted professionals and position yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at JPMorgan... 
    Suggested

    Hackajob

    Palo Alto, CA
    22 hours ago
  • $200k - $260k

     ...Site Reliability Engineering Lead Glean is seeking a Site Reliability Engineering Lead to foster a culture of engineering excellence, drive technical strategy, and develop a high-performing, collaborative team. Your role is pivotal in ensuring our services meet stringent... 
    Suggested
    Work at office
    Home office

    Glean - Mountain View, CA, US

    Mountain View, CA
    2 days ago
  • $217.57k - $260k

     ...explicitly states otherwise, all roles are on-site five days per week at one of our...  ...here. Role Overview The Staff Site Reliability Engineer, Infrastructure role is building a high...  ...experience operating at this scale and leading infrastructure through significant... 
    Suggested
    Full time
    Temporary work
    Work at office
    Remote work
    Flexible hours
    Shift work

    ID.me

    Mountain View, CA
    3 days ago
  •  ...Team: Infra Reliability • SF Bay Area / Remote (US) You'll own the GPU infrastructure Luma's research and product run on - thousands...  ...hands-on, close-to-the-metal role for a first-principles Linux engineer. You'll be the final escalation for the hardest GPU, networking... 
    Work experience placement
    Remote work

    Luma

    Redwood City, CA
    4 days ago
  •  ...world running. Location: 5 on-site days a week in Sunnyvale, CA...  ...Our Team's Vision: Our Engineering team is shaping the future of...  ...an experienced Senior Site Reliability Engineer (SRE) with a strong...  ...and infrastructure updates Lead incident response and resolution... 
    Work experience placement
    Immediate start

    Illumio

    Sunnyvale, CA
    22 hours ago
  • $230k - $250k

     ...Site Reliability Engineer Forward was founded in 2013 by four Stanford Ph.D.s, building the industry's first network digital twin: a mathematically...  ...team always knows what's happening before customers do Lead incident response: on-call rotations, runbooks, post-... 
    Night shift

    Forward Networks Inc

    Santa Clara, CA
    4 days ago
  • $65 - $85 per hour

     ...Site Reliability Engineer Sustainable Talent is partnering with a global leader who's been transforming computer graphics, PC gaming, and accelerated computing for over 25 years. We are looking for a Site Reliability Engineer to support our client's team based out of... 
    Full time
    Contract work
    Worldwide

    Sustainable Talent

    Santa Clara, CA
    1 day ago
  • • Design, implement, and maintain complex data systems supporting millions of customers with Cloud Native principles and best practices to ensure highly available, secure, performant and scalable database systems • Build and maintain CI/CD pipelines in Jenkins • Build...

    United IT Solutions

    Mountain View, CA
    2 days ago
  • $160k - $240k

     ...another millions of times a day - quickly, reliably, and securely. Any time you swipe your...  ...at Fiserv. Job Title Senior Site Reliability Engineer What does a successful Site...  ...Participate in on-call rotations and lead incident response activities; run and... 

    Fiserv

    Sunnyvale, CA
    4 days ago
  •  ...Job Title : Senior Site Reliability Engineer Location : Santa Clara, CA Contract ENGAGEMENT SUMMARY The Candidate will...  ...cluster dependencies, and shared infrastructure components. Lead or support incident triage for service degradation... 
    Contract work

    VDart

    Santa Clara, CA
    2 days ago
  •  ...Site Reliability Engineer (SRE) Share Contractual Sunnyvale, CA PDT - 8450 8-10 Overview: *Must have Apple experience* • At least 8+ years in a Reliability Engineering, DevOps or infrastructure focused role • Advanced experience with programming languages (Python... 

    Purple Drive

    Sunnyvale, CA
    3 days ago
  •  ...Senior Site Reliability Engineer LeanData helps the world's fastest-growing companies automate, simplify, and accelerate revenue. We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly... 
    Full time
    Work at office
    Flexible hours
    2 days per week

    LeanData

    Santa Clara, CA
    3 days ago
  • $132.6k - $214.5k

     ...As part of this role, you will collaborate closely with our engineering teams to develop innovative solutions that provide clear and...  ...team to influence the operability of the product and ensure the reliability and availability of our services. Qualifications DevOps... 
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks

    Santa Clara, CA
    2 days ago
  •  ...Site Reliability Engineer (SRE) Location: Santa Clara Valley (Cupertino), California, Hybrid. Duration: 6+ Months Job Description Deploy, support and monitor new and existing services, platforms, and application stacks. Use scale testing to measure, tune... 

    Zortech Solutions

    Cupertino, CA
    4 days ago
  •  ...Senior Site Reliability Engineer Location: Remote Duration: 12 month contract to start IV Process: 1-3 Round IV process International...  ..., or service operations and quality • Participate in, or lead design reviews with peers and stakeholders to decide... 
    Contract work
    Local area
    Remote work

    My3Tech Inc

    Sunnyvale, CA
    2 days ago
  • $128k - $216k

     ...consumers to one another millions of times a day - quickly, reliably, and securely. Any time you swipe your credit card, pay...  ...a global scale, come make a difference at Fiserv. Sr. Site Reliability Engineer About Clover Clover is a pioneer in the fintech space... 
    Worldwide

    BentoBox

    Sunnyvale, CA
    4 days ago
  • $170k - $200k

     ...Site Reliability Engineer We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high... 
    Full time
    Worldwide

    Edelman

    Sunnyvale, CA
    2 days ago
  • $28 per hour

     ...Position Title: Lead Premium Supervisor  Location: Stanford University Athletics Pay Range : $28.00 We Make Applying Easy! Want to apply to this job via text messaging? Text JOB to 75000  and search requisition ID number 1567464 . The advertised... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Compass Group

    Stanford, CA
    12 days ago
  • $61k - $101k

     ...Requirements: We require formal training or certification in site reliability engineering, along with 5+ years of hands-on experience. We need...  ..., communities of practice, guilds, and conferences. We lead reuse-first adoption of AI-assisted reliability workflows across... 
    Full time

    J.P. Morgan

    Palo Alto, CA
    8 days ago
  • $255.7k - $300k

     ...Manager, Software Engineer, Site Reliability Engineering Share Manager, Software Engineer, Site Reliability Engineering ~ link Copy link...  ...hybrid schedule as per Google policy. Responsibilities Lead a team of engineers to maintain service uptime while managing... 
    Full time
    Work at office

    Google Inc.

    Sunnyvale, CA
    4 days ago
  •  ...Job Description Responsibility •       Lead the effort of global expansion of Huobi...  ...infrastructure. •       Work with engineering teams to make sure new features and changes...  ...Constantly improve our system performance and reliability through better tools, process and... 
    Worldwide

    Cryptoware Technologies Inc

    Santa Clara, CA
    17 days ago
  •  ...infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,...  ...and deployment workflows for accuracy and reliability. Work with AWS, Azure, GCP,...  ...Azure DevOps Cloud Infrastructure Site Reliability Engineering (SRE) Platform... 
    Remote job
    For contractors

    YO AI Labs

    Palo Alto, CA
    18 days ago
  •  ...Overview We are seeking a highly motivated Systems Reliability Engineer (SRE) to lead the design and implementation of operational excellence across...  ...supporting sensitive and cleared workforces. The Site Reliability Engineer (SRE) - SecOps will embrace our commitment... 
    For contractors
    Work at office
    Flexible hours

    Arkenstone Defense

    Menlo Park, CA
    18 days ago
  • $165k - $190k

    Obsidian Security is the leading SaaS security platform, trusted by global enterprises...  ...DevOps/SRE team at Obsidian ensures that engineering excellence translates into stable,...  ...complex challenges around scalability, reliability, observability, and cost efficiencyCollaborate... 
    Work from home

    Obsidian Security

    Palo Alto, CA
    6 days ago
  • $165k - $280k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s most... 
    Permanent employment
    Temporary work
    Worldwide
    Weekend work

    SpaceX

    Palo Alto, CA
    5 days ago
  • $115.5k - $189.75k

     ...driving by turning on-road signals and incidents into actionable engineering insights.  The Release & Triage Tooling sub-team builds AI-...  ...end-to-end workflows. Design and implement scalable, reliable internal services used by release and triage teams, ensuring maintainability... 
    Full time
    Temporary work
    Work at office
    Flexible hours

    Woven

    Palo Alto, CA
    22 hours ago
  • $200k - $247k

     ...AI Enablement Lead Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since...  ...business processes. The Data Intelligence team is the strategic engine driving the democratization of AI across Waymo G&A. Operating at... 
    Full time
    Remote work

    Waymo

    Mountain View, CA
    1 day ago
  • $22 - $26 per hour

    Job Summary: Opens and closes the store in the absence of store management, including all required systems startups, required cash handling, and ensuring the floor and stock room are ready for the business day. Responsible for opening back door of store for deliveries...
    Hourly pay
    Work experience placement
    Seasonal work
    Local area
    Flexible hours
    Shift work
    Afternoon shift

    Walgreens

    Mountain View, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!