Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineering - Network

JP Morgan Chase

Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.As a Lead Site Reliability Engineer at JPMorgan Chase within the Network Product, you hold a leadership role in your team, demonstrate strong knowledge across multiple technical domains, and advise others on the technical and business issues facing them. Take lead and conduct resiliency design reviews, break up complex problems into digestible work for other engineers, act as a technical lead for medium to large-sized products, and provide advice and mentoring to other engineers.Job responsibilitiesApplies network reliability principles (Permit to Operate, FMEA, operational readiness), balancing feature delivery, efficiency, and stability.Partners with network engineering domains (Datacenter, Firewall, Proxies, DMZ, Load Balancing, etc.) and Lines of Business to align goals and outcomes.Drives adoption of reliability best practices and observability, demonstrating impact through stability/reliability metrics.; Bridges Engineering, Operations, DevOps, and customers to build resilient, scalable, and secure network services.Provides Tier-3 network support, leading major incident response, rapid restoration, RCA, and follow-through on corrective actions.Leads reliability and stability initiatives using data-driven analysis to improve service levels and reduce recurring failure modes.Defines SLI/SLOs and error budgets with stakeholders and customers, ensuring measurable performance targets and trade-off clarity. Identifies and removes technical bottlenecks within core domains of expertise, proactively preventing reliability and capacity risks.Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.Runs blameless, data-driven post-mortems and debriefs, converting learnings (successes and failures) into actionable improvements.Fosters continuous improvement and strong knowledge sharing, soliciting real-time feedback, avoiding duplicated work, and promoting innovation via internal communities.Produces and packages thought leadership with specialists/product/engineering teams—documenting best practices and lessons learned for internal assets and industry forums/conferences.Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls.Required qualifications, capabilities, and skillsFormal training or certification in network engineering concepts and 5+ years of applied experience.10+ years of experience leading technologists to manage and solve complex technical items within your domain of expertise.Advanced proficiency in network reliability engineering, including Permit to Operate, FMEA, and operational readiness processes.Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations.Experience leading technologists to manage and solve complex network issues at a firmwide level.Ability to influence team culture by championing innovation and change for success.Proficiency in SD-WAN, cloud platforms (AWS, Azure, etc.), and major network technologies (Palo Alto, Juniper, F5, Broadcom, Arista, Cisco, etc.).Proficiency in observability and monitoring tools such as Grafana, SevOne, Prometheus, Kibana, ThousandEyes, and Splunk.Preferred qualifications, capabilities, and skillsCCIE, Load-balancing, SD-WAN, Observability tools, eBPF, Cloud certsDemonstrated proficiency in troubleshooting and supporting complex networking environments, including Tier-3 operational support for major incidents.Experience with continuous integration and delivery tools (e.g., Jenkins, GitLab, Terraform, etc.).Experience in scalable networking design, including high availability, redundancy, failover, and load balancing.Experience troubleshooting networking protocols such as TCP/IP, and BGP.Experience in customer-facing migration, including service discovery, assessment, planning, execution, and operations.This position is subject to Section 19 of the Federal Deposit Insurance Act. As such, an employment offer for this position is contingent on JPMorganChase’s review of criminal conviction history, including pretrial diversions or program entries. JPMorganChase, one of the oldest financial institutions, offers innovative financial solutions to millions of consumers, small businesses and many of the world’s most prominent corporate, institutional and government clients under the J.P. Morgan and Chase brands. Our history spans over 200 years and today we are a leader in investment banking, consumer and small business banking, commercial banking, financial transaction processing and asset management. We offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may receive commission-based pay and/or discretionary incentive compensation, paid in the form of cash and/or forfeitable equity, awarded in recognition of individual achievements and contributions. We also offer a range of benefits and programs to meet employee needs, based on eligibility. These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare, tuition reimbursement, mental health support, financial coaching and more. Additional details about total compensation and benefits will be provided during the hiring process. We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants’ and employees’ religious practices and beliefs, as well as mental health or physical disability needs. Visit our FAQs for more information about requesting an accommodation.JPMorgan Chase & Co. is an Equal Opportunity Employer, including Disability/VeteransOur professionals in our Corporate Functions cover a diverse range of areas from finance and risk to human resources and marketing. Our corporate teams are an essential part of our company, ensuring that we’re setting our businesses, clients, customers and employees up for success.Full timePosting Date: 2026-06-16

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineering - Network in Palo Alto, CA vacancy
  • Lead Cloud ArchitectCooley is seeking a Lead Cloud Architect to join the Innovation team...  ...owning the technical architecture and engineering standards for a greenfield SOC2-...  ...Extensive hands-on background with AWS (IAM, networking, Organizations) and Terraform at scaleExperience... 
    Suggested
    Full time
    Work at office
    Local area
    Immediate start
    Remote work
    Work from home
    Worldwide
    Weekend work

    Cooley

    Palo Alto, CA
    2 days ago
  •  ...everyone. Role Summary We are seeking an experienced Site Reliability Engineer to help design, build, and operate the infrastructure that...  ...at scale. ~ Deep understanding of Kubernetes internals, networking, storage and service mesh architectures. ~ Proven track... 
    Suggested
    Full time
    Contract work

    Rivian and Volkswagen Group Technologies

    Palo Alto, CA
    5 days ago
  • $165k - $280k

     ...with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building...  ...HBase, HDFS, FlinkExperience troubleshooting hardware and network-layer issuesProgramming experience in Python, C#, Java,... 
    Suggested
    Permanent employment
    Temporary work
    Worldwide
    Weekend work

    SpaceX

    Palo Alto, CA
    4 hours ago
  • $140k - $205k

    Senior Technology Site Reliability EngineerCooley is seeking a Senior Site Reliability Engineer to join the Infrastructure & Development Operations team.Position summary: The Senior Technology Site Reliability Engineer (“SRE”) is responsible for ensuring the reliability... 
    Suggested
    Full time
    Temporary work
    Work at office
    Flexible hours
    Weekend work

    Cooley

    Palo Alto, CA
    5 days ago
  • $165k - $190k

    Obsidian Security is the leading SaaS security platform, trusted by global enterprises...  ...DevOps/SRE team at Obsidian ensures that engineering excellence translates into stable,...  ...complex challenges around scalability, reliability, observability, and cost efficiencyCollaborate... 
    Suggested
    Work from home

    Obsidian Security

    Palo Alto, CA
    2 days ago
  •  ...role focuses on translating requirements into scalable designs, leading architectural initiatives and guiding delivery teams to high-...  ...solutions meet business outcomes and best practices across compute, network, and storage. A relevant VMware certification and extensive VCF... 
    Remote work

    Entelligence

    Palo Alto, CA
    5 days ago
  • $232k - $263k

    Obsidian Security is the leading SaaS security platform, trusted by global enterprises...  ...growth and IPO readiness.Sr. Staff Site Reliability EngineerAs a Sr. Staff SRE at Obsidian...  ...strategic partner to DevOps and Platform Engineering leadership, shaping a unified... 
    Work from home

    Obsidian Security

    Palo Alto, CA
    4 days ago
  • $210.6k - $305.1k

     ...digital experiences across every network - even the ones they don’t...  ...insights within Cisco’s leading Networking, Security, Collaboration...  ...led a distributed team of 5+ engineers, can demonstrate strong...  ...Please see the Cisco careers site to discover more benefits and... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    Los Altos, CA
    2 days ago
  • Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally...  ...yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at...  ...community, and continues to expand network and leads evaluation sessions with... 

    JP Morgan Chase

    Palo Alto, CA
    4 hours ago
  • $186.9k - $267.7k

     ...approximately 2 days per week on-site at Cisco offices in...  ...intended, improving reliability and reducing risks....  ...Site Reliability Engineer (SRE), you will provide...  ...reliability strategy, lead major infrastructure initiatives...  ..., databases, and networking.Drive automation to... 
    Full time
    Temporary work
    Local area
    Flexible hours
    2 days per week

    CISCO Systems

    Palo Alto, CA
    1 hour ago
  • $100k - $200k

     ...OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about... 
    Full time

    OPPO

    Palo Alto, CA
    1 day ago
  • $170k - $250k

     ...Site Reliability Engineer (SRE) Location: San Francisco, CA / Palo Alto, CA Company Stage of Funding: Growth-Stage AI Infrastructure Company...  ...distributed systems with a strong understanding of networking fundamentals. ~ Proficiency with Python and/or Go for building... 
    Work at office
    Visa sponsorship
    Flexible hours

    Recruiting from Scratch

    Palo Alto, CA
    5 days ago
  •  ...Site Reliability Engineer, Data Platform - USDS Responsibilities Engage in and improve the whole lifecycle...  ...cloud native technologies and networking. Experience developing tools and APIs...  ...related benefits. About USDS TikTok is the leading destination for short-form mobile... 

    Tik Tok

    Mountain View, CA
    5 days ago
  •  ...Senior Site Reliability Engineer Latitude AI is building the future of Ford's autonomy roadmap to...  ...Latitude team, you'll work alongside leading experts across machine learning and robotics...  ...operating system internals, TCP/IP networking, and storage subsystems Hands on... 
    Work at office
    Immediate start

    Latitude AI

    Palo Alto, CA
    4 days ago
  • $150k - $175k

     ...Site Reliability Engineer At ASAPP, our mission is simple: deliver the best AI-powered customer experience—faster than anyone else. To achieve...  ...our product engineers Design secure and performant networking solutions in our production systems What You'll Need... 
    Remote work

    ASAPP

    Mountain View, CA
    5 days ago
  • $137.77k - $194.59k

     ...distributed team of roughly 80 scientists and engineers building and operating Rubin's petascale...  .... Your role: You will own the reliability and robustness of Rubin Observatory's...  ...nature of this position, SLAC is open to on-site, hybrid, and remote work options.... 
    Remote work
    Flexible hours
    Night shift

    SLAC National Accelerator Laboratory

    Menlo Park, CA
    5 days ago
  •  ...role. JOB DESCRIPTION Elevate your engineering prowess to unprecedented levels by...  ...yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at JPMorgan...  ...community, and continues to expand network and leads evaluation sessions with... 

    J.P. Morgan

    Palo Alto, CA
    5 days ago
  •  ...together safely, transparently, and at scale. Join Wand in leading the Agentic Shift Wand is building a high-performing...  ...a hands‑on Head of SRE to establish, lead, and scale our Site Reliability Engineering function. This role combines strategic ownership with deep... 
    Shift work

    Wand AI

    Palo Alto, CA
    5 days ago
  • $170k - $230k

     ...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About Mithril Mithril is...  ...accessible and affordable for the world's leading enterprises, AI startups, and the AI...  ...practical understanding of cloud networking fundamentals (VPC, DNS, load balancing... 
    Work at office
    Local area
    1 day per week

    Mithril

    Palo Alto, CA
    4 days ago
  •  ...Acryl Data seeks a Site Reliability Engineering (SRE) Tech Lead to enhance the reliability and scalability of its DataHub platform. The role involves leading infrastructure design, optimizing system performance, and driving continuous improvement across cloud deployments... 

    Acryl Data

    Palo Alto, CA
    5 days ago
  •  ...ActiveHours is looking for an experienced DevOps Engineer to enhance platform automation and site reliability in a collaborative environment. The role involves automating key systems, improving visibility through metrics, and troubleshooting critical problems. You'll... 

    ActiveHours

    Palo Alto, CA
    5 days ago
  • $168k - $264.5k

     ...impact on the world.Are you ready to be part of something outstanding? NVIDIA's Digital Marketing Organization seeks a senior Site Reliability Engineer (SRE) to join our Santa Clara, CA team. As an SRE at NVIDIA, you will have a meaningful role in keeping our Digital... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $252k - $308k

     ...Staff Site Reliability Engineer Mountain View, US About EarnIn As one of the first pioneers of earned wage access, our passion at EarnIn...  ...without increasing operational risk. This role exists to lead EarnIn's next stage of reliability maturity: an AI-first operating... 
    Full time
    Work at office
    2 days per week

    Earnin

    Mountain View, CA
    3 days ago
  • $200k - $260k

     ...every company. About the Role: Glean is seeking a Site Reliability Engineering Lead to foster a culture of engineering excellence, drive...  ...like Terraform is essential. ~ Solid understanding of networking, security principles, and best SRE and security practices... 
    Work at office
    Home office
    Flexible hours

    Glean.info

    Mountain View, CA
    1 day ago
  • $272k - $431.25k

    NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning and Process) Cloud Infrastructure...  ...and cost efficiency of AI development and testing systems.Leading software development projects and technically direct a team... 
    Full time
    Work experience placement
    Worldwide

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

    At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability. This demanding position merges software and systems engineering efforts to guarantee flawless service operation... 
    Full time

    Nvidia

    Santa Clara, CA
    4 hours ago
  •  ...beyond. Together, we advance your career. PLATFORM THERMAL LEADER  THE ROLE:  We are looking for a dynamic, energetic platform Thermal Lead to join our growing team. THE PERSON:   As a leader of the Server Platform Thermal Mechanical (SPTM) group, you will represent,... 

    AMD

    Santa Clara, CA
    4 days ago
  • $174k - $253k

     ...changes that improve reliability and velocity.Practice...  ...in Computer Science, Engineering, a related field, or equivalent...  ...2 years of experience leading projects and providing...  ...or Engineering.Site Reliability Engineering...  ...rebuild them. We keep our networks up and running,... 

    Google

    Sunnyvale, CA
    3 days ago
  • $170k - $200k

    We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining...  ....Manage datacenter infrastructure (Linux servers, network devices, databases etc,.) Improve observability with logging... 
    Full time
    Worldwide

    Fortinet

    Sunnyvale, CA
    2 days ago
  • $147k - $211k

     ...Review code developed by other engineers and provide feedback to...  ...and the impact on hardware, network, or service operations and quality. Participate in, or lead design reviews with peers and...  ...large-scale distributed systems. Site Reliability Engineering (SRE) is what you... 

    Google

    Sunnyvale, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineering - Network. Be the first to apply!