Lead Site Reliability Engineering - Network
JP Morgan Chase
Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.As a Lead Site Reliability Engineer at JPMorgan Chase within the Network Product, you hold a leadership role in your team, demonstrate strong knowledge across multiple technical domains, and advise others on the technical and business issues facing them. Take lead and conduct resiliency design reviews, break up complex problems into digestible work for other engineers, act as a technical lead for medium to large-sized products, and provide advice and mentoring to other engineers.Job responsibilitiesApplies network reliability principles (Permit to Operate, FMEA, operational readiness), balancing feature delivery, efficiency, and stability.Partners with network engineering domains (Datacenter, Firewall, Proxies, DMZ, Load Balancing, etc.) and Lines of Business to align goals and outcomes.Drives adoption of reliability best practices and observability, demonstrating impact through stability/reliability metrics.; Bridges Engineering, Operations, DevOps, and customers to build resilient, scalable, and secure network services.Provides Tier-3 network support, leading major incident response, rapid restoration, RCA, and follow-through on corrective actions.Leads reliability and stability initiatives using data-driven analysis to improve service levels and reduce recurring failure modes.Defines SLI/SLOs and error budgets with stakeholders and customers, ensuring measurable performance targets and trade-off clarity. Identifies and removes technical bottlenecks within core domains of expertise, proactively preventing reliability and capacity risks.Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.Runs blameless, data-driven post-mortems and debriefs, converting learnings (successes and failures) into actionable improvements.Fosters continuous improvement and strong knowledge sharing, soliciting real-time feedback, avoiding duplicated work, and promoting innovation via internal communities.Produces and packages thought leadership with specialists/product/engineering teams—documenting best practices and lessons learned for internal assets and industry forums/conferences.Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls.Required qualifications, capabilities, and skillsFormal training or certification in network engineering concepts and 5+ years of applied experience.10+ years of experience leading technologists to manage and solve complex technical items within your domain of expertise.Advanced proficiency in network reliability engineering, including Permit to Operate, FMEA, and operational readiness processes.Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations.Experience leading technologists to manage and solve complex network issues at a firmwide level.Ability to influence team culture by championing innovation and change for success.Proficiency in SD-WAN, cloud platforms (AWS, Azure, etc.), and major network technologies (Palo Alto, Juniper, F5, Broadcom, Arista, Cisco, etc.).Proficiency in observability and monitoring tools such as Grafana, SevOne, Prometheus, Kibana, ThousandEyes, and Splunk.Preferred qualifications, capabilities, and skillsCCIE, Load-balancing, SD-WAN, Observability tools, eBPF, Cloud certsDemonstrated proficiency in troubleshooting and supporting complex networking environments, including Tier-3 operational support for major incidents.Experience with continuous integration and delivery tools (e.g., Jenkins, GitLab, Terraform, etc.).Experience in scalable networking design, including high availability, redundancy, failover, and load balancing.Experience troubleshooting networking protocols such as TCP/IP, and BGP.Experience in customer-facing migration, including service discovery, assessment, planning, execution, and operations.This position is subject to Section 19 of the Federal Deposit Insurance Act. As such, an employment offer for this position is contingent on JPMorganChase’s review of criminal conviction history, including pretrial diversions or program entries. JPMorganChase, one of the oldest financial institutions, offers innovative financial solutions to millions of consumers, small businesses and many of the world’s most prominent corporate, institutional and government clients under the J.P. Morgan and Chase brands. Our history spans over 200 years and today we are a leader in investment banking, consumer and small business banking, commercial banking, financial transaction processing and asset management. We offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may receive commission-based pay and/or discretionary incentive compensation, paid in the form of cash and/or forfeitable equity, awarded in recognition of individual achievements and contributions. We also offer a range of benefits and programs to meet employee needs, based on eligibility. These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare, tuition reimbursement, mental health support, financial coaching and more. Additional details about total compensation and benefits will be provided during the hiring process. We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants’ and employees’ religious practices and beliefs, as well as mental health or physical disability needs. Visit our FAQs for more information about requesting an accommodation.JPMorgan Chase & Co. is an Equal Opportunity Employer, including Disability/VeteransOur professionals in our Corporate Functions cover a diverse range of areas from finance and risk to human resources and marketing. Our corporate teams are an essential part of our company, ensuring that we’re setting our businesses, clients, customers and employees up for success.Full timePosting Date: 2026-06-16
- Lead Cloud ArchitectCooley is seeking a Lead Cloud Architect to join the Innovation team... ...owning the technical architecture and engineering standards for a greenfield SOC2-... ...Extensive hands-on background with AWS (IAM, networking, Organizations) and Terraform at scaleExperience...SuggestedFull timeWork at officeLocal areaImmediate startRemote workWork from homeWorldwideWeekend work
- ...everyone. Role Summary We are seeking an experienced Site Reliability Engineer to help design, build, and operate the infrastructure that... ...at scale. ~ Deep understanding of Kubernetes internals, networking, storage and service mesh architectures. ~ Proven track...SuggestedFull timeContract work
$165k - $280k
...with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building... ...HBase, HDFS, FlinkExperience troubleshooting hardware and network-layer issuesProgramming experience in Python, C#, Java,...SuggestedPermanent employmentTemporary workWorldwideWeekend work$140k - $205k
Senior Technology Site Reliability EngineerCooley is seeking a Senior Site Reliability Engineer to join the Infrastructure & Development Operations team.Position summary: The Senior Technology Site Reliability Engineer (“SRE”) is responsible for ensuring the reliability...SuggestedFull timeTemporary workWork at officeFlexible hoursWeekend work$165k - $190k
Obsidian Security is the leading SaaS security platform, trusted by global enterprises... ...DevOps/SRE team at Obsidian ensures that engineering excellence translates into stable,... ...complex challenges around scalability, reliability, observability, and cost efficiencyCollaborate...SuggestedWork from home- ...role focuses on translating requirements into scalable designs, leading architectural initiatives and guiding delivery teams to high-... ...solutions meet business outcomes and best practices across compute, network, and storage. A relevant VMware certification and extensive VCF...Remote work
$232k - $263k
Obsidian Security is the leading SaaS security platform, trusted by global enterprises... ...growth and IPO readiness.Sr. Staff Site Reliability EngineerAs a Sr. Staff SRE at Obsidian... ...strategic partner to DevOps and Platform Engineering leadership, shaping a unified...Work from home$210.6k - $305.1k
...digital experiences across every network - even the ones they don’t... ...insights within Cisco’s leading Networking, Security, Collaboration... ...led a distributed team of 5+ engineers, can demonstrate strong... ...Please see the Cisco careers site to discover more benefits and...Full timeTemporary workLocal areaFlexible hours- Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally... ...yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at... ...community, and continues to expand network and leads evaluation sessions with...
$186.9k - $267.7k
...approximately 2 days per week on-site at Cisco offices in... ...intended, improving reliability and reducing risks.... ...Site Reliability Engineer (SRE), you will provide... ...reliability strategy, lead major infrastructure initiatives... ..., databases, and networking.Drive automation to...Full timeTemporary workLocal areaFlexible hours2 days per week$100k - $200k
...OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about...Full time$170k - $250k
...Site Reliability Engineer (SRE) Location: San Francisco, CA / Palo Alto, CA Company Stage of Funding: Growth-Stage AI Infrastructure Company... ...distributed systems with a strong understanding of networking fundamentals. ~ Proficiency with Python and/or Go for building...Work at officeVisa sponsorshipFlexible hours- ...Site Reliability Engineer, Data Platform - USDS Responsibilities Engage in and improve the whole lifecycle... ...cloud native technologies and networking. Experience developing tools and APIs... ...related benefits. About USDS TikTok is the leading destination for short-form mobile...
- ...Senior Site Reliability Engineer Latitude AI is building the future of Ford's autonomy roadmap to... ...Latitude team, you'll work alongside leading experts across machine learning and robotics... ...operating system internals, TCP/IP networking, and storage subsystems Hands on...Work at officeImmediate start
$150k - $175k
...Site Reliability Engineer At ASAPP, our mission is simple: deliver the best AI-powered customer experience—faster than anyone else. To achieve... ...our product engineers Design secure and performant networking solutions in our production systems What You'll Need...Remote work$137.77k - $194.59k
...distributed team of roughly 80 scientists and engineers building and operating Rubin's petascale... .... Your role: You will own the reliability and robustness of Rubin Observatory's... ...nature of this position, SLAC is open to on-site, hybrid, and remote work options....Remote workFlexible hoursNight shift- ...role. JOB DESCRIPTION Elevate your engineering prowess to unprecedented levels by... ...yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at JPMorgan... ...community, and continues to expand network and leads evaluation sessions with...
- ...together safely, transparently, and at scale. Join Wand in leading the Agentic Shift Wand is building a high-performing... ...a hands‑on Head of SRE to establish, lead, and scale our Site Reliability Engineering function. This role combines strategic ownership with deep...Shift work
$170k - $230k
...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About Mithril Mithril is... ...accessible and affordable for the world's leading enterprises, AI startups, and the AI... ...practical understanding of cloud networking fundamentals (VPC, DNS, load balancing...Work at officeLocal area1 day per week- ...Acryl Data seeks a Site Reliability Engineering (SRE) Tech Lead to enhance the reliability and scalability of its DataHub platform. The role involves leading infrastructure design, optimizing system performance, and driving continuous improvement across cloud deployments...
- ...ActiveHours is looking for an experienced DevOps Engineer to enhance platform automation and site reliability in a collaborative environment. The role involves automating key systems, improving visibility through metrics, and troubleshooting critical problems. You'll...
$168k - $264.5k
...impact on the world.Are you ready to be part of something outstanding? NVIDIA's Digital Marketing Organization seeks a senior Site Reliability Engineer (SRE) to join our Santa Clara, CA team. As an SRE at NVIDIA, you will have a meaningful role in keeping our Digital...Full time$252k - $308k
...Staff Site Reliability Engineer Mountain View, US About EarnIn As one of the first pioneers of earned wage access, our passion at EarnIn... ...without increasing operational risk. This role exists to lead EarnIn's next stage of reliability maturity: an AI-first operating...Full timeWork at office2 days per week$200k - $260k
...every company. About the Role: Glean is seeking a Site Reliability Engineering Lead to foster a culture of engineering excellence, drive... ...like Terraform is essential. ~ Solid understanding of networking, security principles, and best SRE and security practices...Work at officeHome officeFlexible hours$272k - $431.25k
NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning and Process) Cloud Infrastructure... ...and cost efficiency of AI development and testing systems.Leading software development projects and technically direct a team...Full timeWork experience placementWorldwide$184k - $287.5k
At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability. This demanding position merges software and systems engineering efforts to guarantee flawless service operation...Full time- ...beyond. Together, we advance your career. PLATFORM THERMAL LEADER THE ROLE: We are looking for a dynamic, energetic platform Thermal Lead to join our growing team. THE PERSON: As a leader of the Server Platform Thermal Mechanical (SPTM) group, you will represent,...
$174k - $253k
...changes that improve reliability and velocity.Practice... ...in Computer Science, Engineering, a related field, or equivalent... ...2 years of experience leading projects and providing... ...or Engineering.Site Reliability Engineering... ...rebuild them. We keep our networks up and running,...$170k - $200k
We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining... ....Manage datacenter infrastructure (Linux servers, network devices, databases etc,.) Improve observability with logging...Full timeWorldwide$147k - $211k
...Review code developed by other engineers and provide feedback to... ...and the impact on hardware, network, or service operations and quality. Participate in, or lead design reviews with peers and... ...large-scale distributed systems. Site Reliability Engineering (SRE) is what you...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineering - Network. Be the first to apply!
- lead operating engineer Palo Alto, CA
- lead engineer Palo Alto, CA
- site reliability engineer Palo Alto, CA
- site reliability engineer sre Palo Alto, CA
- network engineer level Palo Alto, CA
- network infrastructure engineer Palo Alto, CA
- network engineer - transport Palo Alto, CA
- data center network engineer Palo Alto, CA
- network implementation engineer Palo Alto, CA
- ip network engineer Palo Alto, CA

