Lead Site Reliability Engineer, Platforms
$124k - $271.2kZoom Video Communications, Inc.
What You Can ExpectAs a Lead Staff Site Reliability Engineer, you will be one of the technical leads for our DevOps Platforms organization. This group is responsible for DevOps Platforms including cloud infrastructure, physical data center orchestration, critical security services, and our Zoom for Government (ZfG) environment. You will be an uber tech lead working across a broad area, defining projects and guiding work across various teams. Your scope of work is wide and you will have the opportunity to improve our datacenter kubernetes infrastructure, our cloud infrastructure, our security posture, and our operation of ZfG environments. Broadly speaking, you are an exemplary SRE and you will guide all of our teams toward SRE best practices (automation, monitoring, infrastructure as code, etc).About the TeamThe DevOps Platforms organization owns the full infrastructure stack: cloud infrastructure on AWS and OCI, physical data center orchestration, critical security services (identity, authentication, authorization), and Zoom's FedRAMP-rated federal environment, Zoom for Government (ZfG).The team is currently working on one of the most technically interesting work in the org, hardening our security posture, and building the automation and reliability systems that underpin Zoom's global services. If you want broad visibility, real cross-team influence, and the chance to define how infrastructure gets built, this is the seat.ResponsibilitiesDesign and scale DevOps platform services including Kubernetes infrastructure, cloud systems, and compliance-ready environmentsDefine technical roadmaps and architectural direction for infrastructure automation and securityPartner with service teams to understand platform needs and deliver solutions that improve reliability and efficiencyEstablish and advocate for SRE best practices including infrastructure as code, monitoring, and incident managementMentor team members through design, implementation, and production deployment of complex systemsWhat We're Looking ForBring 8+ years of SRE or DevOps experience building and operating production infrastructure at scaleCode proficiently in at least one programming language beyond scripting (e.g., Python, Go, Java)Deploy and manage CI/CD pipelines using tools like Git, Jenkins, Argo CD, or JFrogOperate cloud infrastructure on AWS, OCI, or similar platforms using Terraform and KubernetesImplement observability solutions with logging and monitoring tools such as ELK, Prometheus, or GrafanaCommunicate complex technical concepts clearly to diverse audiences including security teams, senior leadership, and external auditorsParticipate in on-call rotations and lead incident response to maintain system reliabilityHold a degree in Computer Science or related field, or equivalent practical experienceHold US citizenship, or Greencard statusPreferredHave experience with security from an SRE perspectiveHave experience with Identity security (e.g. IAM, workload identity, zero trust) and tools (e.g. Teleport, Okta)Have experience operating Government environments and understanding their compliance requirementsHave experience with system design and distributed computing at scaleAbility to speak Chinese/Mandarin is a plus, but not requiredSalary Range or On Target Earnings:Minimum:$124,000.00Maximum:$271,200.00In addition to the base salary and/or OTE listed Zoom has a Total Direct Compensation philosophy that takes into consideration; base salary, bonus and equity value.Note: Starting pay will be based on a number of factors and commensurate with qualifications & experience.We also have a location based compensation structure; there may be a different range for candidates in this and other locationsAt Zoom, we offer a window of at least 5 days for you to apply because we believe in giving you every opportunity. Below is the potential closing date, just in case you want to mark it on your calendar. We look forward to receiving your application!Anticipated Position Close Date:09/17/26Ways of WorkingOur structured hybrid approach is centered around our offices and remote work environments. The work style of each role, Hybrid, Remote, or In-Person is indicated in the job description/posting.BenefitsAs part of our award-winning workplace culture and commitment to delivering happiness, our benefits program offers a variety of perks, benefits, and options to help employees maintain their physical, mental, emotional, and financial health; support work-life balance; and contribute to their community in meaningful ways. Click Learnfor more information.About UsZoomies help people stay connected so they can get more done together. We set out to build the best collaboration platform for the enterprise, and today help people communicate better with products like Zoom Contact Center, Zoom Phone, Zoom Events, Zoom Apps, Zoom Rooms, and Zoom Webinars.We’re problem-solvers, working at a fast pace to design solutions with our customers and users in mind. Find room to grow with opportunities to stretch your skills and advance your career in a collaborative, growth-focused environment.Our CommitmentAt Zoom, we believe great work happens when people feel supported and empowered. We’re committed to fair hiring practices that ensure every candidate is evaluated based on skills, experience, and potential. If you require an accommodation during the hiring process, let us know—we’re here to support you at every step.If you need assistance navigating the interview process due to a medical disability, please submit an Accommodations Request Form and someone from our team will reach out soon. This form is solely for applicants who require an accommodation due to a qualifying medical disability. Non-accommodation-related requests, such as application follow-ups or technical issues, will not be addressed.Our interviews are supported by BrightHire, a tool that helps us create a consistent and thoughtful interview experience and may include recordings. Please refer to our candidate privacy statement for more information of how we use your data.SummaryLocation: San Jose (CA)Type: Full time
- ...Tuesday.About the RoleLambda’s Core Cloud Platform powers compute provisioning and... ...centers. We are looking for a Senior Site Reliability Engineer to improve the reliability, scalability... ...radius and prevent cascading failures.Lead production incident response, postmortems...SuggestedWork at officeLocal areaWork from homeFlexible hours
$298k - $368k
...applied to a range of vehicle platforms and product use cases. The Waymo... .... states. As the Pipeline SRE Lead, you will play a key role in driving the reliability of our most critical release pipelines... ...as proactively partnering with engineering to evolve our software system...SuggestedFull timeRemote work$101k - $161k
...prestigious awards, such as Best Engineering Team, Best Company for... ...Work WithWe’re looking for Site Reliability Engineers to join our growing... ...Familiarity with GCP (Google Cloud Platform) and GKE (Google Kubernetes... ...to be drive, develop, and lead projects in any of the...Suggested$230k - $250k
...foundation for autonomous networking, giving engineers and AI agents the ability to know the... ...comfortable, building a groundbreaking platform that transforms how teams run and... ...always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "...SuggestedNight shift- ...simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure.... ...have deep experience configuring New Relic (or similar platforms) to create meaningful dashboards, SLIs, and SLOs.Automation...SuggestedFull timeWork at office2 days per week
$148k - $235.75k
...the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for... ...checks, post-deploy validation), and lead rollbacks/remediations when needed....Full time- Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by... ...security and operational data through its Intelligent Operations Platform. Built to address the increasing complexity of modern...Flexible hours
$168k - $270.25k
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization... ...intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing...Full time$267k - $356k
...currently Tuesday.Lambda's Storage Engineering team is the backbone behind... ...spectrum of Lambda's data platform services—from low-level... ...in the industry, which means reliability and performance aren't just... ...storage across new and existing sites using tools such as Ansible,...Work experience placementWork at officeLocal areaWork from homeFlexible hours$152k - $241.5k
...and amazing people. NVIDIA is leading the way in groundbreaking... ...generation of our global services platform. At NVIDIA, you’ll keep... ...lifecycle management, fleet reliability/auto-healing, E2E observability... ..., or Ruby.Mentored other engineers and influenced technical direction...Full time- ...home day is currently Tuesday.Engineering at Lambda is responsible for... ...multi-tenant cloud networking platform and SDN infrastructureOperate... ...teams to improve service reliability and deployment workflowsDeploy... ...rotationYouHave 5+ years of experience in Site Reliability Engineering,...Work at officeLocal areaWork from homeFlexible hours
- ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building... ...tooling and automate the validation of platform quality.Design, build, and maintain scalable... ...services, workloads, and platform reliability.You6+ years of experience in a SRE, operations...Work at officeLocal areaWork from homeFlexible hours
$207k - $300k
...consulting, developing software platforms and frameworks, capacity... ...for changes that improve reliability and velocity.Practice sustainable... ....3 years of experience leading projects.3 years of... ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE...$146.7k - $339.3k
...available for this positionWhat you can expect As a Senior Lead Site Reliability Engineer, you can anticipate opportunities to work on our hybrid... ...design patterns. Partner with Security, Networking, and Platform teams on architecture roadmaps. Influence vendor and hardware...Full timeWork at officeRemote workWorldwideShift workWeekend work$192.4k - $275.8k
...demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service... ...automation and frameworks that make the whole platform more resilient. If you are the kind... ...audiences 4+ years experience leading post-mortems and root cause analysis...Full timeTemporary workLocal areaFlexible hours$207k - $300k
...team members to enhance system reliability and efficiency.Initiate, own, and lead large-scale, complex projects... ...techniques.3 years of experience as a Site Reliability Engineer.3 years of experience leading... ...grow.Semantic Understanding Platform (SUP) is the standard platform...$207.4k - $259.2k
...are seeking a highly experienced and passionate Sr. Staff Site Reliability Engineer (SRE) to join our growing team. In this critical role, you... ...internal LLM-powered chat service, potentially leveraging platforms like OpenRouter or similar alternatives.implement and maintain...Permanent employmentLocal area$122.5k - $175k
...security data lake to power our cloud-native Zero Trust Exchange platform. This innovation protects our customers from cyberattacks... ...future of cybersecurity.RoleWe are looking for a Staff Site Reliability Engineer to join our team. This is a hybrid role going into the San...Full timeWork at officeLocal area3 days per week- ServiceNow in Santa Clara, CA is seeking a Senior Software Engineer - SRE & AIOps to advance infrastructure automation, resilience, and toil reduction across hybrid cloud operations. You will deploy and manage production Kubernetes clusters, build auto-remediation, and...
$200k - $322k
...endeavor!What you will be doing:Lead initiatives to transform IT... ...building for performance and reliability at global scale, covering... ...with NVIDIA leadership, senior engineers, program managers, and product... ...proven experience in compute platform engineering with a focus on automation...Full timeRemote work$272k - $431.25k
NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure,... ...Linux, and Android. It supports hardware platforms including NVIDIA GPUs and Tegra... ...of AI development and testing systems.Leading software development projects and technically...Full timeWork experience placementWorldwide$187.04k - $359.72k
...changes that improve reliability and velocity. Qualifications... ...Science, Electrical Engineering, Computer Engineering... ...USDS TikTok is the leading destination for short-... ...protection of the TikTok platform and U.S. user data, so... ...and more. On-site presence across teams...Temporary workLocal areaOverseasShift work- ...responsibilities of a Technical Support Engineer within a SaaS (Software as a Service) environment with a growing focus on Site Reliability Engineering (SRE). The ideal candidate has... ..., and scale an AI Security Public SaaS platform, operating AI inference workloads at...Work at officeLocal areaRemote workWork from home
- ...A leading technology firm is in search of a Senior Wireless Network Site Reliability Engineer to manage and enhance their wireless network infrastructure. The ideal candidate has over 8 years of experience in wireless network operations and a strong background in wireless...
- ...We are seeking a Senior Database Reliability Engineer (DBRE) to design, operate, and improve reliable, scalable, secure, and highly available database and data platforms. The role combines database engineering, site reliability engineering, Linux systems administration...
$100k - $170k
...We are seeking a DevOps Engineer who is eager to have an immediate... ..., testing, and release platforms, taking us from square one to... ...on Amazon EKS with focus on reliability and performance Build and... ...Participate in on-call rotation and lead incident response for production...Full timeWork at officeImmediate startVisa sponsorshipNight shift$207k - $300k
...the next generation of Indexing Engine platform, Web Data Service (for rapid prototyping... ..., YouTube, Shopping and Lens.Lead key projects to ensure and improve reliability of the pipeline we supported... ...programming in Golang or C++.Site Reliability Engineering (SRE) combines...$214.1k - $309.8k
...global team of software engineers and SREs responsible... ...operations partners to ensure reliability and performance.Webex... ...distributed data platforms across AWS and private... ...OpenSearch. You’ll lead cloud engineering initiatives... ...see the Cisco careers site to discover more...Permanent employmentFull timeTemporary workLocal areaWorldwideFlexible hoursShift work- ...Job Title : Senior Site Reliability Engineer Location : Santa Clara, CA Contract ENGAGEMENT... ...will provide SRE services for AI platforms and supporting infrastructure with emphasis... ...infrastructure components. Lead or support incident triage for service...Contract work
- Lendistry, LLC. is seeking a Senior AI Engineer to lead the delivery of AI solutions, including document intelligence and risk assessment tools. In this role, you will be responsible for mentoring junior engineers and shaping AI-driven workflows, improving the borrower...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer, Platforms. Be the first to apply!
- lead operating engineer San Jose, CA
- lead engineer San Jose, CA
- site reliability engineer San Jose, CA
- site reliability engineer sre San Jose, CA
- platform engineer San Jose, CA
- platform developer San Jose, CA
- senior platform engineer San Jose, CA
- platform engineering manager San Jose, CA
- website coordinator San Jose, CA
- on-site clinical research associate (traveling/remote) San Jose, CA


