Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineer, Platforms

$124k - $271.2k

Zoom Video Communications, Inc.

What You Can ExpectAs a Lead Staff Site Reliability Engineer, you will be one of the technical leads for our DevOps Platforms organization. This group is responsible for DevOps Platforms including cloud infrastructure, physical data center orchestration, critical security services, and our Zoom for Government (ZfG) environment. You will be an uber tech lead working across a broad area, defining projects and guiding work across various teams. Your scope of work is wide and you will have the opportunity to improve our datacenter kubernetes infrastructure, our cloud infrastructure, our security posture, and our operation of ZfG environments. Broadly speaking, you are an exemplary SRE and you will guide all of our teams toward SRE best practices (automation, monitoring, infrastructure as code, etc).About the TeamThe DevOps Platforms organization owns the full infrastructure stack: cloud infrastructure on AWS and OCI, physical data center orchestration, critical security services (identity, authentication, authorization), and Zoom's FedRAMP-rated federal environment, Zoom for Government (ZfG).The team is currently working on one of the most technically interesting work in the org, hardening our security posture, and building the automation and reliability systems that underpin Zoom's global services. If you want broad visibility, real cross-team influence, and the chance to define how infrastructure gets built, this is the seat.ResponsibilitiesDesign and scale DevOps platform services including Kubernetes infrastructure, cloud systems, and compliance-ready environmentsDefine technical roadmaps and architectural direction for infrastructure automation and securityPartner with service teams to understand platform needs and deliver solutions that improve reliability and efficiencyEstablish and advocate for SRE best practices including infrastructure as code, monitoring, and incident managementMentor team members through design, implementation, and production deployment of complex systemsWhat We're Looking ForBring 8+ years of SRE or DevOps experience building and operating production infrastructure at scaleCode proficiently in at least one programming language beyond scripting (e.g., Python, Go, Java)Deploy and manage CI/CD pipelines using tools like Git, Jenkins, Argo CD, or JFrogOperate cloud infrastructure on AWS, OCI, or similar platforms using Terraform and KubernetesImplement observability solutions with logging and monitoring tools such as ELK, Prometheus, or GrafanaCommunicate complex technical concepts clearly to diverse audiences including security teams, senior leadership, and external auditorsParticipate in on-call rotations and lead incident response to maintain system reliabilityHold a degree in Computer Science or related field, or equivalent practical experienceHold US citizenship, or Greencard statusPreferredHave experience with security from an SRE perspectiveHave experience with Identity security (e.g. IAM, workload identity, zero trust) and tools (e.g. Teleport, Okta)Have experience operating Government environments and understanding their compliance requirementsHave experience with system design and distributed computing at scaleAbility to speak Chinese/Mandarin is a plus, but not requiredSalary Range or On Target Earnings:Minimum:$124,000.00Maximum:$271,200.00In addition to the base salary and/or OTE listed Zoom has a Total Direct Compensation philosophy that takes into consideration; base salary, bonus and equity value.Note: Starting pay will be based on a number of factors and commensurate with qualifications & experience.We also have a location based compensation structure; there may be a different range for candidates in this and other locationsAt Zoom, we offer a window of at least 5 days for you to apply because we believe in giving you every opportunity. Below is the potential closing date, just in case you want to mark it on your calendar. We look forward to receiving your application!Anticipated Position Close Date:09/17/26Ways of WorkingOur structured hybrid approach is centered around our offices and remote work environments. The work style of each role, Hybrid, Remote, or In-Person is indicated in the job description/posting.BenefitsAs part of our award-winning workplace culture and commitment to delivering happiness, our benefits program offers a variety of perks, benefits, and options to help employees maintain their physical, mental, emotional, and financial health; support work-life balance; and contribute to their community in meaningful ways. Click Learnfor more information.About UsZoomies help people stay connected so they can get more done together. We set out to build the best collaboration platform for the enterprise, and today help people communicate better with products like Zoom Contact Center, Zoom Phone, Zoom Events, Zoom Apps, Zoom Rooms, and Zoom Webinars.We’re problem-solvers, working at a fast pace to design solutions with our customers and users in mind. Find room to grow with opportunities to stretch your skills and advance your career in a collaborative, growth-focused environment.Our CommitmentAt Zoom, we believe great work happens when people feel supported and empowered. We’re committed to fair hiring practices that ensure every candidate is evaluated based on skills, experience, and potential. If you require an accommodation during the hiring process, let us know—we’re here to support you at every step.If you need assistance navigating the interview process due to a medical disability, please submit an Accommodations Request Form and someone from our team will reach out soon. This form is solely for applicants who require an accommodation due to a qualifying medical disability. Non-accommodation-related requests, such as application follow-ups or technical issues, will not be addressed.Our interviews are supported by BrightHire, a tool that helps us create a consistent and thoughtful interview experience and may include recordings. Please refer to our candidate privacy statement for more information of how we use your data.SummaryLocation: San Jose (CA)Type: Full time

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineer, Platforms in San Jose, CA vacancy
  •  ...Tuesday.About the RoleLambda’s Core Cloud Platform powers compute provisioning and...  ...centers. We are looking for a Senior Site Reliability Engineer to improve the reliability, scalability...  ...radius and prevent cascading failures.Lead production incident response, postmortems... 
    Suggested
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  • $298k - $368k

     ...applied to a range of vehicle platforms and product use cases. The Waymo...  .... states. As the Pipeline SRE Lead, you will play a key role in driving the reliability of our most critical release pipelines...  ...as proactively partnering with engineering to evolve our software system... 
    Suggested
    Full time
    Remote work

    Waymo

    Mountain View, CA
    3 days ago
  • $101k - $161k

     ...prestigious awards, such as Best Engineering Team, Best Company for...  ...Work WithWe’re looking for Site Reliability Engineers to join our growing...  ...Familiarity with GCP (Google Cloud Platform) and GKE (Google Kubernetes...  ...to be drive, develop, and lead projects in any of the... 
    Suggested

    Arista Networks

    Santa Clara, CA
    2 days ago
  • $230k - $250k

     ...foundation for autonomous networking, giving engineers and AI agents the ability to know the...  ...comfortable, building a groundbreaking platform that transforms how teams run and...  ...always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "... 
    Suggested
    Night shift

    Forward Networks

    Santa Clara, CA
    4 days ago
  •  ...simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure....  ...have deep experience configuring New Relic (or similar platforms) to create meaningful dashboards, SLIs, and SLOs.Automation... 
    Suggested
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    4 days ago
  • $148k - $235.75k

     ...the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for...  ...checks, post-deploy validation), and lead rollbacks/remediations when needed.... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by...  ...security and operational data through its Intelligent Operations Platform. Built to address the increasing complexity of modern... 
    Flexible hours

    Sumo Logic

    San Jose, CA
    2 days ago
  • $168k - $270.25k

    NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization...  ...intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $267k - $356k

     ...currently Tuesday.Lambda's Storage Engineering team is the backbone behind...  ...spectrum of Lambda's data platform services—from low-level...  ...in the industry, which means reliability and performance aren't just...  ...storage across new and existing sites using tools such as Ansible,... 
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  • $152k - $241.5k

     ...and amazing people. NVIDIA is leading the way in groundbreaking...  ...generation of our global services platform. At NVIDIA, you’ll keep...  ...lifecycle management, fleet reliability/auto-healing, E2E observability...  ..., or Ruby.Mentored other engineers and influenced technical direction... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...home day is currently Tuesday.Engineering at Lambda is responsible for...  ...multi-tenant cloud networking platform and SDN infrastructureOperate...  ...teams to improve service reliability and deployment workflowsDeploy...  ...rotationYouHave 5+ years of experience in Site Reliability Engineering,... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    1 day ago
  •  ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building...  ...tooling and automate the validation of platform quality.Design, build, and maintain scalable...  ...services, workloads, and platform reliability.You6+ years of experience in a SRE, operations... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    3 days ago
  • $207k - $300k

     ...consulting, developing software platforms and frameworks, capacity...  ...for changes that improve reliability and velocity.Practice sustainable...  ....3 years of experience leading projects.3 years of...  ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE... 

    Google

    San Jose, CA
    4 days ago
  • $146.7k - $339.3k

     ...available for this positionWhat you can expect As a Senior Lead Site Reliability Engineer, you can anticipate opportunities to work on our hybrid...  ...design patterns. Partner with Security, Networking, and Platform teams on architecture roadmaps. Influence vendor and hardware... 
    Full time
    Work at office
    Remote work
    Worldwide
    Shift work
    Weekend work

    Zoom

    San Jose, CA
    4 days ago
  • $192.4k - $275.8k

     ...demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service...  ...automation and frameworks that make the whole platform more resilient. If you are the kind...  ...audiences 4+ years experience leading post-mortems and root cause analysis... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    4 days ago
  • $207k - $300k

     ...team members to enhance system reliability and efficiency.Initiate, own, and lead large-scale, complex projects...  ...techniques.3 years of experience as a Site Reliability Engineer.3 years of experience leading...  ...grow.Semantic Understanding Platform (SUP) is the standard platform... 

    Google

    San Jose, CA
    4 days ago
  • $207.4k - $259.2k

     ...are seeking a highly experienced and passionate Sr. Staff Site Reliability Engineer (SRE) to join our growing team. In this critical role, you...  ...internal LLM-powered chat service, potentially leveraging platforms like OpenRouter or similar alternatives.implement and maintain... 
    Permanent employment
    Local area

    Archer Aviation

    San Jose, CA
    2 days ago
  • $122.5k - $175k

     ...security data lake to power our cloud-native Zero Trust Exchange platform. This innovation protects our customers from cyberattacks...  ...future of cybersecurity.RoleWe are looking for a Staff Site Reliability Engineer to join our team. This is a hybrid role going into the San... 
    Full time
    Work at office
    Local area
    3 days per week

    Zscaler

    San Jose, CA
    17 hours ago
  • ServiceNow in Santa Clara, CA is seeking a Senior Software Engineer - SRE & AIOps to advance infrastructure automation, resilience, and toil reduction across hybrid cloud operations. You will deploy and manage production Kubernetes clusters, build auto-remediation, and... 

    Servicenow

    Santa Clara, CA
    1 day ago
  • $200k - $322k

     ...endeavor!What you will be doing:Lead initiatives to transform IT...  ...building for performance and reliability at global scale, covering...  ...with NVIDIA leadership, senior engineers, program managers, and product...  ...proven experience in compute platform engineering with a focus on automation... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  • $272k - $431.25k

    NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure,...  ...Linux, and Android. It supports hardware platforms including NVIDIA GPUs and Tegra...  ...of AI development and testing systems.Leading software development projects and technically... 
    Full time
    Work experience placement
    Worldwide

    Nvidia

    Santa Clara, CA
    2 days ago
  • $187.04k - $359.72k

     ...changes that improve reliability and velocity. Qualifications...  ...Science, Electrical Engineering, Computer Engineering...  ...USDS TikTok is the leading destination for short-...  ...protection of the TikTok platform and U.S. user data, so...  ...and more. On-site presence across teams... 
    Temporary work
    Local area
    Overseas
    Shift work

    Tik Tok

    San Jose, CA
    4 days ago
  •  ...responsibilities of a Technical Support Engineer within a SaaS (Software as a Service) environment with a growing focus on Site Reliability Engineering (SRE). The ideal candidate has...  ..., and scale an AI Security Public SaaS platform, operating AI inference workloads at... 
    Work at office
    Local area
    Remote work
    Work from home

    F5 Networks

    San Jose, CA
    5 days ago
  •  ...A leading technology firm is in search of a Senior Wireless Network Site Reliability Engineer to manage and enhance their wireless network infrastructure. The ideal candidate has over 8 years of experience in wireless network operations and a strong background in wireless... 

    TechDigital Group

    Santa Clara, CA
    4 days ago
  •  ...We are seeking a Senior Database Reliability Engineer (DBRE) to design, operate, and improve reliable, scalable, secure, and highly available database and data platforms. The role combines database engineering, site reliability engineering, Linux systems administration... 

    Neshent Technologies

    Los Gatos, CA
    3 days ago
  • $100k - $170k

     ...We are seeking a  DevOps Engineer who is eager to have an immediate...  ..., testing, and release platforms, taking us from square one to...  ...on Amazon EKS with focus on reliability and performance Build and...  ...Participate in on-call rotation and lead incident response for production... 
    Full time
    Work at office
    Immediate start
    Visa sponsorship
    Night shift

    E-Space

    Saratoga, CA
    10 days ago
  • $207k - $300k

     ...the next generation of Indexing Engine platform, Web Data Service (for rapid prototyping...  ..., YouTube, Shopping and Lens.Lead key projects to ensure and improve reliability of the pipeline we supported...  ...programming in Golang or C++.Site Reliability Engineering (SRE) combines... 

    Google

    San Jose, CA
    4 days ago
  • $214.1k - $309.8k

     ...global team of software engineers and SREs responsible...  ...operations partners to ensure reliability and performance.Webex...  ...distributed data platforms across AWS and private...  ...OpenSearch. You’ll lead cloud engineering initiatives...  ...see the Cisco careers site to discover more... 
    Permanent employment
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours
    Shift work

    CISCO Systems

    Milpitas, CA
    1 day ago
  •  ...Job Title : Senior Site Reliability Engineer Location : Santa Clara, CA Contract ENGAGEMENT...  ...will provide SRE services for AI platforms and supporting infrastructure with emphasis...  ...infrastructure components. Lead or support incident triage for service... 
    Contract work

    VDart Inc

    Santa Clara, CA
    2 days ago
  • Lendistry, LLC. is seeking a Senior AI Engineer to lead the delivery of AI solutions, including document intelligence and risk assessment tools. In this role, you will be responsible for mentoring junior engineers and shaping AI-driven workflows, improving the borrower... 

    Lendistry, LLC.

    Santa Clara, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineer, Platforms. Be the first to apply!