Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineer, Platforms

$124k - $271.2k

Zoom Video Communications, Inc.

What You Can ExpectAs a Lead Staff Site Reliability Engineer, you will be one of the technical leads for our DevOps Platforms organization. This group is responsible for DevOps Platforms including cloud infrastructure, physical data center orchestration, critical security services, and our Zoom for Government (ZfG) environment. You will be an uber tech lead working across a broad area, defining projects and guiding work across various teams. Your scope of work is wide and you will have the opportunity to improve our datacenter kubernetes infrastructure, our cloud infrastructure, our security posture, and our operation of ZfG environments. Broadly speaking, you are an exemplary SRE and you will guide all of our teams toward SRE best practices (automation, monitoring, infrastructure as code, etc).About the TeamThe DevOps Platforms organization owns the full infrastructure stack: cloud infrastructure on AWS and OCI, physical data center orchestration, critical security services (identity, authentication, authorization), and Zoom's FedRAMP-rated federal environment, Zoom for Government (ZfG).The team is currently working on one of the most technically interesting work in the org, hardening our security posture, and building the automation and reliability systems that underpin Zoom's global services. If you want broad visibility, real cross-team influence, and the chance to define how infrastructure gets built, this is the seat.ResponsibilitiesDesign and scale DevOps platform services including Kubernetes infrastructure, cloud systems, and compliance-ready environmentsDefine technical roadmaps and architectural direction for infrastructure automation and securityPartner with service teams to understand platform needs and deliver solutions that improve reliability and efficiencyEstablish and advocate for SRE best practices including infrastructure as code, monitoring, and incident managementMentor team members through design, implementation, and production deployment of complex systemsWhat We're Looking ForBring 8+ years of SRE or DevOps experience building and operating production infrastructure at scaleCode proficiently in at least one programming language beyond scripting (e.g., Python, Go, Java)Deploy and manage CI/CD pipelines using tools like Git, Jenkins, Argo CD, or JFrogOperate cloud infrastructure on AWS, OCI, or similar platforms using Terraform and KubernetesImplement observability solutions with logging and monitoring tools such as ELK, Prometheus, or GrafanaCommunicate complex technical concepts clearly to diverse audiences including security teams, senior leadership, and external auditorsParticipate in on-call rotations and lead incident response to maintain system reliabilityHold a degree in Computer Science or related field, or equivalent practical experienceHold US citizenship, or Greencard statusPreferredHave experience with security from an SRE perspectiveHave experience with Identity security (e.g. IAM, workload identity, zero trust) and tools (e.g. Teleport, Okta)Have experience operating Government environments and understanding their compliance requirementsHave experience with system design and distributed computing at scaleAbility to speak Chinese/Mandarin is a plus, but not requiredSalary Range or On Target Earnings:Minimum:$124,000.00Maximum:$271,200.00 In addition to the base salary and/or OTE listed Zoom has a Total Direct Compensation philosophy that takes into consideration; base salary, bonus and equity value.Note: Starting pay will be based on a number of factors and commensurate with qualifications & experience.We also have a location based compensation structure; there may be a different range for candidates in this and other locationsAt Zoom, we offer a window of at least 5 days for you to apply because we believe in giving you every opportunity. Below is the potential closing date, just in case you want to mark it on your calendar. We look forward to receiving your application!Anticipated Position Close Date:10/02/26Ways of WorkingOur structured hybrid approach is centered around our offices and remote work environments. The work style of each role, Hybrid, Remote, or In-Person is indicated in the job description/posting.BenefitsAs part of our award-winning workplace culture and commitment to delivering happiness, our benefits program offers a variety of perks, benefits, and options to help employees maintain their physical, mental, emotional, and financial health; support work-life balance; and contribute to their community in meaningful ways. Click Learnfor more information.About UsZoomies help people stay connected so they can get more done together. We set out to build the best collaboration platform for the enterprise, and today help people communicate better with products like Zoom Contact Center, Zoom Phone, Zoom Events, Zoom Apps, Zoom Rooms, and Zoom Webinars.We’re problem-solvers, working at a fast pace to design solutions with our customers and users in mind. Find room to grow with opportunities to stretch your skills and advance your career in a collaborative, growth-focused environment.Our CommitmentAt Zoom, we believe great work happens when people feel supported and empowered. We’re committed to fair hiring practices that ensure every candidate is evaluated based on skills, experience, and potential. If you require an accommodation during the hiring process, let us know—we’re here to support you at every step.If you need assistance navigating the interview process due to a medical disability, please submit an Accommodations Request Form and someone from our team will reach out soon. This form is solely for applicants who require an accommodation due to a qualifying medical disability. Non-accommodation-related requests, such as application follow-ups or technical issues, will not be addressed.Our interviews are supported by BrightHire, a tool that helps us create a consistent and thoughtful interview experience and may include recordings. Please refer to our candidate privacy statement for more information of how we use your data.SummaryLocation: San Jose (CA)Type: Full time

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineer, Platforms in San Jose, CA vacancy
  • $168k - $270.25k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline to design, build and maintain...  ...learn and grow.What you’ll be doing:Lead the technical strategy and roadmap for...  ...develop AI Agents, AI Skills to accelerate platform operationsDrive automation and... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $104.9k - $174.7k

     ...responsible for improving the reliability, availability,...  ...completion.Follow up with engineering, development, security...  ...and operational tasks.Lead or contribute to...  ...years of experience in Site Reliability Engineering...  ...PlanWellbeing: Wellness platform with incentives,... 
    Suggested
    Full time
    Local area

    LexisNexis Risk Solutions Group

    San Jose, CA
    2 days ago
  • $168k - $270.25k

    NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization...  ...intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building...  ...tooling and automate the validation of platform quality.Design, build, and maintain scalable...  ...services, workloads, and platform reliability.You6+ years of experience in a SRE, operations... 
    Suggested
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    3 days ago
  •  ...home day is currently Tuesday.Engineering at Lambda is responsible for...  ...multi-tenant cloud networking platform and SDN infrastructureOperate...  ...teams to improve service reliability and deployment workflowsDeploy...  ...rotationYouHave 5+ years of experience in Site Reliability Engineering,... 
    Suggested
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    1 day ago
  • $148k - $235.75k

     ...the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for...  ...checks, post-deploy validation), and lead rollbacks/remediations when needed.... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $230k - $250k

     ...foundation for autonomous networking, giving engineers and AI agents the ability to know the...  ...comfortable, building a groundbreaking platform that transforms how teams run and...  ...always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "... 
    Night shift

    Forward Networks

    Santa Clara, CA
    4 days ago
  • $101k - $161k

     ...prestigious awards, such as Best Engineering Team, Best Company for...  ...Work WithWe’re looking for Site Reliability Engineers to join our growing...  ...Familiarity with GCP (Google Cloud Platform) and GKE (Google Kubernetes...  ...to be drive, develop, and lead projects in any of the... 

    Arista Networks

    Santa Clara, CA
    2 days ago
  • $267k - $356k

     ...currently Tuesday.Lambda's Storage Engineering team is the backbone behind...  ...spectrum of Lambda's data platform services—from low-level...  ...in the industry, which means reliability and performance aren't just...  ...storage across new and existing sites using tools such as Ansible,... 
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  • $152k - $241.5k

     ...and amazing people. NVIDIA is leading the way in groundbreaking...  ...generation of our global services platform. At NVIDIA, you’ll keep...  ...lifecycle management, fleet reliability/auto-healing, E2E observability...  ..., or Ruby.Mentored other engineers and influenced technical direction... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $230k - $250k

     ...the future of network reliability, security, and AI‑ready...  ...the reliability engineering function at Forward —...  ...complex, distributed SaaS platform. You will work closely...  ...before customers do Lead incident response: on‑...  ...years of experience in site reliability engineering... 
    Night shift

    Forward

    Santa Clara, CA
    2 days ago
  • $132.6k - $214.5k

     ...delivers the industry’s most advanced SecOps platform, consisting of XDR, XSIAM, XSOAR, and...  ...you will collaborate closely with our engineering teams to develop innovative solutions...  ...of the product and ensure the reliability and availability of our services.... 
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks

    Santa Clara, CA
    3 days ago
  • $150.4k - $277.6k

     ...United States Software and Services The Media Platforms SRE team under the Apple Service Engineering division is one of the most exciting examples of Apple...  ...with 4+ years experience At least 6 years in a Reliability Engineering, DevOps or infrastructure focused role... 
    Relocation
    Day shift

    Apple

    Cupertino, CA
    4 days ago
  • $65 - $85 per hour

     ...graphics, PC gaming, and accelerated computing, to bring a Site Reliability Engineer (Contract) to the team based in Santa Clara, CA. This...  ...and resolve infrastructure issues in collaboration with the platform engineering team. Experience maintaining and setting up... 
    Full time
    Contract work
    Worldwide

    Sustainable Talent

    Santa Clara, CA
    3 days ago
  •  ...Location: 5 on-site days a week in Sunnyvale, CA...  ...infrastructure updates Lead incident response and...  ...to enhance reliability, scalability, and efficiency...  ...in computer science, Engineering, or related field; or...  ...AWS and/or Azure cloud platform ~ Hands‑on experience... 
    Work experience placement

    Illumio

    San Jose, CA
    3 days ago
  • $207k - $300k

     ...team members to enhance system reliability and efficiency.Initiate, own, and lead large-scale, complex projects...  ...techniques.3 years of experience as a Site Reliability Engineer.3 years of experience leading...  ...grow.Semantic Understanding Platform (SUP) is the standard platform... 

    Google

    San Jose, CA
    4 days ago
  • $192.4k - $275.8k

     ...demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service...  ...automation and frameworks that make the whole platform more resilient. If you are the kind...  ...audiences 4+ years experience leading post-mortems and root cause analysis... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    20 hours ago
  • $122.5k - $175k

     ..., resilient, and secure. The Zscaler Zero Trust Exchange️ platform protects thousands of customers from cyberattacks and data...  ...AI era? Join us at Zscaler.RoleWe are looking for a Staff Site Reliability Engineer to join our team. This is a hybrid role going into the San... 
    Full time
    Work at office
    Local area
    3 days per week

    Zscaler

    San Jose, CA
    20 hours ago
  • $207k - $300k

     ...consulting, developing software platforms and frameworks, capacity...  ...for changes that improve reliability and velocity.Practice sustainable...  ....3 years of experience leading projects.3 years of...  ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE... 

    Google

    San Jose, CA
    20 hours ago
  • $248k - $396.75k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing, building...  ...system performance, build scalable platforms, and continuously strengthen the reliability...  ...of NVIDIA’s AI Platform Runtime and lead reliability engineering initiatives... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $207.4k - $259.2k

     ...aerospace sector building an end-to-end advanced air mobility platform that delivers air taxis, unmanned aircraft systems (“UAS”)...  ...are seeking a highly experienced and passionate Sr. Staff Site Reliability Engineer (SRE) to join our growing team. In this critical role, you... 
    Permanent employment
    Local area
    Worldwide
    Visa sponsorship

    Archer Aviation

    San Jose, CA
    2 days ago
  • $60 - $62 per hour

     ...to enhance overall system reliability. Key Responsibilities: Deploying...  ...to customers facing platform challenges. Must-Have Skills...  ...3+ years of experience in Site Reliability Engineering. Proficiency with...  ...Join Akraya Today! Let us lead you to your dream career and... 
    Hourly pay
    Contract work
    Remote work

    Akraya

    Santa Clara, CA
    2 hours ago
  •  ...responsibilities of a Technical Support Engineer within a SaaS (Software as a Service) environment with a growing focus on Site Reliability Engineering (SRE). The ideal candidate has...  ..., and scale an AI Security Public SaaS platform, operating AI inference workloads at... 
    Work at office
    Local area
    Remote work
    Work from home

    F5 Networks

    San Jose, CA
    3 days ago
  •  ...Site Reliability Engineer Location – San Jose, CA What You'll Do - Responsibilities Engage in and improve the whole lifecycle of...  ...activities such as system design consulting, developing software platforms and frameworks, capacity planning and launch reviews.... 

    Netpace

    San Jose, CA
    3 days ago
  •  ...Site Reliability Engineer Foxconn Industrial Internet (Fii) is a world leading professional design and manufacturing service provider of communication network equipment...  ...centered on the Industrial Internet platform. Foxconn is currently seeking a Site Reliability... 
    Full time
    Work at office
    Local area

    FII

    San Jose, CA
    3 days ago
  • $65 - $85 per hour

     ...Site Reliability Engineer Sustainable Talent is partnering with a global leader who's been transforming computer graphics, PC gaming, and...  ...systems (Windows/Linux/Android), a multitude of hardware platforms both NVIDIA GPUs and Tegra Processors. Are you passionate... 
    Full time
    Contract work
    Worldwide

    Sustainable Talent

    Santa Clara, CA
    1 day ago
  • $200k - $300k

     ...are looking for an SDK/API & Developer Platform Lead to own the customer-facing software experience...  ...discussions. Hire, mentor, and lead engineers focused on SDK APIs, developer tools,...  ...developer portals, documentation sites, SDK installers, CLI tools, sample catalogs... 
    Flexible hours

    Jobleads-US

    Santa Clara, CA
    3 days ago
  •  ...Inc . [NYSE: IONQ] is the world's leading quantum platform and merchant supplier - delivering...  ...1874 The Role The Platform Engineering team builds, secures, and operates...  ...premises components deployed at customer sites. The Site Reliability Engineering discipline keeps the... 
    Work at office
    Remote work

    Physics World

    Santa Clara, CA
    3 days ago
  • $120k - $150k

     ...interested in working with the World's leading AI-first Quality Engineering Company? Ready to advance your...  ...! We are looking for a Site Reliability Engineer to join our growing team...  ...Skillsets - Expectations from Data Platform team: Expertise in Message Broker... 
    Casual work
    Local area
    Flexible hours

    QualiTest Group

    Santa Clara, CA
    4 days ago
  •  ...Velaura is seeking an SDK/API & Developer Platform Lead to own the customer-facing software experience for Velaura’s AI SoC, defining public...  ..., containers, and reference applications, while mentoring engineers and shaping the developer portal. #J-18808-Ljbffr Jobleads... 
    Flexible hours

    Jobleads-US

    Santa Clara, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineer, Platforms. Be the first to apply!