Lead Site Reliability Engineer, Platforms
$124k - $271.2kZoom Video Communications, Inc.
What You Can ExpectAs a Lead Staff Site Reliability Engineer, you will be one of the technical leads for our DevOps Platforms organization. This group is responsible for DevOps Platforms including cloud infrastructure, physical data center orchestration, critical security services, and our Zoom for Government (ZfG) environment. You will be an uber tech lead working across a broad area, defining projects and guiding work across various teams. Your scope of work is wide and you will have the opportunity to improve our datacenter kubernetes infrastructure, our cloud infrastructure, our security posture, and our operation of ZfG environments. Broadly speaking, you are an exemplary SRE and you will guide all of our teams toward SRE best practices (automation, monitoring, infrastructure as code, etc).About the TeamThe DevOps Platforms organization owns the full infrastructure stack: cloud infrastructure on AWS and OCI, physical data center orchestration, critical security services (identity, authentication, authorization), and Zoom's FedRAMP-rated federal environment, Zoom for Government (ZfG).The team is currently working on one of the most technically interesting work in the org, hardening our security posture, and building the automation and reliability systems that underpin Zoom's global services. If you want broad visibility, real cross-team influence, and the chance to define how infrastructure gets built, this is the seat.ResponsibilitiesDesign and scale DevOps platform services including Kubernetes infrastructure, cloud systems, and compliance-ready environmentsDefine technical roadmaps and architectural direction for infrastructure automation and securityPartner with service teams to understand platform needs and deliver solutions that improve reliability and efficiencyEstablish and advocate for SRE best practices including infrastructure as code, monitoring, and incident managementMentor team members through design, implementation, and production deployment of complex systemsWhat We're Looking ForBring 8+ years of SRE or DevOps experience building and operating production infrastructure at scaleCode proficiently in at least one programming language beyond scripting (e.g., Python, Go, Java)Deploy and manage CI/CD pipelines using tools like Git, Jenkins, Argo CD, or JFrogOperate cloud infrastructure on AWS, OCI, or similar platforms using Terraform and KubernetesImplement observability solutions with logging and monitoring tools such as ELK, Prometheus, or GrafanaCommunicate complex technical concepts clearly to diverse audiences including security teams, senior leadership, and external auditorsParticipate in on-call rotations and lead incident response to maintain system reliabilityHold a degree in Computer Science or related field, or equivalent practical experienceHold US citizenship, or Greencard statusPreferredHave experience with security from an SRE perspectiveHave experience with Identity security (e.g. IAM, workload identity, zero trust) and tools (e.g. Teleport, Okta)Have experience operating Government environments and understanding their compliance requirementsHave experience with system design and distributed computing at scaleAbility to speak Chinese/Mandarin is a plus, but not requiredSalary Range or On Target Earnings:Minimum:$124,000.00Maximum:$271,200.00 In addition to the base salary and/or OTE listed Zoom has a Total Direct Compensation philosophy that takes into consideration; base salary, bonus and equity value.Note: Starting pay will be based on a number of factors and commensurate with qualifications & experience.We also have a location based compensation structure; there may be a different range for candidates in this and other locationsAt Zoom, we offer a window of at least 5 days for you to apply because we believe in giving you every opportunity. Below is the potential closing date, just in case you want to mark it on your calendar. We look forward to receiving your application!Anticipated Position Close Date:10/02/26Ways of WorkingOur structured hybrid approach is centered around our offices and remote work environments. The work style of each role, Hybrid, Remote, or In-Person is indicated in the job description/posting.BenefitsAs part of our award-winning workplace culture and commitment to delivering happiness, our benefits program offers a variety of perks, benefits, and options to help employees maintain their physical, mental, emotional, and financial health; support work-life balance; and contribute to their community in meaningful ways. Click Learnfor more information.About UsZoomies help people stay connected so they can get more done together. We set out to build the best collaboration platform for the enterprise, and today help people communicate better with products like Zoom Contact Center, Zoom Phone, Zoom Events, Zoom Apps, Zoom Rooms, and Zoom Webinars.We’re problem-solvers, working at a fast pace to design solutions with our customers and users in mind. Find room to grow with opportunities to stretch your skills and advance your career in a collaborative, growth-focused environment.Our CommitmentAt Zoom, we believe great work happens when people feel supported and empowered. We’re committed to fair hiring practices that ensure every candidate is evaluated based on skills, experience, and potential. If you require an accommodation during the hiring process, let us know—we’re here to support you at every step.If you need assistance navigating the interview process due to a medical disability, please submit an Accommodations Request Form and someone from our team will reach out soon. This form is solely for applicants who require an accommodation due to a qualifying medical disability. Non-accommodation-related requests, such as application follow-ups or technical issues, will not be addressed.Our interviews are supported by BrightHire, a tool that helps us create a consistent and thoughtful interview experience and may include recordings. Please refer to our candidate privacy statement for more information of how we use your data.SummaryLocation: San Jose (CA)Type: Full time
$168k - $270.25k
Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline to design, build and maintain... ...learn and grow.What you’ll be doing:Lead the technical strategy and roadmap for... ...develop AI Agents, AI Skills to accelerate platform operationsDrive automation and...SuggestedFull time$104.9k - $174.7k
...responsible for improving the reliability, availability,... ...completion.Follow up with engineering, development, security... ...and operational tasks.Lead or contribute to... ...years of experience in Site Reliability Engineering... ...PlanWellbeing: Wellness platform with incentives,...SuggestedFull timeLocal area$168k - $270.25k
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization... ...intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing...SuggestedFull time- ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building... ...tooling and automate the validation of platform quality.Design, build, and maintain scalable... ...services, workloads, and platform reliability.You6+ years of experience in a SRE, operations...SuggestedWork at officeLocal areaWork from homeFlexible hours
- ...home day is currently Tuesday.Engineering at Lambda is responsible for... ...multi-tenant cloud networking platform and SDN infrastructureOperate... ...teams to improve service reliability and deployment workflowsDeploy... ...rotationYouHave 5+ years of experience in Site Reliability Engineering,...SuggestedWork at officeLocal areaWork from homeFlexible hours
$148k - $235.75k
...the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for... ...checks, post-deploy validation), and lead rollbacks/remediations when needed....Full time$230k - $250k
...foundation for autonomous networking, giving engineers and AI agents the ability to know the... ...comfortable, building a groundbreaking platform that transforms how teams run and... ...always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "...Night shift$101k - $161k
...prestigious awards, such as Best Engineering Team, Best Company for... ...Work WithWe’re looking for Site Reliability Engineers to join our growing... ...Familiarity with GCP (Google Cloud Platform) and GKE (Google Kubernetes... ...to be drive, develop, and lead projects in any of the...$267k - $356k
...currently Tuesday.Lambda's Storage Engineering team is the backbone behind... ...spectrum of Lambda's data platform services—from low-level... ...in the industry, which means reliability and performance aren't just... ...storage across new and existing sites using tools such as Ansible,...Work experience placementWork at officeLocal areaWork from homeFlexible hours$152k - $241.5k
...and amazing people. NVIDIA is leading the way in groundbreaking... ...generation of our global services platform. At NVIDIA, you’ll keep... ...lifecycle management, fleet reliability/auto-healing, E2E observability... ..., or Ruby.Mentored other engineers and influenced technical direction...Full time$230k - $250k
...the future of network reliability, security, and AI‑ready... ...the reliability engineering function at Forward —... ...complex, distributed SaaS platform. You will work closely... ...before customers do Lead incident response: on‑... ...years of experience in site reliability engineering...Night shift$132.6k - $214.5k
...delivers the industry’s most advanced SecOps platform, consisting of XDR, XSIAM, XSOAR, and... ...you will collaborate closely with our engineering teams to develop innovative solutions... ...of the product and ensure the reliability and availability of our services....Full timeWork at officeVisa sponsorshipWork visa$150.4k - $277.6k
...United States Software and Services The Media Platforms SRE team under the Apple Service Engineering division is one of the most exciting examples of Apple... ...with 4+ years experience At least 6 years in a Reliability Engineering, DevOps or infrastructure focused role...RelocationDay shift$65 - $85 per hour
...graphics, PC gaming, and accelerated computing, to bring a Site Reliability Engineer (Contract) to the team based in Santa Clara, CA. This... ...and resolve infrastructure issues in collaboration with the platform engineering team. Experience maintaining and setting up...Full timeContract workWorldwide- ...Location: 5 on-site days a week in Sunnyvale, CA... ...infrastructure updates Lead incident response and... ...to enhance reliability, scalability, and efficiency... ...in computer science, Engineering, or related field; or... ...AWS and/or Azure cloud platform ~ Hands‑on experience...Work experience placement
$207k - $300k
...team members to enhance system reliability and efficiency.Initiate, own, and lead large-scale, complex projects... ...techniques.3 years of experience as a Site Reliability Engineer.3 years of experience leading... ...grow.Semantic Understanding Platform (SUP) is the standard platform...$192.4k - $275.8k
...demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service... ...automation and frameworks that make the whole platform more resilient. If you are the kind... ...audiences 4+ years experience leading post-mortems and root cause analysis...Full timeTemporary workLocal areaFlexible hours$122.5k - $175k
..., resilient, and secure. The Zscaler Zero Trust Exchange️ platform protects thousands of customers from cyberattacks and data... ...AI era? Join us at Zscaler.RoleWe are looking for a Staff Site Reliability Engineer to join our team. This is a hybrid role going into the San...Full timeWork at officeLocal area3 days per week$207k - $300k
...consulting, developing software platforms and frameworks, capacity... ...for changes that improve reliability and velocity.Practice sustainable... ....3 years of experience leading projects.3 years of... ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE...$248k - $396.75k
Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing, building... ...system performance, build scalable platforms, and continuously strengthen the reliability... ...of NVIDIA’s AI Platform Runtime and lead reliability engineering initiatives...Full time$207.4k - $259.2k
...aerospace sector building an end-to-end advanced air mobility platform that delivers air taxis, unmanned aircraft systems (“UAS”)... ...are seeking a highly experienced and passionate Sr. Staff Site Reliability Engineer (SRE) to join our growing team. In this critical role, you...Permanent employmentLocal areaWorldwideVisa sponsorship$60 - $62 per hour
...to enhance overall system reliability. Key Responsibilities: Deploying... ...to customers facing platform challenges. Must-Have Skills... ...3+ years of experience in Site Reliability Engineering. Proficiency with... ...Join Akraya Today! Let us lead you to your dream career and...Hourly payContract workRemote work- ...responsibilities of a Technical Support Engineer within a SaaS (Software as a Service) environment with a growing focus on Site Reliability Engineering (SRE). The ideal candidate has... ..., and scale an AI Security Public SaaS platform, operating AI inference workloads at...Work at officeLocal areaRemote workWork from home
- ...Site Reliability Engineer Location – San Jose, CA What You'll Do - Responsibilities Engage in and improve the whole lifecycle of... ...activities such as system design consulting, developing software platforms and frameworks, capacity planning and launch reviews....
- ...Site Reliability Engineer Foxconn Industrial Internet (Fii) is a world leading professional design and manufacturing service provider of communication network equipment... ...centered on the Industrial Internet platform. Foxconn is currently seeking a Site Reliability...Full timeWork at officeLocal area
$65 - $85 per hour
...Site Reliability Engineer Sustainable Talent is partnering with a global leader who's been transforming computer graphics, PC gaming, and... ...systems (Windows/Linux/Android), a multitude of hardware platforms both NVIDIA GPUs and Tegra Processors. Are you passionate...Full timeContract workWorldwide$200k - $300k
...are looking for an SDK/API & Developer Platform Lead to own the customer-facing software experience... ...discussions. Hire, mentor, and lead engineers focused on SDK APIs, developer tools,... ...developer portals, documentation sites, SDK installers, CLI tools, sample catalogs...Flexible hours- ...Inc . [NYSE: IONQ] is the world's leading quantum platform and merchant supplier - delivering... ...1874 The Role The Platform Engineering team builds, secures, and operates... ...premises components deployed at customer sites. The Site Reliability Engineering discipline keeps the...Work at officeRemote work
$120k - $150k
...interested in working with the World's leading AI-first Quality Engineering Company? Ready to advance your... ...! We are looking for a Site Reliability Engineer to join our growing team... ...Skillsets - Expectations from Data Platform team: Expertise in Message Broker...Casual workLocal areaFlexible hours- ...Velaura is seeking an SDK/API & Developer Platform Lead to own the customer-facing software experience for Velaura’s AI SoC, defining public... ..., containers, and reference applications, while mentoring engineers and shaping the developer portal. #J-18808-Ljbffr Jobleads...Flexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer, Platforms. Be the first to apply!
- lead engineer San Jose, CA
- lead operating engineer San Jose, CA
- site reliability engineer sre San Jose, CA
- site reliability engineer San Jose, CA
- client platform engineer San Jose, CA
- senior platform engineer San Jose, CA
- platform engineering manager San Jose, CA
- platform developer San Jose, CA
- platform engineer San Jose, CA
- official site San Jose, CA


