Staff Site Reliability Engineer, Waymo Fleet
$251k - $310kWaymo
Software Reliability Engineering : Waymo’s software reliability engineers (SRE) are responsible for the stable operation of Waymo’s fully autonomous systems and supporting infrastructure. As an SRE, you combine software and systems engineering techniques to build and run large-scale, fault-tolerant, reliable systems. You focus on optimizing existing systems, building new infrastructure, eliminating manual, error-prone or time-consuming work through automation, and ensuring products that are fast, efficient, and effective.
You Will- Serve as a Waymo production expert, and collaborate with other engineers to build reliable systems for autonomous vehicle operations, including depot logistics, automation flow, and critical vehicle state infrastructure.
- Manage end-to-end availability and performance for core fleet services, ensuring we have enough usable vehicle supply available to meet targeted user demand, and developing observability and automation to support this goal.
- Write designs and implement software to improve system architecture, telemetry or deployment for fleet-specific mission-critical services, preventing outages that could hinder vehicle launch or maintenance.
- Provide technical leadership and direction while troubleshooting highly complex system reliability issues and managing technical debt.
- Arbitrate in cases of technical disagreement among team members, SREs, SWE partners, and PMs to set global architectural guidelines.
- Lead cross-functional and cross-organizational collaborations to integrate disparate projects and processes into widely-reusable components and services.
- Champion operational excellence and blameless retrospectives, mentoring and developing leadership within the team to ensure efficient delivery of the vision.
- Drive operational standards for fleet and supply infrastructure by leading incident response efforts. You’ll participate in a sustainable on-call rotation, while championing a culture of blameless retrospectives to drive continuous improvement.
- 8+ years of experience architecting and maintaining mission-critical systems in C++, Java, or Python.
- Proven depth and breadth of knowledge in software reliability, acting as the "go-to" expert for resolving highly complex, high-impact system failures..
- Demonstrated track record of successfully leading the reliability program of a complex system (e.g., autonomous systems, cloud infrastructure, or large-scale web service.)
- Experience developing and communicating a shared vision, strategy, and priorities across multiple organizations and stakeholders.
- Proven ability to influence beyond the immediate area of responsibility, providing input to key decision makers that directly impact the future direction of the program.
- Demonstrated experience leading cross-functional initiatives between Engineering and Dev.
- A history of developing leadership within teams, including delegating, and recruiting/mentoring engineers to improve team scaling and execution.
- A Bachelor's degree in a relevant field or 10+ years similar experience in a high-growth environment with leadership experience.
- Demonstrated track record of leading multiple large projects or a mission-critical program from concept to delivery.
- Proven ability to anticipate and solve problems across functional groups, translating business needs into site reliability initiatives.
- Strong background in managing technical debt and system evolution, balancing business velocity with best reliability practices.
- Proven ability to lead cross-functional architectural investigations and resolve complex, high-impact system failures in distributed systems.
- A Master’s or PhD in Computer Science, Computer Engineering, or a related field.
- Once or twice a year travel to the Bay Area, US is preferred (but not required)
- ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and... ..., resizing, and incident response using fleet management toolsParticipate in a well-... ...Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE,...FleetWork at officeLocal areaWork from homeFlexible hours
- ...Plus meaningful equity All roles San Francisco, CA Site Reliability Engineer San Francisco, CAFull-timeMid to SeniorOn-site Zof AI... ...Reliability Engineer to run the infrastructure that lets fleets of sandboxed agents execute customer code safely and cheaply...FleetFull time
$148k - $235.75k
...lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer to operate the...FleetFull time$267k - $356k
...currently Tuesday.Lambda's Storage Engineering team is the backbone behind... ...in the industry, which means reliability and performance aren't just... ...Lambda's production storage fleet across all data centers,... ...storage across new and existing sites using tools such as Ansible,...FleetWork experience placementWork at officeLocal areaWork from homeFlexible hours$101k - $161k
...several prestigious awards, such as Best Engineering Team, Best Company for Diversity,... ...DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s... ...for our global CloudVision service fleet, ensuring scalability, reliability, and...Fleet$127k - $249k
Platform Engineering is the department within SRE that is responsible for a range of critical... ...observability and alerting systems.The Fleet Management team provides the core runtime... ...critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-...FleetWork at officeLocal areaRemote workWorldwideFlexible hours$152k - $241.5k
...infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (... ...languages such as Python, Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through design reviews,...FleetFull time- ...Anduril, Tesla, Uber, and the U.S. Special Forces. The Role We're hiring a Site Reliability Engineer to own the operational health of our connected sensor platform — spanning a live fleet of edge hardware deployed at customer sites and the cloud infrastructure...FleetRemote work
- ...Description Position Overview We are seeking a highly skilled Site Reliability Engineer (SRE) with hands-on experience in warehouse automation,... ..., including Concinity WES, PLC libraries, vision systems, fleet management tools, and real-time data exchange platforms....FleetRemote work
$207k - $300k
...mentoring team members to enhance system reliability and efficiency.Initiate, own, and... ...techniques.3 years of experience as a Site Reliability Engineer.3 years of experience leading... ...Experience in large-scale and secure fleet management of servers and components....Fleet$65 - $85 per hour
...graphics, PC gaming, and accelerated computing, to bring a Site Reliability Engineer (Contract) to the team based in Santa Clara, CA. This is... ...uncover real problems, and resolve them. Responsibilities Fleet monitoring and recovery of assets in the private cloud...FleetFull timeContract workWorldwide$170k - $195k
...States and its allies. About the Role As Picogrid's first Site Reliability Engineer you will own production reliability across cloud and edge,... ...response through node lifecycle, stateful workloads, and a fleet of hardware edge devices in the field. You will help build...FleetPermanent employmentWork at officeRemote workRelocation package- ...Global, and prominent developer-tool founders, including the founder of Snyk and former CTO of GitHub. About The Role Own compute fleet health end to end. Build the metrics pipelines, alerting, and unified health view that tell you the true state of every GPU in...Fleet
$164k - $270k
...looking for.The Role What You’ll DoOwn the reliability of our robotics systems, from PLCs... ...Partner with controls, robotics, and platform engineering teams to bake reliability in early.... ...OPC UA, EtherCAT, motion controllers, or fleet management for autonomous systemsAn individual...FleetPermanent employmentFull timeRelocation packageFlexible hours$150.4k - $277.6k
...Platforms SRE team under the Apple Service Engineering division is one of the most exciting... ...platforms - many run on our bare metal fleet which you will help transition to modern... ...years experience At least 6 years in a Reliability Engineering, DevOps or infrastructure focused...FleetRelocationDay shift- ...infrastructure that has to perform under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the foundation it runs on. This... ...ensuring consistency and reliability across the fleet.Improve system and network reliability and...Fleet
- ...About the Role We're looking for a Senior Site Reliability Engineer who is equally at home writing production software and running the infrastructure... ...problems on our platform: intelligently managing a large fleet of GPU-backed models. We run nearly 30 models across...FleetShift work
- ...actually happened, and it is ours. We’re looking for a Site Reliability Engineer to help advance MaintainX’s reliability, observability, and... ...mission is to keep the physical world running. Factories, fleets, hospitals and campuses stay up because the people who...FleetRemote work
$105k - $127k
...Site Reliability Engineer At ICEYE, we design, build and operate the largest fleet of Synthetic Aperture Radar (SAR) satellites in the world. Using advanced technology, our constellation collects topographical data about any location on Earth, day or night, through...FleetWork at officeLocal areaFlexible hoursNight shift$175k - $285k
...deployments, security configurations, and efficient operational workflows for our end-users.Own, administer, and optimize MDM platforms (Fleet DM, Intune, Workspace ONE) to enforce configuration and drive self-healing by writing OS-level scripts and lightweight tools that...FleetPermanent employmentFull timeRemote workRelocation packageFlexible hours$65 - $85 per hour
...Site Reliability Engineer Sustainable Talent is partnering with a global leader who's been transforming computer graphics, PC gaming, and accelerated... ...to have you onboard. What you'll be doing: Fleet monitoring & recovery of assets in our private cloud...FleetFull timeContract workWorldwide$157k - $239k
...Infrastructure /Full time /On-siteWanna join the adventure?As a Site Reliability Engineer with strong networking skills in our Cloud Infrastructure (... ...automating away toil so on-call load doesn’t grow with the fleet.Beyond the network, you’ll help the SRE team across its...FleetFull timeTemporary work$75.7k - $136.3k
...and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages... ...reliability, security, and usability at scale for our global fleet while maintaining Akamai's mission at the forefront of what...FleetWork experience placementWork at office$121.4k - $218.6k
...and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages... ...reliability, security, and usability at scale for our global fleet while maintaining Akamai's mission at the forefront of what...FleetWork experience placementWork at office$195k - $285k
...the infrastructure underpinning our engineering organization must be as reliable and scalable as the chips we build.... ...role builds and leads d-Matrix's Site Reliability Engineering function from... ...— host lifecycle automation, fleet auto-remediation, and AIOps-driven...FleetRemote work$175k - $265k
...owns the infrastructure layer that every engineering team and customer depends on —... ...core member of that team, responsible for reliability, automation, and observability across colo... ...reliability and availability across colo server fleets, on-premises lab clusters, cloud...Fleet- ...to keep the electric grid secure and reliable, even during extended periods of stress... ...Description Form Energy is hiring a Manager, Site Reliability Engineer to lead the operational function... ...own the mechanisms used to monitor fleet health, respond to incidents, triage field...FleetFull timeRemote workRelocation package
- ...services that empower utilities, cities, fleets, transit agencies, and automakers to... .... We are looking for a Manager of Site Reliability Engineering to join our team. The role requires a... ...enterprise. What You Will Do Help staff and lead a team of enthusiastic,...FleetCasual workFlexible hours
$119k - $170k
...in the AI era? Join us at Zscaler.RoleWe are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid... ...billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write...FleetFull timeWork at officeLocal areaRemote workShift work3 days per week- ...for an experienced Technical Program Manager to lead capacity planning and fleet strategy for our Inference Service organization. This is a highly visible role working directly with Engineering, Product, Infrastructure, SRE, Operations, and executive leadership to maximize...FleetRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Site Reliability Engineer, Waymo Fleet. Be the first to apply!
- assistant engineer California
- senior staff systems engineer California
- technology administrator California
- engineering aide California
- senior staff engineer California
- software engineer staff California
- staff engineer California
- official site California
- site services specialist California
- construction site safety California


