Staff Site Reliability Engineer-Production Operations
Socket.dev
About Us Rivian and Volkswagen Group Technologies is a joint venture between two industry leaders with a clear vision for automotive’s next chapter. From operating systems to zonal controllers to cloud and connectivity solutions, we’re addressing the challenges of electric vehicles through technology that will set the standards for software-defined vehicles around the world. The road to the future is uncharted. By combining our expertise across connectivity, AI, security and more, we’ll map a new way forward. Working together, we’ll create a future that’s more connected, more intelligent, more sustainable for everyone. Role Summary We are looking for an SRE Lead to serve as the senior technical leader and player-coach for ProdOps. This is a hybrid role: you will set the technical direction of the team and lead from the front during incidents, while also growing and managing a small group of exceptional engineers as the function scales. As the calm center during a crisis, you will maintain a high-level mental model of the entire production ecosystem, freeing engineers to focus strictly on debugging and mitigation. You will own the reliability feedback loop end to end, which includes running blameless post-incident reviews, coordinating major incidents across Cloud systems, vehicle development pipelines, and Product Security, driving the systemic action-item backlog with TPMs and development teams, and verifying that fixes hold in production. This role suits a senior systems engineer who still writes code, thinks in terms of systems and failure modes rather than single root causes, and wants to build a lean, automation-first reliability practice rather than staff a support queue. Responsibilities Lead incident coordination and communication Drive incident progress and coordinate cross-functional response as the central nervous system during a crisis. Own executive, customer, and parent-company communications, providing production expertise and communication leadership so engineers can concentrate on the technical problem. Maintain an accurate, high-level model of the full production ecosystem spanning Cloud, vehicle development pipelines, and Product Security. Facilitate modern, blameless post-incident learning Run post-incident reviews using Learning From Incidents (LFI) principles and HOWIE-style reporting. Move the organization away from the search for a single root cause and toward understanding how tooling, context, and multiple latent conditions combined to produce failure. Surface weaknesses in observability, process, testing, and tooling that impaired our ability to detect, mitigate, and recover. Own and prioritize systemic action items Manage the backlog of action items generated by reviews. While ProdOps does not write the fixes itself, you will prioritize, track, and drive these items to closure in partnership with TPMs and development teams, keeping leadership focused on customer impact. Measure and verify efficacy Once the fixes ship, measure and verify that they actually prevent recurrence. Drive accountability for outcomes and feed the results back into the Novel Incident Rate. Build the automation and observability backbone Design and build the systems, tooling, and AI-agent workflows that automate incident triage and administrative toil. Improve observability, including instrumentation, alerting, dashboards, and SLOs, so incidents are detected faster and understood more deeply. Write and review code where it multiplies the team's impact. Lead and grow the team Set technical standards and operating rhythm for ProdOps. Mentor and develop engineers, and as the team scales, take on hiring and people management while preserving the minimal-headcount, maximum-automation philosophy. Qualifications Minimum Qualifications: 6+ years of experience in SRE, production/platform engineering, or systems engineering for large-scale distributed systems, including senior technical leadership or lead responsibilities. Proven incident command experience: you have coordinated high-severity, cross-functional incidents and led communication with executives and external stakeholders under pressure. Strong systems engineering fundamentals, including distributed systems, networking, cloud infrastructure, and an instinct for how complex systems fail. Hands-on coding ability (e.g., Python, Go, or similar) sufficient to build automation, tooling, and integrations. This is not a code-free management role. Deep observability expertise: instrumentation, metrics, logging, tracing, alerting, dashboards, and SLO/SLI design (Datadog or comparable platforms). Fluency in modern reliability and post-incident practice: blameless reviews, Learning From Incidents (LFI), HOWIE, and systemic (non-single-root-cause) analysis. A demonstrated bias toward eliminating toil through automation, and enthusiasm for using AI agents as force multipliers. Excellent written and verbal communication; able to translate technical detail for executive and cross-company audiences. Experience mentoring engineers, with the judgment and appetite to grow into formal people management. Preferred Qualifications Experience across both cloud services and hardware/vehicle or embedded development pipelines. Exposure to Product Security or working closely with security teams during incidents. Track record building an SRE or reliability function from an early stage. Experience deploying LLM- or agent-based automation into production operational workflows. Total Rewards We build the exceptional — and we believe the people doing that work should be rewarded accordingly. In addition to a competitive base salary, full-time positions may be is eligible to participate in our annual company performance bonus program. Payments are discretionary and not guaranteed; actual amounts depend on company results and the terms of the plan in effect, and require active employment at the time of payout. This role is also eligible for equity in the form of Restricted Stock Units (RSUs), subject to board approval and the terms of our equity incentive plans, including applicable vesting requirements. In addition to our compensation programs, we invest in our people with a comprehensive benefits package designed to support the health, wellbeing, and financial future for full-time employees — including health coverage, retirement savings, time off, and family planning programs. Offerings vary by country. Learn more about our global benefit programs . Equal Opportunity Rivian and Volkswagen Group Technologies is committed to creating a diverse environment and is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, ancestry, sex, sexual orientation, gender, gender expression, gender identity, genetic information or characteristics, physical or mental disability, marital/domestic partner status, age, military/veteran status, medical condition, or any other characteristic protected by law. We are also committed to ensuring compliance with all applicable fair employment practice laws regarding citizenship and immigration status. Rivian and Volkswagen Group Technologies is committed to ensuring that our hiring process is accessible for persons with disabilities. If you have a disability or limitation, such as those covered by the Americans with Disabilities Act, that requires accommodations to assist you in the search and application process, please email us at View email address on click.appcast.io. Candidate Data Privacy Rivian and Volkswagen Group Technologies may collect, use and disclose your personal information or personal data (within the meaning of the applicable data protection laws) when you apply for employment and/or participate in our recruitment processes (“Candidate Personal Data”). This data includes contact, demographic, communications, educational, professional, employment, social media/website, network/device, recruiting system usage/interaction, security and preference information. Rivian and Volkswagen Group Technologies may use your Candidate Personal Data for the purposes of (i) tracking interactions with our recruiting system; (ii) carrying out, analyzing and improving our application and recruitment process, including assessing you and your application and conducting employment, background and reference checks; (iii) establishing an employment relationship or entering into an employment contract with you; (iv) complying with our legal, regulatory and corporate governance obligations; (v) record keeping; (vi) ensuring network and information security and preventing fraud; and (vii) as otherwise required or permitted by applicable law. Rivian and Volkswagen Group Technologies may share your Candidate Personal Data with (i) internal personnel who have a need to know such information in order to perform their duties, including individuals on our People Team, Finance, Legal, and the team(s) with the position(s) for which you are applying; (ii) Rivian and Volkswagen Group Technologies affiliates; and (iii) Rivian and Volkswagen Group Technologies’ service providers, including providers of background checks, staffing services, and cloud services. Rivian and Volkswagen Group Technologies may transfer or store internationally your Candidate Personal Data, including to or in the United States, Canada, and the European Union and in the cloud, and this data may be subject to the laws and accessible to the courts, law enforcement and national security authorities of such jurisdictions. Rivian and Volkswagen Group Technologies may use your personal information as part of your application or during the recruitment process. Please see our Candidate Data Privacy Notice (English) and Candidate Data Privacy Notice (Serbian) for more information. -- #J-18808-Ljbffr Socket.dev
$147k - $210k
Write product or system development code.Review code developed by other engineers and provide feedback to ensure best practices (e... ...hardware, network, or service operations and quality. Participate... ...scale distributed systems. Site Reliability Engineering (SRE) is what...Operations$174k - $252k
...through to deployment, operation, and refinement.... ...changes that improve reliability and velocity.Practice... ...in Computer Science, Engineering, a related field, or equivalent... ...or Engineering.Site Reliability Engineering... ...billions. We own those products in production. We drive...Operations$184k - $287.5k
At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability. This demanding position... ...efforts to guarantee flawless service operation with consistent reliability and uptime...OperationsFull time$189k - $274k
...accessible for all. We’re searching for a Senior Staff Vehicle Hardware Verification and Validation Engineer Lead.As the Senior Staff Vehicle Hardware Verification... ...with design, manufacturing, software, operations and safety teams to align on test requirements, schedules...OperationsWork experience placementWork at officeLocal area3 days per week$165k - $280k
...goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging... .... We design, build, test, and operate all parts of the system - thousands of... ...operable, scalable, and maintainable products Engage throughout the whole software...OperationsPermanent employmentTemporary workWorldwideWeekend work$210.6k - $305.1k
...ThousandEyes FedRAMP team builds and operates our US GovCloud platform.... ..., including development, product management, and security to define... ...led a distributed team of 5+ engineers, can demonstrate strong... ...Please see the Cisco careers site to discover more benefits and...OperationsFull timeTemporary workLocal areaFlexible hours- Elevate your engineering prowess to unprecedented levels by joining... ...among the top echelon in site reliability. As a Senior Lead Site Reliability... ...in your application and product lines. You will ensure... ...accelerate reliability design and operational decisioning (e.g., incident...Operations
$203.45k - $344.3k
...expansion, and safe, compliant, and auditable operations. Lead the design of a highly available... ..., data architecture and closed-loop engineering system.Job ResponsbilitiesResponsible... ...full-process traceability of data from production, processing to use; establish strict...OperationsFull timeTemporary workWork experience placement- ...in a realm tailored for top achievers in site reliability.As a Lead Site Reliability Engineer at JPMorgan Chase within the Network Product, you hold a leadership role in your team,... ...network reliability principles (Permit to Operate, FMEA, operational readiness), balancing...Operations
$253k - $416k
...of the global workforce. Our products help people make powerful... ...DescriptionDistinguished Software Engineer, Systems Infrastructure -... ...balancing business impact, operational impact and cost benefits of... ...architecture, balancing scalability, reliability, and costHelp evolve...OperationsFor contractorsWork experience placementWork at officeFlexible hours$230k - $250k
...mathematically accurate model of the production network. It's the... ...autonomous networking, giving engineers and AI agents the ability... ....Forward is looking for a Site Reliability EngineerAbout the Role This... ...observability, incident response, and operational excellence across a complex...Night shift$150k - $177.5k
Senior Reliability Engineer - Power Electronics Reliability Engineers at Lunar Energy will be responsible for ensuring product reliability throughout the entire lifecycle of our revolutionary... ...Working knowledge in theory of operation of common power converter topologies...OperationsFull time$100k - $200k
...Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you... ..., automation, and building reliable, production-grade environments. Key Responsibilities Operate and maintain Linux-based systems, ensuring high...OperationsFull time- ...strategic and execution-focused Sr/Staff Business Enablement... ...digital transformation of our production environments. You will architect... ...driven, and integrated with Engineering and Supply Chain. The ideal... ...with both technical and operational teams to ensure business processes...OperationsTemporary workImmediate startRelocation package
$144k - $209k
...standards are enhanced for production.Lead and guide Subject... ....Ensure manufacturing sites are prepared for... ...with the Product Quality Engineer to establish product quality and reliability goals. Validate product... ...world. The Manufacturing Operations team is responsible for...OperationsContract workWorldwide$142.8k - $274.8k
...yearEmployment type: Full-TimeWork site: 3 days / week in-... ...AI Infrastructure Engineering Systems team at... ...to develop, deploy, and operate AI services at scale. We... ...closely with AI researchers, product engineers, and hardware teams, we provide reliable and consistent paths...OperationsOngoing contractLocal area3 days per week$145k - $200k
...for data-driven decisions and operations. By bringing the right data to... ...more.The Role We are a software engineering team with expertise in enabling ML models in production. We deploy AI models to run in... ...enabling fast, secure, and reliable model rollouts across on-premises...OperationsFull timeWork experience placementWork at officeRemote workWork from homeRelocation package$209k - $313k
...have fun together.The Company operates Snapchat, a visual messaging... ...services.We are looking for a Staff, Revenue Growth to join Snap... ...with a lens akin to driving product-led growth. They will have a... ...Marketing, Operations, Product, Engineering, Advertiser Support, and Data...OperationsFull timeLive inWork at officeLocal areaShift work- ...right place.As a Principal Engineer at JPMorganChase within the... ...up an isolated, composable, production-representative virtual environment... ...Argo CDEnsure the platform operates reliably across hybrid infrastructure... ...health care coverage, on-site health and wellness centers,...OperationsShift work
$170k - $199k
...do. Expectations are high, and so are the rewards.As a Staff Product Manager on the international team, you'll manage global... ...experiences in new geos, and work with experts in data, design, engineering, marketing, operations, and research to bring ideas to life that will help us...OperationsWork at officeLocal areaFlexible hoursShift work3 days per week$174k - $252k
...inputs and physical synthesis tool runs.Engineer modular carve-outs of our flow to run seamlessly... ..., maintaining, or launching software products, and 1 year of experience with software... ...in multiple areas of machine learning, operation research, compiler construction,...Operations$169k
...reputation for being a leading and reliable force in global commerce.We... ...and experienced Senior Staff Technical Program Manager to... ...members of company-wide business & Engineering teams to help define,... ...Define, iterate and improve operation excellence in large organization...OperationsTemporary work- ...together some of the strongest AI Engineers and Machine Learning... ...systems, foundation models, and production ML infrastructure. We build... ..., and production-grade reliability. A data-centric mindset... ...R&D prototyping to scalable operation in multiple production stacks...OperationsWork at office
$98k - $113.5k
...s Applications team is looking for a versatile, collaborative engineer to join our group. You will help lead EPRI's efforts to apply... ...expertise to client-specific challenges in grid planning and operations, distributed energy resources, electric transportation, and other...OperationsFull timeWork at officeRemote workWork from homeHome officeRelocation packageFlexible hours$175k - $186k
...progressing toward general release and scalable production in early 2026. Mobility is one of the... ...grow with us. As part of the Firmware Engineering team, you will drive the development... ...decisions and implementation Translate operational needs into intuitive interface behavior...OperationsWork at office$170k - $198k
...electronics. We capture digital exhaust and engineering context from assembly lines—images,... ...rely on Instrumental to accelerate new product introduction and production. As a... ...Depth: Work in depth across engineering, operations, and quality to drive solutions within...OperationsImmediate start$170k - $240k
...Commure, we're building the AI Operating System for healthcare, the... ...ownership from early thinking to production. If you're energized by hard... ...matters if we can't reliably translate a clinical encounter... ...wrong person. Scale a rule engine that runs hundreds of configurable...OperationsImmediate start$200k - $287.5k
...done. Snowflake’s Release Engineering team builds and operates the systems that safely... ...infrastructure, platform, and product changes to production at... ..., distributed systems reliability, and large-scale multi-cloud... ...on the Snowflake Careers Site for salary and benefits information...OperationsFlexible hours$170k - $240k
...Commure, we're building the AI Operating System for healthcare, the... ...ownership from early thinking to production. If you're energized by hard... ...production. Scale a rule engine that runs hundreds of... ...insurance type, denial category, and site. Generate appeal letters...OperationsWork at officeImmediate start- ...we're building the AI Operating System for healthcare,... ...from early thinking to production. If you're energized... ...Insights—provides practice staff and administrators... ...As a Senior Software Engineer Engineer, you will play... ...improvement, enhancing the reliability, scalability, and...OperationsImmediate start
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Site Reliability Engineer-Production Operations. Be the first to apply!
- assistant engineer Palo Alto, CA
- staff engineer Palo Alto, CA
- staff data engineer Palo Alto, CA
- software engineer staff Palo Alto, CA
- senior staff engineer Palo Alto, CA
- senior staff systems engineer Palo Alto, CA
- technology administrator Palo Alto, CA
- engineering aide Palo Alto, CA
- site reliability engineer Palo Alto, CA
- site reliability engineer sre Palo Alto, CA


