Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Software Engineer - Reliability (US Citizen Only)

$218.3k - $327.5k

Rubrik

About Team & About RoleThe Site Reliability Engineering (SRE) team at Rubrik ensures the absolute reliability, availability, performance, and security of our enterprise infrastructure services, spanning both global SaaS platforms and government-compliant environments. We operate at the intersection of software development and systems engineering, prioritizing hyperscale platform automation, self-healing architectures, and structural resiliency. As a Staff Site Reliability Engineer, you will serve as a primary technical leader and architect across our broader distributed cloud systems. You will drive long-term technical roadmaps, establish cross-organizational reliability standards, and solve complex distributed systems challenges that safeguard both enterprise and public sector environments. Beyond the core SRE charter, this Staff role also leads the Application-SRE team — a US-based group that partners closely with engineering, Sales, and Support to unblock POCs, drive complex customer escalations to resolution, and convert recurring field signals into engineering and reliability roadmap items. You will be the technical leader and project owner for Application-SRE: setting direction, tracking commitments, and ensuring the team operates as a high-leverage bridge between the field and the broader engineering org.What You'll DoAs a Staff Site Reliability Engineer, you will possess engineering-wide influence and take ownership of the following critical areas:Infrastructure Strategy & Architecture: Formulate and execute the architectural vision for Rubrik's Cloud Platform, optimizing backend infrastructure systems like Kubernetes, MySQL, and cloud-native services for performance, security, and multi-region scale.Hyperscale Automation & Platform Tooling: Build, scale, and maintain sophisticated custom internal tools, platform controllers, and automation frameworks in Go or Python to systematically eliminate operational toil.AI Infrastructure for SaaS: Deploy, scale, and operate the AI infrastructure that powers Rubrik's SaaS offerings, owning the reliability, performance, cost, and security controls required to run AI workloads in multi-tenant, compliance-bound environments.AI for SRE & Engineering Productivity: Drive the adoption of AI-driven solutions across the SRE charter to compress toil and multiply the org - applying agentic and LLM-based approaches to automated triage, incident response, operational analysis, and developer productivity.AI Adoption Guardrails for SaaS Reliability: Build the guardrails, controls, and platform patterns that keep Rubrik's SaaS reliable as AI adoption accelerates across product and engineering, ensuring new AI capabilities ship without eroding availability, performance, security, or cost posture.Cross-Functional Leadership: Wield engineering-wide influence to create technical consensus among component, platform, and security engineering teams, effectively "shifting left" to embed structural resilience, capacity guards, and compliance from initial feature designs.Reliability Governance: Define, audit, and enforce robust Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets across all critical enterprise platform services, translating telemetry insights into actionable product roadmaps during executive reviews.Incident Command & Operations Review: Serve as a primary Incident Commander for high-severity cloud outages, establishing roles, directing mitigation vectors under pressure, and orchestrating comprehensive, blameless post-mortems that drive durable systemic fixes.Cost Governance & Capacity Modeling: Architect cost-observability tools and attribution frameworks, leading cloud infrastructure capacity forecasting, resource quota optimization, and vendor SLA management.Application-SRE Leadership: Set the technical direction for the Application-SRE team, raising the bar on how the team diagnoses, mitigates, and durably resolves the most complex customer-impacting issues across our platform.Technical Multiplier & Mentorship: Champion SRE best practices, mentoring senior and junior individual contributors across the organization, participating in interview frameworks, and actively raising the collective technical bar.On-Call Rotations: Participate in on-call rotationsExperience You'll NeedCitizenship & Residency: Must be a US Citizen currently residing on CONUS soil (strict regulatory requirement to enable support for federal and FedRAMP environments when required).Education: BS, MS, or PhD in Computer Science, Computer Engineering, or a highly related technical discipline.Industry Experience: A minimum of 8–12+ years of software engineering and production cloud infrastructure experience, with at least 5+ years dedicated to a formal SRE, DevOps, or Platform engineering role operating hyperscale SaaS products.Technical Depth: Comprehensive, hands-on programming expertise in Golang, Python, or Java with a deep grasp of concurrency models, data structures, and test-driven software design patterns.Distributed Systems Expertise: Proven proficiency designing, deploying, analyzing, and auditing complex, large-scale distributed systems, database topologies, and high-availability public cloud meshes.Systems Internals: Authoritative operational command of Unix/Linux operating system environments (process models, file systems, kernels), systems administration, and advanced L4/L7 networking protocols.AI Systems Fluency: Working knowledge of operating AI systems in production — including model serving, cost trade-offs, and the reliability and safety considerations of LLM- and agent-based workloads. Practical judgment on when AI is the right tool versus deterministic automation.Field-to-Product Feedback Loop: Institutionalize the channel that converts patterns from customer escalations and POCs into prioritized product and reliability feedback, partnering directly with Product, Sales Engineering and Support leadership.Customer & Field Fluency: Track record of partnering directly with Sales, Support, and customers on escalations and POCs, and translating field signals into engineering action.Leadership Capability: Demonstrated history of technical leadership, mapping architectural dependencies, managing multi-team technical projects, and guiding organizations through critical platform shifts with high technical judgment.Preferred QualificationsExtensive production experience provisioning, lifecycle-managing, and recovering enterprise-scale Kubernetes (GKE, EKS) deployments and large-scale relational/non-relational databases (MySQL).Prior experience building, certifying, or auditing infrastructure environments under compliance structures such as FedRAMP (High/Moderate), SOC 2, ISO 27001, or CJIS.Fluency in Infrastructure-as-Code (Terraform, Pulumi) module design, multi-tenant state isolation, and enterprise observability fabrics (Prometheus, Grafana, OpenTelemetry).Exposure to building AI- or LLM-powered internal tooling and applying it to SRE, operations, or engineering productivity use cases.Familiarity with the operational considerations of running AI workloads on cloud and Kubernetes platforms.The minimum and maximum base salaries for this role are posted below; additionally, the role is eligible for bonus potential, equity and benefits. The range displayed reflects the minimum and maximum target for new hire salaries for the role based on U.S. location. Within the range, the salary offered will be determined by work location and additional factors, including job-related skills, experience, and relevant education or training.US Pay Range$218,300—$327,500 USDJoin Us in Securing and Accelerating the World's AI TransformationRubrik (RBRK), the Security and AI Operations Company, leads at the intersection of data protection, cyber resilience, and enterprise AI acceleration. Rubrik Security Cloud delivers complete cyber resilience by securing, monitoring, and recovering data, identities, and workloads across clouds. Rubrik Agent Cloud accelerates trusted AI agent deployments at scale by monitoring and auditing agentic actions, enforcing real-time guardrails, fine-tuning for accuracy and undoing agentic mistakes. Linkedin | X (formerly Twitter) | Instagram | Rubrik.comInclusion @ RubrikAt Rubrik, we are dedicated to fostering a culture where people from all backgrounds are valued, feel they belong, and believe they can succeed. Our commitment to inclusion is at the heart of our mission to secure the world’s data.Our goal is to hire and promote the best talent, regardless of background. We continually review our hiring practices to ensure fairness and strive to create an environment where every employee has equal access to opportunities for growth and excellence. We believe in empowering everyone to bring their authentic selves to work and achieve their fullest potential.Our inclusion strategy focuses on three core areas of our business and culture:Our Company: We are committed to building a merit-based organization that offers equal access to growth and success for all employees globally. Your potential is limitless here.Our Culture: We strive to create an inclusive atmosphere where individuals from all backgrounds feel a strong sense of belonging, can thrive, and do their best work. Your contributions help us innovate and break boundaries.Our Communities: We are dedicated to expanding our engagement with the communities we operate in, creating opportunities for underrepresented talent and driving greater innovation for our clients. Your impact extends beyond Rubrik, contributing to safer and stronger communities.Equal Opportunity Employer/Veterans/DisabledRubrik is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, or protected veteran status and will not be discriminated against on the basis of disability.Rubrik provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability or genetics. In addition to federal law requirements, Rubrik complies with applicable state and local laws governing nondiscrimination in employment in every location in which the company has facilities. This policy applies to all terms and conditions of employment, including recruiting, hiring, placement, promotion, termination, layoff, recall, transfer, leaves of absence, compensation and training. Federal law requires employers to provide reasonable accommodation to qualified individuals with disabilities. Please contact us at View email address on us.fitly.work if you require a reasonable accommodation to apply for a job or to perform your job. Examples of reasonable accommodation include making a change to the application process or work procedures, providing documents in an alternate format, using a sign language interpreter, or using specialized equipment.EEO IS THE LAWNOTIFICATION OF EMPLOYEE RIGHTS UNDER FEDERAL LABOR LAWS

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Staff Software Engineer - Reliability (US Citizen Only) in Palo Alto, CA vacancy
  • $158k - $237k

    About The TeamThe Rubrik Engineering team is comprised of people who produce...  ...driven to build efficient, reliable, and cost effective products....  ...at Rubrik are systems/software engineers who ensure that Rubrik...  ...relevant education or training.US Pay Range$158,000—$237,000... 
    Resident
    Local area

    Rubrik

    Palo Alto, CA
    7 days ago
  • $156k - $255k

     ...everyone can succeed. Join us to transform the way...  ...the core of LinkedIn's Reliability Infrastructure...  ...as "always on", every engineer to benefit from a more...  ...advanced degree (MS/PhD) for Staff-level roles. ~6+ years...  ...experience in software development, distributed... 
    Suggested
    For contractors
    Work at office
    Flexible hours

    LinkedIn

    Mountain View, CA
    13 hours ago
  •  ...Job Description Federal US Citizens ONLY - NO SUBS, NO Subcontracting...  ..., NO outside firms. AI Engineer – AI Agents & Generative AI...  ...candidate combines strong Python/software engineering skills with...  ...experimentation and into usable, reliable applications. The... 
    Resident

    United Global Technologies

    Menlo Park, CA
    23 days ago
  • $251k - $310k

     ...+ U.S. states. The Planner/Perception Reliability team’s goal is to build out architectures...  ...identify, and guide fixes of reliability and software integrity issues.  We focus the...  ...range for this full-time position across US locations is listed below. Actual starting... 
    Suggested
    Full time
    Remote work

    Waymo

    Mountain View, CA
    13 hours ago
  • $262k - $364k

     ...decision-maker and escalation point for all reliability, production health, and application...  ...; drive consensus across disparate software engineering and platform teams, acting as a neutral...  ...experience, and relevant education or training. US: $262000 - $364000 (USD) + 25% bonus... 
    Suggested

    Google

    Sunnyvale, CA
    7 days ago
  • $185k - $275k

     ...workloads run seamlessly, reliably, and efficiently...  ...What You'll Do As a Staff Engineer, you will be a technical...  ...You Are ~8+ years of software engineering experience...  ...from you, too. Come join us!  The base salary...  ...defined as a (i) U.S. citizen or national, (ii) U.S.... 
    Resident
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    18 days ago
  •  ...Job Description About Us Array Labs is building a...  ...unprecedented scale, speed, and reliability. Our technology is built to...  ...About the Job As a Staff Software Engineer for Computer Vision, you will...  ...(ITAR) you must be a U.S. citizen, lawful permanent resident... 
    Resident
    Permanent employment
    Full time
    Remote work
    Night shift

    Array Labs

    Redwood City, CA
    24 days ago
  • $150k - $226k

     ...Harness is the AI Software Delivery Platform company, led by technologist...  ..., application security, reliability, compliance, and cost...  ...for exceptional talent to help us move even faster. Position...  ...-time risk detection to help engineering teams ensure software integrity... 
    Local area
    Immediate start
    Flexible hours
    Shift work

    Harness

    Mountain View, CA
    1 day ago
  • $230k - $275k

     ...technology, AI, and analytics to help us scale fast. What began with...  ...is a deeply technical growth engineer who combines hands-on...  ...optimization engine, building a reliable creative generation pipeline,...  ...building and shipping high-impact software systems. ~ Proven experience... 
    Local area

    Quince

    Palo Alto, CA
    3 days ago
  • $143.8k - $230k

     ...operations and management engineering team in VCF Business...  ...cost. In our team, software engineers are focused...  ...completing? As a Staff Software Engineer, you...  ...management to ensure reliability, observability, and deterministic...  ...to work in the US Compensation and... 
    Full time
    Local area

    Broadcom Corporation

    Palo Alto, CA
    2 days ago
  • $279k - $341k

     ...Senior Staff Software Engineer Mountain View, US About EarnIn As one of the first pioneers of earned wage access, our passion at EarnIn is building...  ...for security, privacy, correctness, and operational reliability. The base salary range for this full-time position is... 
    Full time
    2 days per week

    Earnin

    Mountain View, CA
    3 days ago
  • $228k - $285k

     ...generations. Role Summary As a Software Engineer specializing in safety-critical self-...  ...successful implementation of robust and reliable self-driving solutions. This role is...  ...and application process, please email us at ****@*****.***.... 
    Full time
    Contract work
    Local area

    Rivian

    Palo Alto, CA
    13 hours ago
  • $209.7k - $283.8k

     ...ownership to ensure our ML pipelines remain reliable, scalable, and architecturally sound. We are seeking a staff ML engineer to design and evolve the large-scale offline...  ...law. Our differences are strengths that enable us to support the growing and evolving needs of... 
    Work at office
    Worldwide
    Relocation package

    Unity

    Mountain View, CA
    4 days ago
  • $165k - $280k

     ...enabling human life on Mars. SR. SITE RELIABILITY ENGINEER (STARLINK) At SpaceX we’re...  ...connect within minutes of unboxing, and the software that brings it all together. We’ve only...  ...regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful,... 
    Resident
    Permanent employment
    Temporary work
    Worldwide
    Weekend work

    InvestedintheMission

    Palo Alto, CA
    4 days ago
  •  ...About the Role We are looking for a Software Engineer to commercialize the MVP and expand the...  ...infrastructure to ensure security and reliability Assist our efforts to recruit and onboard...  ...happy while they change the world with us. Work remote, or work with us in... 
    Remote work

    Turnblock.io

    Mountain View, CA
    26 days ago
  • $144k - $216k

     ...Staff Software Engineer Onto Innovation is a leader in process control, combining global scale...  ...yield, device performance, quality, and reliability issues. Onto Innovation strives to...  ...necessary for applicants who are not U.S. Citizens, Permanent Residents, or other... 
    Resident
    Permanent employment

    Onto

    Milpitas, CA
    1 day ago
  •  ...It all started when engineer Fred Luddy wrote code that automated...  ...just getting started. Join us to put AI to work for people....  ...on — bias toward generality, reliability, clean abstractions Drive measurable...  ...What We Look For at Senior Staff Level You're a multiplier.... 
    Full time
    Work at office
    Immediate start
    Remote work
    Flexible hours

    ServiceNow

    Mountain View, CA
    3 days ago
  •  ...with Moveworks’ Reasoning Engine and natural language...  ...world-class talent to help us extend agentic AI to...  ...RoleAre you a software engineer who has honed...  ...role for you.As a senior staff software engineer on our...  ...entrusted to agents to perform reliably at scale. You will have... 
    Work at office
    Immediate start
    Remote work
    Flexible hours

    Moveworks

    Mountain View, CA
    7 days ago
  • $152k

     ...reputation for being a dominant and reliable force in South Korean...  ...public company. This fuels us to continue our growth and launch...  ...The Resource Fabric Engineering team builds and operates the...  ...scale. We are looking for a Staff Software Engineer to lead the design... 
    Temporary work
    Flexible hours

    Coupang

    Mountain View, CA
    a month ago
  • $175k - $220k

     ...team member drives our success. Joining us means more than filling a role—you’re...  ...workplace.We are looking for an exceptional Staff Software Engineer to join our Growth Engineering team — a...  ...— turning LLM-powered features into reliable, high-quality product surfaces that... 
    Work experience placement
    Local area
    Flexible hours
    Shift work

    Pebl

    Palo Alto, CA
    7 days ago
  • $207k - $300k

     ...consumption, and rock-solid reliability under peak global traffic.Minimum...  ....8 years of experience in software development in mobile systems...  ...high-performance application engineering.5 years of experience testing...  ...relevant education or training. US: $207000 - $300000 (USD) + 20... 
    Worldwide

    Google

    Mountain View, CA
    14 days ago
  • $180k - $205k

     ...develops the algorithms and software needed to make these systems...  ...path to make it real. Come join us. Job Summary:The OS Core Team...  ...works closely alongside engineers and scientists in the electronics...  ...record of developing and shipping reliable, performant software in a... 
    Full time
    Shift work

    PsiQuantum

    Palo Alto, CA
    a month ago
  • $210k - $240k

     ...creativity and AI with real impact. Join us to help shape the future of enterprise marketing. What You’ll Do As a Staff Software Engineer, you will be responsible for building features...  ...of experience in developing scalable, reliable, performant, and secure full-stack... 
    Work at office
    3 days per week

    Typeface

    Palo Alto, CA
    14 days ago
  • $207k - $300k

    Drive the long-term engineering strategy, incorporating organizational...  ...experience.8 years of experience in software development.7 years of...  ...systems scalable, safer, more reliable, and highly robust.Google Ads...  ...education or training. US: $207000 - $300000 (USD) + 20... 
    Temporary work
    Shift work

    Google

    Mountain View, CA
    11 days ago
  • $207k - $300k

     ...experience.8 years of experience in software development.5 years of...  ...qualifications:Master’s degree or PhD in Engineering, Computer Science, or a...  .../O performance, latency, and reliability for Pixel devices....  ...relevant education or training. US: $207000 - $300000 (USD) + 20... 

    Google

    Mountain View, CA
    14 days ago
  • $207k - $300k

     ...more visible, understandable, reliable, affordable, abundant, and...  ...together experts in energy, AI, software engineering, and products to build tools...  ...About the role As a staff backend software engineer...  ...infrastructure complexity. Come talk to us if you are excited about... 
    Full time
    Flexible hours

    Tapestry

    Mountain View, CA
    3 days ago
  • $207k - $300k

     ...in Computer Science, Computer Engineering, Electrical Engineering,...  ...experience.8 years of experience in software development.Experience in...  ...Experience with observability and reliability for large distributed systems...  ...education or training. US: $207000 - $300000 (USD) + 20... 

    Google

    Mountain View, CA
    14 days ago
  • $229k - $343k

     ...digital services.We’re looking for a Staff Software Engineer to join Snap Inc on our Feature Store...  ...they meet correctness, performance, and reliability requirements. A central part of the role...  ..., please don’t be shy and provide us some information."Default Together" Policy... 
    Full time
    Live in
    Work at office
    Local area

    Snap

    Palo Alto, CA
    a month ago
  • $180k - $211k

     ...Join us in building the future of finance.Our mission is to...  ...customer experience running reliably at scale. We design and operate the systems that give engineers deep visibility into how Robinhood...  ..., not just a tool.As a Staff Software Engineer on the Observability... 
    Work at office
    Immediate start
    Flexible hours
    Shift work
    3 days per week

    Robinhood Financial

    Menlo Park, CA
    a month ago
  • $150.2k - $283.5k

     ...excellence day in and day out. Join us to make positive change by helping build...  ...dreams.In this position... As a Staff Camera Software Engineer, you will take ownership of a Linux camera...  .../CSI bring-up, and production-grade reliability fixes. This is a hands-on development... 
    Immediate start
    Visa sponsorship
    Flexible hours

    Ford

    Palo Alto, CA
    a month ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Software Engineer - Reliability (US Citizen Only). Be the first to apply!