Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

$170k - $219k

Radar

Site Reliability Engineer

RADAR runs data infrastructure across 1,600+ live retail stores, processing tens of billions of real-world events every day. We're hiring a Site Reliability Engineer to own the reliability of that system end to end — leading incident response, running day-to-day NOC operations, and building the observability foundation that lets us catch issues before they hit a store floor. You'll be the steady hand during a live incident, and the engineer making sure there are fewer of them to begin with.

Responsibilities:

  • Own the incident management lifecycle end to end: detection, triage, escalation, communication, resolution, and postmortem for production incidents.
  • Act as Incident Commander for high severity incidents, coordinating across engineering, support, and leadership until resolution.
  • Run day-to-day NOC (Network Operations Center) operations, including 24/7 shift coverage, escalation matrices, and shift handover protocols.
  • Coach and mentor NOC analysts on triage discipline and escalation judgment, and own NOC KPIs like response time and escalation accuracy.
  • Design and maintain observability pipelines across metrics, logs, and traces, and define SLIs/SLOs with engineering and product.
  • Build dashboards and alert that surface true signal from our sensor and platform data, cutting down on noise and alert fatigue.
  • Facilitate blameless postmortems and root cause analysis, and track corrective actions through to closure.
  • Maintain on-call rotations, runbooks, and escalation policies, and report on MTTA/MTTR/MTBF trends to leadership.

About You:

Required:

  • You have 5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure, or Production Operations, with direct incident response and on-call experience.
  • You have experience running or actively contributing to a NOC, including shift scheduling, escalation processes, and performance metrics.
  • You have strong hands-on experience with observability tooling (Prometheus, Grafana, Datadog, New Relic, Splunk, ELK, OpenTelemetry, or similar).
  • You have a solid understanding of SLIs, SLOs, SLAs, and error budgets, and how to use them to drive prioritization.
  • You have hands-on release engineering experience, including CI/CD pipelines, deployment automation, and safe rollout practices like canary releases, feature flags, and automated rollbacks.
  • You are proficient in at least one scripting or programming language (Python, Go, Bash, etc.).
  • You have experience with infrastructure-as-code tools (Terraform, Ansible).
  • You have experience with cloud platforms (AWS, GCP, or Azure) and container orchestration (Kubernetes, Docker).
  • You are a clear, direct communicator who stays calm and organized under pressure during live incidents.

Preferred:

  • You have experience building or scaling a NOC from the ground up.
  • You have a background in distributed systems architecture and microservices troubleshooting.
  • You have familiarity with chaos engineering and resilience testing.
  • You have a certification such as AWS Certified SysOps Administrator, Google Professional Cloud DevOps Engineer, or ITIL.

In your first 30 days, you will:

  • Learn RADAR's mission, technology stack and core values.
  • Complete onboarding and security compliance training.
  • Shadow the NOC across shifts and review recent incident history and open postmortem action items.

In your first 60 days, you will:

  • Take on-call as primary or secondary responder for at least one service area.
  • Tune or consolidate at least one high-volume, low-signal alert source, and audit existing observability coverage for major gaps.
  • Draft or update runbooks for the top recurring incident types, and instrument one under-monitored service.

In your first 90 days, you will:

  • Lead Incident Commander duties for high severity incidents, including full postmortem facilitation.
  • Deliver a reliability report on incident trends, NOC KPIs, and observability maturity gaps.
  • Present a roadmap for the next 2–3 quarters covering NOC process, Automation, SLO definitions, and observability investment.

At RADAR, your base pay is one part of your total compensation package. The expected base salary range for this position is $170,000 - $219,000. Individual pay is determined by work location and additional factors, including job-related skills, experience and relevant education or training. You will also be eligible to receive other benefits including: equity, comprehensive medical and dental coverage, life and disability benefits, 401k plan, flexible time off, and paid parental leave. The pay range listed for this position is a good faith and reasonable estimate of the range of possible base compensation at the time of posting.

Research has shown that women & underrepresented minorities are more likely to read lists of requirements and consider themselves unqualified if they don't meet every single one. This list represents what we're ideally looking for, but everyone has unique strengths & weaknesses, and we hire for strength & potential, not lack of weakness.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in Sunnyvale, CA vacancy
  • $174k - $252k

     ...systems by pushing for changes that improve reliability and velocity.Practice sustainable...  ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical...  ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you... 
    Senior

    Google

    Sunnyvale, CA
    3 days ago
  • $168k - $270.25k

     ...phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  • $160k - $240k

     ...millions of times a day - quickly, reliably, and securely. Any time you...  ...at Fiserv.Job TitleSenior Site Reliability EngineerWhat does a successful Site Reliability Engineer do at Fiserv?You will join our...  ...operations or DevOps at a mid-to-senior level.Strong shell scripting... 
    Senior
    Full time

    Fiserv

    Sunnyvale, CA
    2 days ago
  • $148k - $235.75k

     ...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $101k - $161k

     ...several prestigious awards, such as Best Engineering Team, Best Company for Diversity,...  ...DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s...  ...: EngineeringExperience level: Mid-Senior LevelIndustry: Computer Networking
    Senior

    Arista Networks

    Santa Clara, CA
    5 days ago
  • $152k - $241.5k

     ...artificial intelligence.We’re looking for a Senior SRE to join our Compute Farm team and...  ...host lifecycle management, fleet reliability/auto-healing, E2E observability or data-...  ...Python, Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $222k - $300.5k

     ...OverviewAbout the TeamIntuit's Infrastructure and Site Reliability organization owns the operational...  .... The Fintech Platform Systems Engineering team builds and operates the AWS-based...  ...negotiable.The OpportunityWe're hiring a Senior Manager, Site Reliability Engineering to... 
    Senior
    Worldwide
    Shift work

    Intuit

    Mountain View, CA
    5 days ago
  •  ...Job Title : Senior Site Reliability Engineer Location : Santa Clara, CA Contract ENGAGEMENT SUMMARY The Candidate will provide SRE services for AI platforms and supporting infrastructure with emphasis on reliability engineering, incident response... 
    Senior
    Contract work

    VDart

    Santa Clara, CA
    5 days ago
  • $132.6k - $214.5k

     ...you will collaborate closely with our engineering teams to develop innovative solutions that...  ...’ performance and health. As a Senior Staff SRE with the Cortex Observability...  ...operability of the product and ensure the reliability and availability of our services.... 
    Senior
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks

    Santa Clara, CA
    1 day ago
  • $174k - $252k

     ...Site Reliability Engineering (SRE) Job Site Reliability Engineering (SRE) is what you get when you treat operations as if it's a software problem. Our mission is to progress, protect, and provide for the software and systems behind all of Google's public services -... 
    Senior

    CompliancePoint

    Mountain View, CA
    3 days ago
  • $150.4k - $277.6k

     ...Services The Media Platforms SRE team under the Apple Service Engineering division is one of the most exciting examples of Apple’s long...  ...field with 4+ years experience At least 6 years in a Reliability Engineering, DevOps or infrastructure focused role Advanced... 
    Senior
    Relocation
    Day shift

    Apple

    Cupertino, CA
    2 days ago
  •  ...Own the architecture and design of reliable, scalable, cost-effective, and performant AI...  ...master's degree in computer science or engineering is preferred. Key Skills Software...  ...Machine Learning, Artificial Intelligence, Site Reliability Engineering, AI Infrastructure... 
    Senior

    Jobleads-US

    Sunnyvale, CA
    2 days ago
  •  ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and...  ...and networking teams to improve service reliability and deployment workflowsDeploy and...  ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  •  ...Lambda’s designated work from home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our cloud offering...  ...and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    1 day ago
  • $267k - $356k

     ...day is currently Tuesday.Lambda's Storage Engineering team is the backbone behind our world-...  ...workloads in the industry, which means reliability and performance aren't just goals—they're...  ...defined storage across new and existing sites using tools such as Ansible, Jenkins etc... 
    Senior
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    5 days ago
  • $104.9k - $174.7k

     ...SRE role is responsible for improving the reliability, availability, performance, and...  ...actions through completion.Follow up with engineering, development, security, support, and business...  ...Qualifications5+ years of experience in Site Reliability Engineering, Systems Engineering... 
    Senior
    Full time
    Local area

    LexisNexis Risk Solutions Group

    San Jose, CA
    5 days ago
  •  ...keep the world running. Location: 5 on-site days a week in Sunnyvale, CA Headquarters. Our Team's Vision: Our Engineering team is shaping the future of...  ...: We are looking for an experienced Senior Site Reliability Engineer (SRE) with a strong background... 
    Senior
    Work experience placement

    Illumio

    Sunnyvale, CA
    3 days ago
  • $192.4k - $275.8k

     ...the world's most demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service Engineering disciplines...  ...this is the team for you Your ImpactYou will be the most senior technical individual contributor on the team — setting the... 
    Senior
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    3 days ago
  • $175k - $265k

     ...Overviewd-Matrix's SRE team owns the infrastructure layer that every engineering team and customer depends on — colocation facilities, on-...  .... This role is a core member of that team, responsible for reliability, automation, and observability across colo, on-premises lab,... 
    Senior

    d-Matrix

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

     ...infrastructure for AI workloads. We are looking for Software Engineers with SRE or Production Engineering experience who have worked...  ...initial provisioning through repair.Experience managing production reliability through on-call duties, incident response, observability, and... 
    Senior
    Permanent employment
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $248k - $396.75k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing, building, and operating large-scale production systems...  ..., analytics, and automated anomaly detection.Partner with senior leaders and engineers across Cloud, Platform, Security,... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...ServiceNow is seeking a Staff Voice AI Engineer – SRE/DevOps in Santa Clara. You will design and deliver cloud-native solutions for Voice AI, ensuring high availability and observability, and integrate LLMs into real-time voice platforms. You will mentor teammates,... 
    Senior

    Jobleads-US

    Santa Clara, CA
    2 days ago
  • $255.7k - $300k

     ...designs from peers, providing feedback to ensure best practices in reliability, security, and efficiency.Triage and resolve complex system...  ...execution of software development initiatives.Mentor other engineers and contribute to the engineering community through documentation... 
    Full time

    Google

    Sunnyvale, CA
    4 days ago
  • $230k - $250k

     ...network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change...  ...how things have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on"... 
    Night shift

    Forward Networks

    Santa Clara, CA
    2 days ago
  • $170k - $200k

    We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance,... 
    Full time
    Worldwide

    Fortinet

    Sunnyvale, CA
    4 days ago
  • $152k - $241.5k

     ...the world.Join the Simulation Software team at NVIDIA as a Senior System Software Engineer! This role offers an outstanding opportunity to work on...  ...development by enabling Chips Simulation as a trusted and reliable virtual platform.What you will be doing:Drive early... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  •  ...d-Matrix's SRE team owns the infrastructure layer that every engineering team and customer depends on — colocation facilities, on-premises...  .... This role is a core member of that team, responsible for reliability, automation, and observability across colo, on-premises lab,... 
    Senior

    Jobleads-US

    Santa Clara, CA
    5 days ago
  • $272k - $431.25k

    NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning and Process) Cloud Infrastructure Team. IPP is a global organization within NVIDIA. This group works with various other groups within NVIDIA such as Graphics... 
    Full time
    Work experience placement
    Worldwide

    Nvidia

    Santa Clara, CA
    5 days ago
  • $152k - $241.5k

    NVIDIA is seeking an innovative and highly motivated engineer with deep expertise in systems software to join our GPU Software team. In...  ...environmentsCollaborate with globally distributed teams to deliver scalable, reliable, and high-impact GPU software solutionsWhat we need to see: BS... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...complex, distributed, cloud-native systems. As a Staff Platform Engineer, you will play a critical role in ensuring these systems...  ...hands-on engineering and technical leadership role. You will own reliability for major platform domains, design scalable solutions on Kubernetes... 
    Senior

    Saviynt

    Milpitas, CA
    a month ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!