Senior Site Reliability Engineer
$170k - $219kRadar
Site Reliability Engineer
RADAR runs data infrastructure across 1,600+ live retail stores, processing tens of billions of real-world events every day. We're hiring a Site Reliability Engineer to own the reliability of that system end to end — leading incident response, running day-to-day NOC operations, and building the observability foundation that lets us catch issues before they hit a store floor. You'll be the steady hand during a live incident, and the engineer making sure there are fewer of them to begin with.
Responsibilities:
- Own the incident management lifecycle end to end: detection, triage, escalation, communication, resolution, and postmortem for production incidents.
- Act as Incident Commander for high severity incidents, coordinating across engineering, support, and leadership until resolution.
- Run day-to-day NOC (Network Operations Center) operations, including 24/7 shift coverage, escalation matrices, and shift handover protocols.
- Coach and mentor NOC analysts on triage discipline and escalation judgment, and own NOC KPIs like response time and escalation accuracy.
- Design and maintain observability pipelines across metrics, logs, and traces, and define SLIs/SLOs with engineering and product.
- Build dashboards and alert that surface true signal from our sensor and platform data, cutting down on noise and alert fatigue.
- Facilitate blameless postmortems and root cause analysis, and track corrective actions through to closure.
- Maintain on-call rotations, runbooks, and escalation policies, and report on MTTA/MTTR/MTBF trends to leadership.
About You:
Required:
- You have 5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure, or Production Operations, with direct incident response and on-call experience.
- You have experience running or actively contributing to a NOC, including shift scheduling, escalation processes, and performance metrics.
- You have strong hands-on experience with observability tooling (Prometheus, Grafana, Datadog, New Relic, Splunk, ELK, OpenTelemetry, or similar).
- You have a solid understanding of SLIs, SLOs, SLAs, and error budgets, and how to use them to drive prioritization.
- You have hands-on release engineering experience, including CI/CD pipelines, deployment automation, and safe rollout practices like canary releases, feature flags, and automated rollbacks.
- You are proficient in at least one scripting or programming language (Python, Go, Bash, etc.).
- You have experience with infrastructure-as-code tools (Terraform, Ansible).
- You have experience with cloud platforms (AWS, GCP, or Azure) and container orchestration (Kubernetes, Docker).
- You are a clear, direct communicator who stays calm and organized under pressure during live incidents.
Preferred:
- You have experience building or scaling a NOC from the ground up.
- You have a background in distributed systems architecture and microservices troubleshooting.
- You have familiarity with chaos engineering and resilience testing.
- You have a certification such as AWS Certified SysOps Administrator, Google Professional Cloud DevOps Engineer, or ITIL.
In your first 30 days, you will:
- Learn RADAR's mission, technology stack and core values.
- Complete onboarding and security compliance training.
- Shadow the NOC across shifts and review recent incident history and open postmortem action items.
In your first 60 days, you will:
- Take on-call as primary or secondary responder for at least one service area.
- Tune or consolidate at least one high-volume, low-signal alert source, and audit existing observability coverage for major gaps.
- Draft or update runbooks for the top recurring incident types, and instrument one under-monitored service.
In your first 90 days, you will:
- Lead Incident Commander duties for high severity incidents, including full postmortem facilitation.
- Deliver a reliability report on incident trends, NOC KPIs, and observability maturity gaps.
- Present a roadmap for the next 2–3 quarters covering NOC process, Automation, SLO definitions, and observability investment.
At RADAR, your base pay is one part of your total compensation package. The expected base salary range for this position is $170,000 - $219,000. Individual pay is determined by work location and additional factors, including job-related skills, experience and relevant education or training. You will also be eligible to receive other benefits including: equity, comprehensive medical and dental coverage, life and disability benefits, 401k plan, flexible time off, and paid parental leave. The pay range listed for this position is a good faith and reasonable estimate of the range of possible base compensation at the time of posting.
Research has shown that women & underrepresented minorities are more likely to read lists of requirements and consider themselves unqualified if they don't meet every single one. This list represents what we're ideally looking for, but everyone has unique strengths & weaknesses, and we hire for strength & potential, not lack of weakness.
$190k - $280k
...establishment of our SRE function. As the SRE Lead, you will establish and mature the reliability practices used across our cloud infrastructure and platform services. You will work with Cloud Engineering and product teams to define reliability targets, improve observability, and...SeniorFull timeTemporary workPart timeWorldwide$92.4k - $148.8k
...Senior Site Reliability Engineer San Diego About SHEIN SHEIN is a global online fashion and lifestyle retailer, offering SHEIN branded apparel and products from a global network of vendors, all at affordable prices. Headquartered in Singapore, with more than...SeniorTemporary workWork at officeWorldwideFlexible hours$149.8k - $262.2k
Company DescriptionIt all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried... ...engineers who are tasked with maintaining and developing the reliability, scalability and performance of the ServiceNow infrastructure....SuggestedPermanent employmentWork at officeImmediate startRemote workFlexible hoursShift workNight shift- ...we deploy forward with the people who own the mission. Our engineers, product team, and CTO deploy forward to the point of friction... ..., and counter-narcotics missions. We are looking for an Site Reliability Engineer to own the reliability, health, and deployment automation...SuggestedFull timeRemote workHome office
- ...on one unified cloud. One cloud for compute, inference, and agents. Role Overview We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and...Suggested
- ...space More than 30,000 companies and 100,000 electronics engineers worldwide use Altium We are growing, debt-free, and... ...resources to become #1 in the EDA industry Role Overview: Senior Site Reliability Engineerensures the reliability, availability, and...Worldwide
- ...contributed to the FaceID and FaceKit project in the past and more recently the new LIDAR iPad sensor. We are looking for the right Site Reliability Engineer to help us take our efforts to the next level. In this role, you will help lead our cloud based infrastructure team for...Work experience placement
$52 per hour
...Sr. Software Systems Engineer If you're an experienced Software Systems Engineer who enjoys solving complex manufacturing challenges and building efficient, automated workflows, this is a great opportunity to make an impact. You'll support Manufacturing Execution System...SeniorContract workFlexible hours- ...Semiengineering seeks a Senior Forward Deployment Engineer who works inside a customer’s development program as a software developer, systems engineer, and deployment lead in one role. The role ships production code and drives complex deployments in collaboration with...Senior
- ...Northrop Grumman in Rancho Bernardo, CA is seeking a Principal Engineer RQ-4 Firebee CAM to join our Product Support team. You will lead lifecycle development and logistics activities, ensuring deliverables meet the SOW, budget, and schedule in a high-tempo environment...Senior
- ...Northrop Grumman Aeronautics Systems is seeking a Sr Principal Software Engineer for a hybrid role across multiple California locations, including Rancho Bernardo and El Segundo, with travel. You will translate LiDAR algorithms into high-performance GPU code, develop...Senior
- ...Northrop Grumman Aeronautics Systems is seeking a Sr Principal Software Engineer for a hybrid role with locations in Melbourne, Florida, Rancho Bernardo, CA or El Segundo, CA. The role reports to the R&AD organization Aerial Refueling System Lead and focuses on high-performance...Senior
- ...cloud, or through a hybrid approach. Teradata delivers real business value with AI.What You Will Do We are seeking a Senior Platform Systems Engineer to help architect and evolve next-generation AI and database infrastructure platforms used in enterprise and hyperscale...Senior
- ...Hi, I hope you are doing well! We have an opportunity for Senior Software Engineer - Go Development & Identity Security with one of our clients for San Diego, CA. Please see the job details below and let me know if you would be interested in this role....SeniorLong term contract
- ...IC Resources in the San Diego area is seeking a Sr. Principal DevOps Engineer with a high level of autonomy and impact. The role is hybrid on-site in the San Diego area, focusing on DevOps/SRE leadership across multi-cloud environments. The ideal candidate has 10 years...Senior
$143k - $215k
...consistency.SummaryResmed is seeking a Senior Platform Engineer to help design, build, and operate our... ...capabilities that enable secure, reliable, and scalable software delivery. The ideal... ...of experience in Platform Engineering, Site Reliability Engineering,...SeniorFull timeTemporary workWork experience placementFlexible hours- ...drone operationsDevelopment of new algorithmsModifications and extension of our communications protocolSoftware architecture and engineering process developmentRequired QualificationsMust be a US CitizenStrong development experience with C++Front end web development using...Senior
- ...Senior Software Engineer San Diego, CA Hybrid Sep 16, 2026 Essential Duties and Responsibilities Maintain and extend ASR C++ code base for ground station and drone operations Development of new algorithms Modifications and extension of our communications protocol Software...Senior
- Description SAIC is hiring for a Senior Software Engineer who must be experienced in the full development life cycle using modern methodologies, such as Agile/Scrum and test-driven development. Ideal candidate will be experienced with complex coding in C++, CLI, C#, Java...Senior
$130k - $230k
...operating in complex and contested environments. Forge is the engineering platform used to build, integrate, test, and operate Hivemind... ...develop autonomy software, investigate failures, and deliver reliable capabilities faster. What you'll do:Build and operate a full-stack...SeniorFull timeTemporary workPart timeLocal areaRemote workWorldwide$195k - $235k
...Join our fast-paced and passionate team as a Senior Principal Software Engineer on our Robotic Automation team. As we scale, you will be instrumental... ...on system interfaces, data flow, scalability, and reliability across the full automation stack. Act as the top technical...SeniorFull timeLocal area$106.8k - $133.5k
...make recommendations.Advanced knowledge of plant and equipment reliability.Travel between facilities to support all functional areas.... ...equivalent experience required.Minimum of 7-10+ years of relevant engineering experience including a minimum of 4 years cGMP experience....SeniorFull timeWork at officeFlexible hours- ...Northrop Grumman Mission Systems seeks a Sr. Principal Cyber Software Engineer (Level 2 CNO Analyst/Programmer) to design and develop specialized tools and data flows for secure systems. The role requires working across diverse environments, including Windows and UNIX...Senior
$255k - $300k
.../ San Jose, California( Client Services / Solutions ) - Sales Engineering /Full Time /HybridAHEAD builds platforms for digital business.... ...diversification and enrichment of ideas and perspectives at AHEAD. Senior Client Solutions EngineerThe Senior Client Solutions Engineer...SeniorFull timeWork at office$120k - $230k
...Description:Hivemind is looking for an experienced Full Stack Software Engineer to join our Shared Services Engineering team. The Shared... ..., contribute to architectural decisions, and build secure, reliable systems for external customers and internal engineering teams....SeniorTemporary workPart timeWorldwide$193k - $220k
...Summary:Works independently as a technical project leader, applies engineering principles, procedures and techniques to perform systems... ...links. Expertise with RF and specifically RF used for line of site and beyond line of site communications is essential. Skills should...SeniorFull timeFor subcontractor$130k - $190k
...and warehouse operations. This is a rare senior IC opportunity to help architect and... ...technical standards, and AI-augmented engineering practices from the inside. You will... ...checks may miss. Own observability, reliability, and performance for cloud-hosted services...SeniorFull time- ...Northrop Grumman Corp. in California is seeking a Senior Principal Systems Safety Engineer to support the AISR&T program in Rancho Bernardo, CA.... ...with systems, software, hardware, integration/test, and reliability teams to embed safety considerations into design #J-...Senior
- ...Aeronautics Systems is seeking a Sr. Principal Electronic Warfare Engineer to join our San Diego (Rancho Bernardo) or Palmdale, CA teams. The role focuses on payload subsystem engineering, with on‑site work and up to 25% travel to support test activities. The...Senior
$181.2k - $317.1k
Company DescriptionIt all started when engineer Fred Luddy wrote code that automated a tedious... ....You will be accountable for the reliability, scalability, and operability of the services... ...and trade-offs to both engineers and senior leadership. For positions in this location...SeniorWork at officeImmediate startRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- senior service associate San Diego, CA
- senior safety specialist San Diego, CA
- senior vice president of business development San Diego, CA
- senior service designer San Diego, CA
- senior mulesoft developer San Diego, CA
- senior media manager San Diego, CA
- senior business manager San Diego, CA
- senior linux systems engineer San Diego, CA
- senior mainframe developer San Diego, CA
- senior cloud security engineer San Diego, CA



