Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

$92.7k - $203.94k

Hispanic Alliance for Career Enhancement

We're building a world of health around every individual - shaping a more connected, convenient and compassionate health experience. At CVS Health®, you'll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger - helping to simplify health care one person, one family and one community at a time. Requisition Job Description Position Summary: About the Team Our Site Reliability Engineering team is the execution engine behind the reliability, availability, and performance of distributed store technology powering thousands of retail and pharmacy locations nationwide. We operate across pharmacy platforms, Point of Sale (POS) systems, handheld devices, store servers, dispensing systems, and edge computing infrastructure - spanning hybrid cloud and on-premises environments at massive fleet scale. Our engineering philosophy is grounded in five pillars: Detection, Prevention, Recovery, Learning Loops, and Developer Experience (DevX). Our operating principle is the reliability covenant: our success is not measured by incident response volume - it is measured by the reliability capability we transfer to the engineering teams we serve. Your success in this role is measured by what the engineering teams in your domain can do independently after working with you, not by how indispensable you become to them. An SSE who has enabled a development team to detect, respond to, and learn from production failures without SRE involvement has delivered the highest-value outcome this role can produce. About the Role As a Senior Software Engineer - SRE, you independently own the reliability posture of an assigned engineering domain. You are not waiting for direction - you are setting it for your domain. You design the alerting strategy, own the SLO health, lead incident command for production issues, facilitate postmortems, tune anomaly detection models, and partner directly with engineering domain owners to shift reliability left into design. You are a technical mentor to SE-level engineers and an escalation resource during active incidents. You have the technical depth to diagnose complex distributed system failures, the data instincts to distinguish genuine anomalies from noise in ML-generated signals, and the organizational skills to drive reliability practice adoption in teams that did not necessarily ask for SRE involvement. Scope: Domain ownership - you operate independently and influence adjacent engineering teams. The Environment You Are Joining This role exists inside an active SRE transformation. The domain-based SRE ownership model you will operate within is in its early stages. Some of the toolchains you will work with are being built in parallel with the operational work. Engineering domain owners are simultaneously learning what SRE can offer them. This is an honest description of the role, not a caveat. Success in this role requires patience alongside technical rigor: you will demonstrate value before you demand process change, build credibility before you expect adoption, and earn trust with engineering domain owners through partnership rather than mandate. The operating environment includes an edge computing fleet deployed directly inside store locations - unattended nodes where deployment blast radius is geographic and fleet-wide, not functional and service-scoped. You will develop a fleet operations mindset: the primary failure mode in this environment is deployment and configuration propagation, not service logic. What You Will Do Detection & Observability Own SLI/SLO health for your assigned domain end-to-end; monitor error budget burn rates and drive proactive burn-down actions before incidents reach end users or pharmacy patients Design multi-signal alerting strategies that go well beyond threshold alerts: burn rate alerting, composite health signals, and anomaly-based detection using time-series models - and validate their output against production ground truth Tune and maintain ML-based anomaly detection models in production: adjust sensitivity thresholds, evaluate false positive rates against alert fatigue metrics, and iterate on model configurations based on incident data Build Critical User Journey (CUJ)-anchored dashboards that surface end-to-end business flow health - not just individual service metrics Drive alert signal-to-noise improvement across the domain; own suppression policies during deployments and maintenance windows to protect on-call quality Identify and instrument unmonitored external dependencies - third-party APIs, downstream services, data providers - using tools like Prometheus Blackbox Exporter, OpenTelemetry, or custom health probes Contribute to the observability platform design: understand the hot/warm/cold tier architecture (Apache Kafka → ClickHouse → cold storage) and contribute domain-specific data models and SLI schemas Prevention & Reliability Engineering Lead Production Readiness Reviews (PRR) for services in your assigned domain; own the readiness gate sign-off and be accountable for what makes it into production on your watch Design and execute fault injection experiments at service level using tools such as LitmusChaos, Chaos Toolkit, or Gremlin; validate blast radius assumptions before rollouts reach production Partner with engineering domain owners on reliability requirements during architecture design and sprint planning - reliability is designed in, not bolted on Own dependency risk mapping for your domain: catalog third-party API failure modes, shared infrastructure failure paths, and chain-wide blast radius scenarios Apply a fleet operations mindset to change management: for any configuration or deployment change touching the edge fleet, assess deployment blast radius by node cohort, review the rollback procedure, and contribute to the go/no-go decision on high-risk changes Understand progressive rollout strategies - canary cohorts, staged fleet expansion, blast radius budgets - and apply them in high-risk deployment reviews Incident Response & Recovery Serve as the primary on-call Technical Incident Commander (IC) for domain incidents; drive structured bridge calls from detection to resolution using established incident command frameworks Author and maintain P0/P1 runbooks with validated, step-by-step remediation procedures; own a quarterly review and dry-run testing cadence to ensure runbooks are accurate when they are actually needed Lead post-incident reviews using structured root cause analysis: causal chain documentation, origin layer classification, and contributing factor identification - not just a timeline of what happened Serve as the real-time escalation point and technical decision support for SE engineers during active incidents Pursue Technical Incident Commander (TIC) certification; qualify as a cross-domain IC candidate available to lead major incidents beyond your assigned domain Learning Loops & Continuous Improvement Learning Loops at this level is systems engineering for organizational memory - not postmortem administration. Facilitate domain-level postmortems with rigor and structure: timestamped timelines, contributing factor taxonomy (origin layer + failure pattern classification), systemic findings, and action items with owners, due dates, and measurable definitions of done Measure learning velocity: track not just whether action items close, but whether the incident class frequency decreases after a fix is applied. Report the ratio of incident classes that recurred versus incident classes that were resolved systemically. This is your primary learning loop quality metric Identify recurring incident patterns across your domain and drive architectural or process changes that eliminate the root cause - not suppress the symptom Close the prevention loop: translate postmortem findings directly into PRR requirement updates, new monitoring coverage, runbook improvements, or development-team practices that prevent recurrence. A postmortem that doesn't change something upstream hasn't finished Mentor SE-level engineers on SRE craft - not just on what to do, but on why the learning loop exists: connect the postmortem to the PRR, the PRR to the design review, the design review to the runbook, and the runbook back to the next incident Developer Experience & Automation (DevX) Eliminate toil systematically: identify and automate manual operational work in your domain using Python, Go, or Bash. Track toil volume eliminated per quarter - time recovered, manual touchpoints removed, and error modes eliminated through automation Build self-service observability tooling that enables development teams to own their own service health dashboards and SLO status pages - reducing SRE as a dependency for basic operational visibility Track and report DORA metrics - Deployment Frequency, Lead Time for Change, MTTR, and Change Failure Rate - for your assigned domain on a quarterly basis; use the data to target reliability investments where they will have the most impact Partner with development teams to shift reliability practices left: reliability checklists in design reviews, SLI instrumentation in definition-of-done, runbook templates that developers can author and own Contribute to the team's shared SRE tooling library; write reusable, well-documented automation modules for common operational patterns Organizational Influence Drive reliability practice adoption without mandate: you will work with engineering teams that have operated independently for years. Adoption of SLO culture, PRR gates, and chaos engineering practices depends entirely on your ability to build credibility through demonstrated value, not through authority you don't have Establish working relationships with engineering domain owners; understand their delivery pressures, their quality concerns, and what they need from SRE to say yes to new reliability requirements Identify the smallest, highest-value reliability practice that a skeptical engineering team will adopt first - land that win, document the outcome, and use it to earn the trust needed for the next one When reliability practices are not being adopted, diagnose the real reason: unclear value, unclear ownership, too much friction, or wrong timing - and adapt the approach rather than escalating to mandate Required Qualifications 5+ years of experience in SRE, DevOps, platform engineering, or related production-systems roles 3+ years operating cloud-native distributed systems at production scale with active on-call responsibility Demonstrated experience as an on-call Incident Commander (IC) for P1 or P2 incidents - structured bridge leadership, not just participant involvement Experience tuning and validating time-series anomaly detection models in a production observability context - this is a Required qualification, not a preferred one; anomaly-based detection is a core function of this role Strong programming proficiency in at least one of Python, Go, or Java at production quality - capable of writing operational tooling that other engineers will rely on Hands-on experience designing SLIs, SLOs, and managing error budgets for customer-facing or business-critical services Deep observability platform experience: Prometheus, Grafana, OpenTelemetry, and at least one log aggregation solution (Loki, Splunk, Elasticsearch) Organizational influence without authority: demonstrated track record of driving SRE practice adoption in engineering teams that did not initially request SRE involvement - this is a Required qualification; technical depth alone is insufficient at this level in a transformation environment Fleet-scale deployment awareness: familiarity with progressive rollout strategies, blast radius management, and configuration drift as a reliability risk in large unattended node deployments Strong cloud platform expertise: AWS, Microsoft Azure, or Google Cloud Platform (GCP) Advanced Kubernetes operational experience: debugging, resource management, networking policies, and workload failure modes. Experience with AI-assisted tooling and development. Preferred Qualifications Experience owning Production Readiness Reviews or service launch gates Hands-on chaos or fault injection experience using LitmusChaos, Chaos Toolkit, or Gremlin TIC (Technical Incident Commander) certification or equivalent structured incident command training Experience operating distributed systems in retail, pharmacy, healthcare, or other operationally sensitive environments where failures have direct patient or customer impact LLM integration for operational use cases (alert summarization, runbook suggestion, incident triage assistance) - design or implementation experience Experience with streaming data platforms: Apache Kafka, Redpanda, Apache Flink, or ksqlDB Familiarity with analytical databases for observability workloads: ClickHouse, Apache Druid, or TimescaleDB Experience with service mesh and traffic management: Istio, Envoy, Linkerd Infrastructure-as-code proficiency at production scale: Terraform, Pulumi, or Ansible What Success Looks Like at 6 Months You own the SLO health and alerting strategy for your assigned domain with full independence - including the ML-based anomaly detection configuration - and the false positive rate is measurably lower than when you joined You have led at least five P1/P2 incident bridges as Incident Commander, with structured postmortems published, action items closed, and at least two incident classes eliminated through root cause remediation rather than symptom suppression You have facilitated your first domain-level PRR and signed off on a service launch At least two engineering teams in your domain have adopted a reliability practice - an SLO health check, a runbook standard, or a deployment gate - that you introduced without top-down mandate You have eliminated measurable toil in your domain: at least three recurring manual tasks automated and documented with before/after comparisons The engineering domain owner you partner with describes SRE as a force multiplier for their team's delivery velocity - not as a gating function that slows them down Education: Bachelor's degree in Computer Science, Engineering, or a related field - or equivalent practical experience Anticipated Weekly Hours 40 Time Type Full time Pay Range The typical pay range for this role is: $92,700.00 - $203,940.00 This pay range represents the base hourly rate or base annual full-time salary for all positions in the job grade within which this position falls. The actual base salary offer will depend on a variety of factors including experience, education, geography and other relevant factors. This position is eligible for a CVS Health bonus, commission or short-term incentive program in addition to the base pay range listed above. Great benefits for great people We take pride in offering a comprehensive and competitive mix of pay and benefits that reflects our commitment to our colleagues and their families. This full‑time position is eligible for a comprehensive benefits package designed to support the physical, emotional, and financial well‑being of colleagues and their families. The benefits for this position include medical, dental, and vision coverage, paid time off, retirement savings options, wellness programs, and other resources, based on eligibility. Additional details about available benefits are provided during the application process and on Benefits Moments. We anticipate the application window for this opening will close on: 10/31/2026 Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state and local laws. #J-18808-Ljbffr

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Richardson, TX vacancy
  •  ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the IP, you will solve complex and broad business problems with simple and straightforward solutions... 
    Suggested

    JP Morgan Chase

    Plano, TX
    4 days ago
  • $96.8k - $145.2k

     ...If you want to be part of an inclusive, adaptable, and forward-thinking organization, apply now.We are currently seeking a Site Reliability Engineer (Onsite Hybrid) to join our team in Plano, Texas (US-TX), United States (US).Job Responsibilities Include: Own and manage... 
    Suggested
    Full time
    Temporary work
    Work at office
    Remote work
    Flexible hours

    NTT DATA

    Plano, TX
    4 days ago
  • $117k - $209.33k

    Job Requisition ID #26WD99276Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team supporting... 
    Suggested
    Full time
    For contractors
    Remote work

    Autodesk

    Plano, TX
    1 day ago
  • $152.6k - $191.5k

     ...is responsible for partnering with leaders across engineering and technology to define objective reliability goals for services. Key responsibilities include composing...  ...improvement.Position Summary:The Senior GCP Site Reliability Engineer acts as an advanced senior individual... 
    Suggested
    Full time
    Work at office
    Day shift

    Bank of America

    Plano, TX
    3 days ago
  •  ...: Meghana GorusuCompany: SRI Tech SolutionsJob Title: Senior Site Reliability EngineerLocation: Plano , TX (remote)Years of Experience: 8 to...  ...are seeking a highly skilled Senior Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate will have... 
    Suggested
    Remote work

    SRI Tech

    Plano, TX
    13 hours ago
  • $128.6k - $184.9k

     ...global cloud platform. As a team of six engineers distributed across the US, Canada, and the...  ...with a strong focus on automation, reliability, and operational excellence. We are one...  ...Qualifications7+ years of experience in Site Reliability Engineering, DevOps, Infrastructure... 
    Permanent employment
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours

    CISCO Systems

    Richardson, TX
    4 days ago
  •  ...we are dedicated to connecting talented professionals with your ideal opportunities. We are currently seeking a qualified Site Reliability Engineer (AI & Agentic Systems) to join our client’s organization and contribute to their ongoing success. Job summaryThis role demands... 

    OpenArc

    Plano, TX
    2 days ago
  •  ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.As a Lead Site Reliability Engineer at JPMorgan Chase within the Infrastructure Platforms, Web Hosting team , you hold a leadership role in your... 
    Work experience placement

    JP Morgan Chase

    Plano, TX
    2 days ago
  • Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally gifted professionals and position yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Corporate and Investment... 

    JP Morgan Chase

    Plano, TX
    1 day ago
  •  ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.As a Lead Site Reliability Engineering at JPMorgan Chase within the Chief Technology Office, Identity & Access Management team, you are the non-functional... 
    Work at office

    JP Morgan Chase

    Plano, TX
    4 days ago
  • $104.9k - $174.7k

     ...Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory... 
    Full time
    Work at office
    Local area
    Remote work
    Work from home

    RELX Group

    Allen, TX
    4 days ago
  • Qualifications: 8+ years of Software Engineering experience, or equivalent demonstrated through...  ...implement and maintain scalable and reliable infrastructure on Google Cloud Platform...  ...vendor resources Willingness to work on-site at stated location in the job openingDepartment... 
    Contract work
    For contractors
    Work experience placement

    Cedent Consulting

    Dallas, TX
    13 hours ago
  •  ...Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk.Design and implement processes that...  ...and reliability tooling.Actively participate in reliability engineering and resilience communities of practice, contributing to... 
    Full time

    Vanguard

    Dallas, TX
    3 days ago
  • $138.4k - $173k

     ...infrastructure as well as help improve the reliability, quality of services and overall...  ...recovery. You’ll collaborate or embed with engineering teams, helping them to improve the reliability...  ...about our locations by visiting our site.Compensation & BenefitsThe base salary that... 
    Full time
    Flexible hours

    AppFolio

    Dallas, TX
    13 hours ago
  •  ...Administrator / SRE in Dallas to own production Java environments, middleware, and cloud automation. You will optimize performance, drive reliability, and mentor teammates while aligning with enterprise security and AI-enabled integrations. You will work across Java apps, IBM... 

    Motion Recruitment Partners LLC

    Dallas, TX
    13 hours ago
  •  ...Experience should include full product experience (APM, Logs, setting up monitoring, alerts, dashboards). Overall looking for a good Reliability Engineer that will support our environments by setting up alerting, monitoring strict SLA's and engaging to determine issues. Must... 

    InterSources

    Dallas, TX
    1 day ago
  •  ...Site Reliability Engineer- W2 Role* Technical proficiency: Strong Proficiency in Java, Strong understanding of Database concepts (Oracle, SQL, Dynamo DB etc.) Industry standard SRE Tools like Prometheus, Grafana, Data Dog Etc Good to have skills: Cloud Concepts / AWS,... 

    RSA Tech Group

    Dallas, TX
    13 hours ago
  •  ...Senior Site Reliability Engineer (Permanent Role) Cleveland, OH, Pittsburgh, PA, or Dallas, TX Your future duties and responsibilities . Monitoring distribution systems and notifying them of any potential issues. . Assisting with troubleshooting on call.... 
    Permanent employment
    Temporary work
    Local area
    Flexible hours
    Shift work
    Weekend work

    System One

    Dallas, TX
    12 days ago
  • Site Reliability Engineer - Vice PresidentSite Reliability Engineering (SRE) is an engineering discipline that combines software and systems engineering to build and run scalable, massively distributed, fault-tolerant systems. At Goldman Sachs, SRE is responsible for improving... 

    Goldman Sachs

    Dallas, TX
    2 days ago
  •  ...and continuously improving the platforms that power TI's digital integration, automation and DevOps capabilities. As an IT Site Reliability Engineer within the Enterprise Platforms team, you will serve as the primary technical platform owner for TI's Apigee Edge private... 
    Local area

    Texas Instruments

    Dallas, TX
    13 hours ago
  • Compliance EngineeringWe are Compliance Engineering, a global team of more than 500 engineers and scientists who work on the most complex...  ...systems by pushing for changes that improve capacity and reliability.Practicing sustainable incident management in a blameless postmortem... 

    Goldman Sachs

    Dallas, TX
    1 day ago
  •  ...tasks using scripting and tools Python Bash etc Collaborate with development infrastructure and support teams to improve system reliability Drive adoption of SRE practices like SLIs SLOs and error budgets Ensure performance optimisation capacity planning and... 
    Permanent employment
    Temporary work
    Work experience placement
    Plano, TX
    3 hours ago
  • $113.1k - $232.3k

    Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied AI Site Reliability Engineer II, you will actively engage in your engineering craft, taking a hands-on approach to the reliability, performance, and operational integrity... 
    Work at office
    Local area
    Visa sponsorship
    Flexible hours
    3 days per week

    Deloitte

    Dallas, TX
    13 hours ago
  •  ...technologies and assists Technical Analysts and Infrastructure Engineers to ensure that technology solutions comply with enterprise system...  ...manual and repeatable work or inefficient processesConducts on-site evaluations of third-party products being considered for firm... 
    Full time
    Work at office
    Day shift

    Bank of America

    Plano, TX
    4 days ago
  •  ...product and program backlogDefines norms and agreements for the Agile Release Train and enforces the agreementsAs the Release Train Engineer (RTE) for Technology Infrastructure, you are accountable for how an ~18-team train prepares for and executes the Planning Interval... 
    Full time
    Work at office
    Immediate start
    Day shift

    Bank of America

    Plano, TX
    3 days ago
  • $350 per month

     ...donation to a charity of your choiceWhat to ExpectThe Sr. Release Engineer & Administration Manager (Salesforce) is responsible for owning...  ...serves as a technical expert and execution leader, ensuring reliable deployments, platform stability, and strong governance across a... 
    Full time
    Seasonal work
    Work at office
    Immediate start

    Hyundai Capital America

    Plano, TX
    13 hours ago
  •  ...Pay Rate: $40/Hr. W2 Experience: 3-5 Years Overview We are seeking a remote Junior SRE/DevOps Engineer role. The ideal candidate has foundational knowledge of Site Reliability Engineering (SRE) and Kubernetes, and is enthusiastic about growing in a DevOps‑driven... 
    Long term contract
    Contract work
    Internship
    Remote work

    BayOne Solutions

    Richardson, TX
    13 hours ago
  • $86.8k - $165.2k

     ...the strength of more than 100 years of experience and renowned engineering expertise to meet the needs of today’s mission and stay ahead of...  ...locations, regardless of whether the role is designated as on-site, hybrid or remote.The salary range for this role is 86,800 USD... 
    Temporary work
    Work experience placement
    Work at office
    Local area
    Remote work
    Relocation
    Flexible hours

    Raytheon

    Plano, TX
    13 hours ago
  • CVS Health is hiring a Senior Software Engineer - SRE to own the reliability posture of an engineering domain. You design alerting, own SLOs, lead incident response, and mentor engineers. You will work with edge computing and hybrid cloud to improve system resilience across... 

    Hispanic Alliance for Career Enhancement

    Richardson, TX
    2 days ago
  • $40 per hour

    A technology solutions provider is seeking a remote Junior SRE/DevOps Engineer. The ideal candidate should have foundational knowledge of Site Reliability Engineering (SRE) and Kubernetes. Responsibilities include gaining experience in a DevOps-driven environment. Applicants... 
    Remote job
    Long term contract
    Internship

    BayOne Solutions

    Richardson, TX
    13 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!