Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Observability Platform Engineer (SRE)

$118.45k - $236.9k
Full-time

CVS Health

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time.

POSITION SUMMARY

CVS Health PBM is looking for hands-on, passionate people who want to join a high energy and growing team, who want to be on the forefront of digital innovation that aims to reinvent what a pharmacy and a health care company can be in the digital world.

As a Lead Platform Reliability Engineer , you will design and implement metrics and observability frameworks with a strong focus on service level objectives (SLOs), service level indicators (SLIs), error budgets, and cloud infrastructure scaling and capacity estimation.

This individual contributor role is critical to enhancing our monitoring and observability capabilities, while also driving automation initiatives related to quality gates within the release engineering process. You will work closely with cross‑functional teams to ensure the reliability, performance, and scalable growth of our cloud‑based systems.

Expectations for the Role:

Metrics Development: Define, implement, and maintain key performance metrics, SLOs, and SLIs to measure system reliability and performance. Ensure alignment with business objectives and operational goals.

Error Budgets: Manage error budgets effectively, collaborating with development teams to balance reliability and feature delivery. Analyze incidents and outages to inform adjustments to error budgets.

Monitoring & Observability: Design and implement comprehensive monitoring solutions to provide real-time visibility into system health. Utilize tools such as Prometheus, Grafana, Loki, Temp and other observability platforms to create dashboards and alerts.

Cloud Infrastructure Scaling: Architect, design, and implement scalable cloud infrastructure capable of supporting multiple business applications, ensuring reliability, performance, and future growth.

Quality Gates Automation: Develop and implement automated quality gates that ensure all releases meet defined reliability and performance standards. Lead the release Devops team to integrate these gates into the CI/CD pipeline.

Incident Management: Assist in incident response efforts by providing insights from metrics and monitoring tools. Conduct post-mortem analyses to identify root causes and recommend preventive measures.

AIOps Insight Automation: Use AI to surface what changed / what’s abnormal / next best action from metrics, logs, and traces—minimizing manual dashboard analysis.

AI‑Accelerated Incident Response: Apply GenAI to speed triage and RCA with fast signal summarization and guided investigation paths.

AI/LLM Observability & Governance: Monitor AI workloads for quality, safety, cost, latency, reliability with end‑to‑end tracing (request → prompt → tools → output) and secure logging/redaction.

AI‑Backed Release Quality Gates: Embed AI signal checks into CI/CD to flag SLO risk, latency/error drift, and regression patterns before production release.

REQUIRED QUALIFICATIONS

  • 10+ years of experience in Software Engineering, Platform Engineering, or SRE.
  • 7+ years of experience with observability practices, including SLIs/SLOs/SLAs, alerting, and incident management.
  • 7+ years building production-grade backend services in Java/python.
  • 7+ years implementing and operating OpenTelemetry, including OTLP, semantic conventions, and instrumentation patterns.
  • 7+ years with cloud-native and containerized platforms (Docker, Kubernetes, Argo CD).
  • 7+ years working with public cloud platforms (AWS, GCP, or Azure).
  • 5+ years designing and scaling distributed, high‑volume data pipelines.
  • 5+ years working with Grafana OSS or comparable observability backends (e.g., Grafana, Loki, Tempo, Prometheus).
  • 5+ years with relational databases (PostgreSQL, MySQL).

PREFERRED QUALIFICATIONS

  • Excellent analytical skills and the ability to communicate complex technical concepts to non-technical stakeholders
  • Experience with service meshes and networking technologies such as Envoy and Istio
  • Experience integrating or operating commercial observability platforms (Splunk, AppDynamics, etc.)
  • Experience with streaming and data platforms such as Kafka, Pulsar, or similar technologies
  • Familiarity with time-series, NoSQL, or analytical databases (ClickHouse, Bigtable, Cassandra, etc.)
  • Experience with Infrastructure as Code tools such as Terraform or CloudFormation
  • Experience with cost optimization and capacity planning for large-scale cloud infra
  • Experience with chaos engineering, resiliency testing, or fault injection
  • Background in security‑aware platform design, including secure service‑to‑service communication
  • Experience mentoring senior engineers and influencing platform standards across organizations
  • Strong operational experience supporting 24x7 production systems, including on‑call responsibilities
  • Knowledge of security best practices in cloud environments

EDUCATION

Bachelor’s degree or equivalent experience (HS diploma + 4 years relevant experience)

Pay Range

The typical pay range for this role is:

$118,450.00 - $236,900.00


This pay range represents the base hourly rate or base annual full-time salary for all positions in the job grade within which this position falls. The actual base salary offer will depend on a variety of factors including experience, education, geography and other relevant factors. This position is eligible for a CVS Health bonus, commission or short-term incentive program in addition to the base pay range listed above. This position also includes an award target in the company’s equity award program.

Our people fuel our future. Our teams reflect the customers, patients, members and communities we serve and we are committed to fostering a workplace where every colleague feels valued and that they belong.

Great benefits for great people

We take pride in offering a comprehensive and competitive mix of pay and benefits that reflects our commitment to our colleagues and their families.

This full‑time position is eligible for a comprehensive benefits package designed to support the physical, emotional, and financial well‑being of colleagues and their families. The benefits for this position include medical, dental, and vision coverage, paid time off, retirement savings options, wellness programs, and other resources, based on eligibility.


Additional details about available benefits are provided during the application process and on Benefits Moments .

We anticipate the application window for this opening will close on: 08/31/2026

Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state and local laws.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Staff Observability Platform Engineer (SRE) in Farmers Branch, TX vacancy
  • $56 - $80 per hour

    AWS SRE/Platform EngineerContract to HireAddison, TX (Hybrid)ROLE OVERVIEW Seeking a highly...  ...Technical Lead - SRE and Platform Engineering to provide technical leadership for the reliability, performance, security, observability, and operational management of enterprise... 
    Suggested

    Yoh

    Addison, TX
    1 day ago
  • Site Reliability Engineer - Vice PresidentSite Reliability Engineering (SRE) is an engineering discipline...  ...firm’s most critical platform services and ensures...  ...initiatives for future growth.Observability & Insights: Define and...  ...develop senior and staff-level engineers.... 
    Suggested

    Goldman Sachs

    Dallas, TX
    2 days ago
  • $95 - $125 per hour

    Platform Engineer (DevOps / CI/CD / Automation / Observability)DETAILSLocation: Remote Position Type: 16M ContractHourly / Salary: to $90W2 based on experience)JOB SUMMARYVaco is currently seeking a Platform Engineer (DevOps / CI/CD / Automation / Observability) for a... 
    Suggested
    Hourly pay
    Contract work
    For contractors
    Work at office
    Local area
    Remote work

    VACO

    Addison, TX
    5 days ago
  • $85 - $90 per hour

     ...Role:  Senior SRE Engineer  Location: Dallas / Fort Worth, Texas Rate: up to $85-$9...  ...Experience with container orchestration platforms such as Kubernetes. ~ Experience...  ...Shell scripting ~ Experience managing observability tools such as Grafana, Kibana and Prometheus... 
    Suggested
    Hourly pay
    Contract work
    Work experience placement

    CorGTA

    Dallas, TX
    more than 2 months ago
  •  ...Senior DevOps Platform Engineer Required Candidate Location: Hybrid/Wilmington, DE...  ...No H1 Must Have: # DevOps and SRE experience - Very Hands on # Dallas...  ...SRE practices, with a focus on CI/CD, observability, automation, and reliability across both... 
    Suggested
    Local area
    Relocation

    3B Staffing LLC

    Dallas, TX
    5 days ago
  •  ...EngineeringWe are Compliance Engineering, a global team of more than...  ...build and operate a suite of platforms and applications that...  ...Compliance application portfolio.SRE at Goldman Sachs combines software...  ...design.Proficiency with Observability stacks, including... 

    Goldman Sachs

    Dallas, TX
    1 day ago
  • $155k - $233k

     ...SummaryWe are seeking a Senior DevSecOps / Platform Engineer to design, build, and operate secure CI...  ...’ll collaborate closely with Security, SRE/Infra Platform, and engineering teams...  ...access boundaries and auditability.Observability & Operational ExcellenceEnhance platform... 
    Full time
    Work at office
    Shift work

    Equinix

    Dallas, TX
    4 days ago
  • $150k - $190.7k

     ...impact. Join us!Job Description:The Senior Engineer SIEM Platform Engineering & Operations is...  ...dependency monitoring.Develop unified observability dashboards covering SIEM platform state...  ...regular upgrades.Experience building SRE-style observability and reliability patterns... 
    Full time
    Work at office
    Shift work
    Day shift

    Bank of America

    Addison, TX
    5 days ago
  •  ...Duration: Long Term Contract Pay Rate: $40/Hr. W2 Experience: 3-5 Years Overview We are seeking a remote Junior SRE/DevOps Engineer role. The ideal candidate has foundational knowledge of Site Reliability Engineering (SRE) and Kubernetes, and is enthusiastic about growing... 
    Long term contract
    Contract work
    Internship
    Remote work

    BayOne Solutions

    Richardson, TX
    2 days ago
  • $140k - $155k

     ...career growth potential. Job Title: Platform Network Engineer Location: 100% Remote (U.S.)...  ...Istio and Linkerd — that provide secure, observable, and reliable service-to-service communication...  ...experience in platform engineering, SRE, or networking roles. Hands-on... 
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    Carrollton, TX
    17 days ago
  • $40 per hour

     ...A technology solutions provider is seeking a remote Junior SRE/DevOps Engineer. The ideal candidate should have foundational knowledge of Site Reliability Engineering (SRE) and Kubernetes. Responsibilities include gaining experience in a DevOps-driven environment. Applicants... 
    Long term contract
    Internship
    Remote work

    BayOne Solutions

    Richardson, TX
    2 days ago
  • $133k - $158k

     ...growth potential. Job Title: Azure Platform Engineer Location: 100% Remote (U.S.) Position...  ...security hardening, cost optimization, observability, and ongoing operational excellence for...  ...application development, security, and SRE teams to deliver cloud-native solutions... 
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    Carrollton, TX
    17 days ago
  •  ...currently available and apply online.Job SummaryThe Principal Devops Platform Engineer will be responsible for driving development and strategy for...  ...versed in Devops or Site reliability engineering (SRE) tenets2+ years in an SRE/Operations/DevOps role as part of a... 
    Full time
    Local area

    TXU Energy

    Irving, TX
    5 days ago
  • $138.4k - $173k

     ...reliability, quality of services and overall observability patterns. Along with your team, you’ll...  .... You’ll collaborate or embed with engineering teams, helping them to improve the...  ...components of the AppFolio Real Estate Platform. You’ll help build the future of reliable... 
    Full time
    Flexible hours

    AppFolio

    Dallas, TX
    5 days ago
  • $179.2k - $320.2k

     ...Role:As a Director in Software Engineering, you will provide...  ...management of container deployment platforms and the CI/CD and applications...  ..., platform engineering, and SRE functions.ESSENTIAL DUTIES...  ...networking, storage, data, security, observability), while performing hands-on... 
    Full time
    Work experience placement
    Work at office

    Wolters Kluwer

    Coppell, TX
    5 days ago
  •  ...Infrastructure as Code practicesacross Workplace platforms using Terraform and Atmos, enabling teams...  ...automated validation, testing patterns, observability, and disciplined release governance. Partner closely with platform engineers, security, infrastructure, and networking... 
    Full time
    Work at office

    Vanguard

    Dallas, TX
    3 days ago
  • $139.3k - $203.6k

     ...sufficient number of applications are received.Senior Kubernetes Platform Engineer - AI Infrastructure - hybrid (2013580)***hybrid role...  ..., and self-healing systems for platform reliabilityImprove observability (metrics, logs, traces) and optimize resource utilization,... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    Dallas, TX
    5 days ago
  • $140.88k - $153.75k

    General Information Job Title Senior Platform Engineer Job ID 105799 Work Areas Technology & Engineering Employment Type...  ...Infrastructure to ensure services are secure-by-default, observable from day one, and operable under production on-call expectations... 
    Permanent employment
    Full time
    Contract work
    Work at office
    Local area

    Bain & Company

    Dallas, TX
    1 day ago
  • About LanternLantern is the specialty care platform connecting people with the best care...  ...an experienced Senior Cloud Platform Engineer to drive the evolution of our Azure-based...  ...cycles.Ensure robust monitoring, observability, and alerting for production systems; proactively... 

    Lantern

    Dallas, TX
    3 days ago
  •  ...GorusuCompany: SRI Tech SolutionsJob Title: Platform EngineerLocation: Dallas, TX (Day 1...  ...Face to FaceLooking for a skilled Platform Engineer with deep knowledge and extensive...  ...solutions to ensure secure, resilient, and observable service-to-service communication.Collaborate... 

    SRI Tech

    Dallas, TX
    5 days ago
  •  ...even if the perfect role isn’t open just yet. We recruit Lead Platform Engineers on a rolling basis throughout the year, and this Evergreen...  ...Recognition High problem determination drive Experience in observability metrics and analysis Bachelor's degree or equivalent... 
    Work experience placement
    Immediate start
    Remote work
    Flexible hours

    DTCC- The Depository Trust & Clearing Corporation

    Dallas, TX
    5 days ago
  •  ...Will Have in This RoleAs a Lead Systems Engineer - Messaging Platforms, you will play a key role in...  ...ensuring messaging services are scalable, observable, and aligned with critical business needs...  ...of Site Reliability Engineering (SRE) concepts such as observability, SLAs... 
    Remote work
    Flexible hours

    DTCC- The Depository Trust & Clearing Corporation

    Dallas, TX
    4 days ago
  •  ...in This RoleAs a Principal Systems Engineer - Messaging Platforms, you will serve as the technical authority...  ...a Site Reliability Engineering (SRE) mindset, emphasizing reliability,...  ...promoting best practices in automation, observability, and reliability engineeringAuthor... 
    Remote work
    Flexible hours

    DTCC- The Depository Trust & Clearing Corporation

    Dallas, TX
    1 day ago
  • $93k - $189k

     ...see location options on posting)The z/OS Platform CICS, MQ & zCEE Lead candidate must have...  ...areas of experience include z/OS engineering, WLM (Workload Manager), CF, SMF Processing...  ...TLS, Resource access & RACF)Proficient in SRE concepts and automationProven experience... 
    Full time
    Work at office
    Remote work
    Work from home
    Flexible hours

    Cadence Bank

    Dallas, TX
    1 day ago
  • $122k - $200k

     ...make an impact. Join us!Position Summary:This is a senior platform engineering role responsible for architecting, building, and advancing...  ...development, evaluation, deployment, monitoring, governance, and observability.Define and implement enterprise standards, reference... 
    Full time
    Work at office
    Flexible hours
    Day shift

    Bank of America

    Addison, TX
    3 days ago
  •  ...Position Summary :This is a hands-on software engineering role focused on building enterprise-grade Generative AI, Data Science, and AI Platform capabilities within Bank of America's...  ..., inferencing, monitoring, and observability.Implement event-driven and streaming solutions... 
    Full time
    Work at office
    Flexible hours
    Day shift

    Bank of America

    Addison, TX
    1 day ago
  •  ...Takes food and beverage orders, places orders, delivers orders, checks back after delivery of food to ensure guest satisfaction, observes guests to respond to any additional needs Maintains table appearance by pre-bussing, checks drink levels, removes clutter and provides... 
    Work experience placement
    Casual work
    Local area
    Flexible hours

    Starbucks

    Farmers Branch, TX
    4 days ago
  •  ...cultivated. Our developers and infrastructure engineers are the driving force behind our...  ...and fearless. You understand that great platform work is invisible - when you do it...  ...their infrastructure, deployments, and observability needs without waiting on manual ops work... 

    Gravitate Energy LLC

    Dallas, TX
    4 days ago
  •  ...Senior Veritas eDiscovery Platform (eDP) Engineer Employment Type: Full-Time, Executive-Level Department: Legal CGS is seeking a dedicated...  ..., Dependent Care, and Commuter) ~ Paid Time Off and Observance of State/Federal Holidays Contact Government... 
    Full time
    For contractors
    Remote work
    Flexible hours

    Contact Government Services LLC

    Dallas, TX
    4 days ago
  •  ...a few. About the role We’re looking for a software engineer to join our small but growing platform team. Platform engineering at Knock is the foundation...  ...sent Significantly improved latency and margins of our observability product features by adopting ClickHouse for event... 
    Remote work

    GrabJobs

    Irving, TX
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Observability Platform Engineer (SRE). Be the first to apply!