Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineering Manager

Full-time

NationsBenefits, LLC

NationsBenefits is recognized as one of the fastest-growing companies in America and a Healthcare Fintech provider of supplemental benefits, flex cards, and member engagement solutions. We partner with managed care organizations to provide innovative healthcare solutions that drive growth, improve outcomes, reduce costs, and bring value to their members.Through our comprehensive suite of innovative supplemental benefits, fintech payment platforms, and member engagement solutions, we help health plans deliver high-quality benefits to their members that address the social determinants of health and improve member health outcomes and satisfaction.Our compliance-focused infrastructure, proprietary technology systems, and premier service delivery model allow our health plan partners to deliver high-quality, value-based care to millions of members.We offer a fulfilling work environment that attracts top talent and encourages all associates to contribute to delivering premier service to internal and external customers alike. Our goal is to transform the healthcare industry for the better! We provide career advancement opportunities from within the organization across multiple locations in the US, South America, and India.

Location: Remote (US-based candidates only)

Manager, Site Reliability Engineering (SRE)

Position Overview

We are seeking a Manager, Site Reliability Engineering (SRE) to lead our US-based SRE team and drive operational excellence across our production platforms.

This is a player-coach leadership role that combines people management with hands-on technical leadership. You will mentor and grow a team of Site Reliability Engineers while actively participating in major incident response, reliability initiatives, and operational reviews. The role is a key part of our global follow-the-sun support model and requires close collaboration with SRE leadership in India.

Key Responsibilities

Team Leadership & Development

  • Lead, mentor, and develop a US-based team of Site Reliability Engineers.
  • Conduct regular 1:1s, performance reviews, and career development discussions.
  • Own hiring, onboarding, and retention efforts as the team scales.
  • Foster a culture of ownership, blameless postmortems, and continuous improvement.

Operational Excellence & Incident Management

  • Lead day-to-day production operations and ensure timely incident triage, resolution, and escalation.
  • Serve as an escalation point and incident commander for major production incidents.
  • Drive problem management and root cause analysis processes.
  • Carry PagerDuty on-call escalation responsibilities for critical issues.
  • Track and report operational KPIs, SLAs, and SLOs, including availability, MTTR, and incident trends.

Reliability & Automation

  • Improve system reliability, observability, and resilience using Datadog and related tooling.
  • Drive automation, self-healing capabilities, and runbook maturity.
  • Partner with Development, DevOps, DevSecOps, and Engineering teams to embed reliability into the SDLC.
  • Contribute hands-on to tooling, automation, and technical reviews as needed.

Collaboration & Global Alignment

  • Coordinate closely with SRE leadership in India to ensure seamless follow-the-sun coverage.
  • Represent the US SRE organization in cross-functional planning and operational reviews.
  • Communicate effectively with both technical and non-technical stakeholders.

Documentation & Compliance

  • Maintain high-quality documentation for incidents, postmortems, runbooks, and operational procedures.
  • Ensure adherence to healthcare and fintech compliance standards, including HIPAA, PCI DSS, SOC 2, ISO 27001, and HITRUST.

Required Qualifications

  • 5–8 years of experience in Site Reliability Engineering, DevOps, Production Support, or Platform Engineering.
  • 1–2+ years of experience leading, mentoring, or managing engineers.
  • Demonstrated success operating in a player-coach leadership model.
  • Strong hands-on experience with production incident management and escalation processes.
  • Proficiency with Datadog or similar observability platforms.
  • Hands-on experience with Kubernetes and Docker in production environments.
  • Strong scripting or programming skills in PowerShell, Bash, Python, Java, or C#.
  • Experience with Helm, CI/CD pipelines, and deployment automation.
  • Working knowledge of ITIL processes and Agile methodologies.
  • Experience working with SQL, MySQL, or NoSQL databases.
  • Excellent communication and stakeholder management skills.
  • Willingness to participate in PagerDuty on-call escalation and work within a global follow-the-sun operating model.

Preferred Qualifications

  • Experience with cloud platforms such as AWS, Azure, or GCP.
  • Experience building or scaling SRE teams and on-call programs.
  • Experience defining and managing SLOs, SLIs, and error budgets.
  • Prior experience in the healthcare or fintech industry.
  • Knowledge of security and compliance frameworks relevant to regulated environments.

Why Join NationsBenefits?

  • Competitive compensation and comprehensive benefits.
  • Unlimited PTO.
  • Fully remote work environment (US-based).
  • Opportunity to lead and grow a high-impact SRE organization.
  • Exposure to modern cloud-native technologies and large-scale reliability challenges.
  • Collaborative culture focused on innovation, learning, and continuous improvement.
  • Meaningful work that directly impacts healthcare technology and millions of members.

Ideal Candidate

We are looking for a technically strong SRE leader who enjoys building teams, improving operational maturity, and remaining hands-on during critical production events. The ideal candidate combines leadership, systems thinking, and automation expertise to help scale reliability practices across a fast-growing Healthcare FinTech organization.

NationsBenefits is an Equal Opportunity Employer.
Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineering Manager in Remote vacancy
  • $182k - $250.8k

     ...at Okta is the backbone of our platform's reliability and operational excellence. We are a forward-thinking group of engineers and leaders who believe that great infrastructure...  ...for millions of users worldwide. As a Manager, Site Reliability Engineer, you'll lead this... 
    Suggested
    Permanent employment
    Local area
    Remote work
    Worldwide
    Flexible hours
    Weekend work
    Weekday work

    Okta

    Washington DC
    1 day ago
  • $134.6k - $230.8k

     ...Connecting. Growing together.Are you passionate about reimagining operations through AI? Optum Financial is seeking a Principal Site Reliability Engineer to lead the evolution of our reliability platform by combining modern SRE practices with AI-assisted operations. You'll... 
    Suggested
    Minimum wage
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work

    UnitedHealth Group

    Eden Prairie, MN
    1 day ago
  •  .... Connecting. Growing together.We are seeking a Principal Site Reliability Engineer (SRE) to define and scale reliability practices across large...  ...in:Reliability engineering (SLOs, SLIs, incident management, observability)Distributed systems in cloud environments (... 
    Suggested
    Minimum wage
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work

    UnitedHealth Group

    Minnetonka, MN
    2 days ago
  •  ...Service (S3), and Auto Scaling Groups for dynamic resource management. Designs, develops, and executes performance tests using...  ...Delivery (CI/CD) pipelines and Kubernetes. Supports Site Reliability Engineering (SRE) functions by establishing Service Level Objectives... 
    Suggested
    Full time

    Fidelity Investments

    Durham, NC
    3 days ago
  • $142.8k - $274.8k

     ...yearEmployment type: Full-TimeWork site: 0 days / week in-office -...  ...EngineeringDiscipline: Site Reliability EngineeringCompany:...  ...a Principal Site Reliability Engineer, you will set technical and operational...  ...eligibility requirements.For manager-level roles, a Tier 5 (T5)... 
    Suggested
    Ongoing contract
    Work at office
    Local area

    Microsoft

    Redmond, WA
    4 days ago
  • $169.3k - $304.7k

     ...Our team designs, develops, and manages applications and infrastructure that...  ...maintaining fast, efficient, scalable, and reliable routing software and infrastructure that...  ...global platform. As a Principal Site Reliability Engineer - Network, you will be responsible... 
    Work experience placement
    Work at office
    Remote work

    Akamai

    Eastern, KY
    3 days ago
  •  ...Setting the reliability strategy for the platform, the full-time Principal Site Reliability Engineer will define deployment and operational standards for distributed systems, ensuring reliability and automation across customer environments while working remotely. Key... 
    Full time
    Remote work

    Virtual Vocations Inc

    United States
    4 days ago
  • $159k - $272k

     ...people thrive in an evolving world. As a premier global asset management organization with more than 85 years of experience, we...  ...that matter to you. Role SummaryIn this role as Principal Site Reliability Engineer, Infrastructure Observability you will help formulate, develop... 
    Full time
    Private practice
    Local area
    Remote work
    Work from home
    3 days per week

    T. Rowe Price

    Owings Mills, MD
    2 days ago
  • $175k - $220k

     ...across the U.S., Canada, and India. The Director, Site Reliability Engineering (SRE) will lead reliability, performance, and...  ...CD practices for assigned product families. Directors will manage multiple teams and collaborate with Product Development, Architecture... 
    Contract work
    Temporary work
    Work at office
    Work from home
    Flexible hours

    Vertafore

    Denver, CO
    4 days ago
  •  ...Manager Of Site Reliability Engineering Delinea is looking for a hands-on Manager of Site Reliability Engineering to lead the SRE and DevOps engineers supporting the Delinea products. This is a working manager role. You will be expected to lead people and lead work:... 
    Full time
    For contractors
    H1b
    Local area
    Remote work

    Delinea

    United States
    3 days ago
  •  ...technology designed to keep the electric grid secure and reliable, even during extended periods of stress. By...  ...right place. Role Description Form Energy is hiring a Manager, Site Reliability Engineer to lead the operational function responsible for maintaining... 
    Full time
    Remote work
    Relocation package

    Form Energy, Inc.

    Berkeley, CA
    2 days ago
  • $122k - $207k

     ...services that help people, businesses and governments realize their greatest potential. Title and Summary Manager, Site Reliability Engineering Who is Mastercard? At Mastercard technology, we work to connect and power an inclusive, digital economy that... 
    Full time
    Part time
    Worldwide
    Flexible hours

    MasterCard

    O Fallon, MO
    2 days ago
  • $207k - $300k

    Manage a team of Software/Systems Engineers on projects for users and remain directly responsible for uptime.Own...  ...and establishing sustainable multi-site on-call rotations across...  ....Deep practical expertise in Site Reliability Engineering practices, including SLO... 

    Google

    San Jose, CA
    1 day ago
  • $262k - $364k

    Lead a team of Software/Systems Engineers on projects for users and be directly responsible...  ...through quality technical execution.Manage on-call rotations across continents, using...  ...or Engineering, or a related field.Site Reliability Engineering (SRE) combines software and... 

    Google

    Sunnyvale, TX
    3 days ago
  • $207k - $300k

    Lead a team of Software/Systems Engineers on projects for users and be directly responsible...  ...through quality technical execution.Manage on-call rotations across continents, using...  ...year of people management experience. Site Reliability Engineering (SRE) combines software and... 

    Google

    Mountain View, CA
    4 days ago
  • $207k - $300k

    Lead a team of Software/Systems Engineers on projects for users and be directly responsible...  ...through quality technical execution.Manage on-call rotations across continents, using...  ...of engineers on large-scale projects.Site Reliability Engineering (SRE) combines software and... 

    Google

    New York, NY
    6 hours ago
  •  ...I.T. is actively seeking a Principal Engineer for an immediate full-time opportunity...  ...technology company is seeking experienced Site Reliability Engineers to take ownership of...  ...including SLI/SLO frameworks and error budget management Establish escalation protocols and... 
    Full time
    Contract work
    Immediate start
    Work from home
    Flexible hours

    KēSTA I.T.

    Beverly Hills, CA
    7 days ago
  • $160k - $180k

     ...S., Canada, and India. We are seeking a Principal Site Reliability Engineer to define the strategic vision and own the enterprise-wide...  ...: Establish the governance models for defining and managing SLIs and SLOs across multiple product lines. ~ Delivery... 
    Contract work
    Work from home
    Flexible hours

    Vertafore

    Denver, CO
    28 days ago
  •  ...Information Technology group delivers secure, reliable technology solutions that enable DTCC...  ...enterprise platforms.As a Principal Site Reliability Engineer (SRE), you will drive operational...  ...related monitoring platforms. Define and manage SLIs, SLOs, dashboards, alerts, and... 
    Remote work
    Flexible hours

    DTCC- The Depository Trust & Clearing Corporation

    Jersey City, NJ
    2 days ago
  • $140k - $230k

     ...Zoox is seeking a Site Reliability Engineer to help ensure the availability, performance, and resilience of the services that power the development...  ...or a similar role, with a strong, objective background in managing large-scale distributed systems. Cloud & Infrastructure... 
    Full time

    Zoox

    Remote
    1 day ago
  •  ...Sophos seeks an experienced Manager, Software Engineering (SRE) to lead a distributed team across the U.S. and Canada, focusing on reliability, scalability, and efficient cloud operations. You will guide AWS, Kubernetes/EKS, Terraform/IaC, and automation efforts while... 
    Remote job

    Jobleads-US

    Kentucky
    2 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range...  ...observability and alerting systems.The Fleet Management team provides the core runtime...  ...critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Seattle, WA
    2 days ago
  •  ...GorusuCompany: SRI Tech SolutionsJob Title: Senior Site Reliability EngineerLocation: Plano , TX (remote)...  ...highly skilled Senior Site Reliability Engineer (SRE) to join our dynamic team. The...  ..., Shell, or other scripting languages.Manage and troubleshoot Kubernetes clusters... 
    Remote work

    SRI Tech

    Plano, TX
    6 hours ago
  •  ...education. Client is currently seeking a talented Software Engineer who is able to work into the Site Reliability Engineer role. This candidate is expected to work...  ..., and performance levels. Reporting to the Manager, you will form part of a global team that will drive... 
    Remote work

    Intelliswift

    Durham, NC
    6 hours ago
  •  ...ServicesSelling Points Contribute to the reliability of a high-transaction payment...  ...principles.Job DescriptionSite Reliability Engineer OverviewThe Site Reliability Engineer ensures the...  ...procedures to reduce operational toil.Manage RDS SQL Server deployments, ensuring... 
    Remote work

    Green Key Resources

    Columbus, OH
    2 days ago
  • $134.25k - $214.8k

     ...where you matter.Your ImpactAre you an engineer who gets excited about the challenge...  ...the Observability team within Axon's Site Reliability organization — a focused team responsible...  ...ArgoCD, and Helm — including capacity management, cybersecurity requirements and... 
    Work experience placement
    Work at office
    Remote work

    Axon

    Boston, MA
    3 days ago
  •  ...thinking organization, apply now.We are currently seeking a Site Reliability Engineer to join our team in Westlake, Texas (US-TX), United States...  ...solutions for infrastructure provisioning, configuration management, and operational workflows.Support and enhance CI/CD... 
    Temporary work
    Work at office
    Remote work
    Flexible hours

    NTT DATA

    Texas
    2 days ago
  • $104.9k - $174.7k

     ...Verification, Fraud and Credit Risk mitigation and Customer Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability... 
    Full time
    Work at office
    Local area
    Remote work
    Work from home

    RELX Group

    Buford, GA
    1 day ago
  • $102.1k - $202.2k

     ...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole...  ...EngineeringDiscipline: Site Reliability EngineeringCompany:...  ...secure, resilient, and easier to manage for customers operating in highly...  ..., you will collaborate with engineers across disciplines to deliver... 
    Ongoing contract
    Work experience placement
    Local area
    Remote work
    3 days per week

    Microsoft

    Redmond, WA
    1 day ago
  • $87.12k - $151.25k

     ...organization, apply now.We are currently seeking a Digital Site Reliability Sr Engineer - Remote to join our team in Memphis, Tennessee (US-TN),...  ....5+ years experience with ServiceNow and Jira for change management.4+ years Unix/Linux shell scripting and supporting NoSQL... 
    Temporary work
    Work at office
    Remote work
    Flexible hours

    NTT DATA

    Memphis, TN
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineering Manager. Be the first to apply!