Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineering Manager

Full-time

NationsBenefits, LLC

NationsBenefits is recognized as one of the fastest-growing companies in America and a Healthcare Fintech provider of supplemental benefits, flex cards, and member engagement solutions. We partner with managed care organizations to provide innovative healthcare solutions that drive growth, improve outcomes, reduce costs, and bring value to their members.Through our comprehensive suite of innovative supplemental benefits, fintech payment platforms, and member engagement solutions, we help health plans deliver high-quality benefits to their members that address the social determinants of health and improve member health outcomes and satisfaction.Our compliance-focused infrastructure, proprietary technology systems, and premier service delivery model allow our health plan partners to deliver high-quality, value-based care to millions of members.We offer a fulfilling work environment that attracts top talent and encourages all associates to contribute to delivering premier service to internal and external customers alike. Our goal is to transform the healthcare industry for the better! We provide career advancement opportunities from within the organization across multiple locations in the US, South America, and India.

Location: Remote (US-based candidates only)

Manager, Site Reliability Engineering (SRE)

Position Overview

We are seeking a Manager, Site Reliability Engineering (SRE) to lead our US-based SRE team and drive operational excellence across our production platforms.

This is a player-coach leadership role that combines people management with hands-on technical leadership. You will mentor and grow a team of Site Reliability Engineers while actively participating in major incident response, reliability initiatives, and operational reviews. The role is a key part of our global follow-the-sun support model and requires close collaboration with SRE leadership in India.

Key Responsibilities

Team Leadership & Development

  • Lead, mentor, and develop a US-based team of Site Reliability Engineers.
  • Conduct regular 1:1s, performance reviews, and career development discussions.
  • Own hiring, onboarding, and retention efforts as the team scales.
  • Foster a culture of ownership, blameless postmortems, and continuous improvement.

Operational Excellence & Incident Management

  • Lead day-to-day production operations and ensure timely incident triage, resolution, and escalation.
  • Serve as an escalation point and incident commander for major production incidents.
  • Drive problem management and root cause analysis processes.
  • Carry PagerDuty on-call escalation responsibilities for critical issues.
  • Track and report operational KPIs, SLAs, and SLOs, including availability, MTTR, and incident trends.

Reliability & Automation

  • Improve system reliability, observability, and resilience using Datadog and related tooling.
  • Drive automation, self-healing capabilities, and runbook maturity.
  • Partner with Development, DevOps, DevSecOps, and Engineering teams to embed reliability into the SDLC.
  • Contribute hands-on to tooling, automation, and technical reviews as needed.

Collaboration & Global Alignment

  • Coordinate closely with SRE leadership in India to ensure seamless follow-the-sun coverage.
  • Represent the US SRE organization in cross-functional planning and operational reviews.
  • Communicate effectively with both technical and non-technical stakeholders.

Documentation & Compliance

  • Maintain high-quality documentation for incidents, postmortems, runbooks, and operational procedures.
  • Ensure adherence to healthcare and fintech compliance standards, including HIPAA, PCI DSS, SOC 2, ISO 27001, and HITRUST.

Required Qualifications

  • 5–8 years of experience in Site Reliability Engineering, DevOps, Production Support, or Platform Engineering.
  • 1–2+ years of experience leading, mentoring, or managing engineers.
  • Demonstrated success operating in a player-coach leadership model.
  • Strong hands-on experience with production incident management and escalation processes.
  • Proficiency with Datadog or similar observability platforms.
  • Hands-on experience with Kubernetes and Docker in production environments.
  • Strong scripting or programming skills in PowerShell, Bash, Python, Java, or C#.
  • Experience with Helm, CI/CD pipelines, and deployment automation.
  • Working knowledge of ITIL processes and Agile methodologies.
  • Experience working with SQL, MySQL, or NoSQL databases.
  • Excellent communication and stakeholder management skills.
  • Willingness to participate in PagerDuty on-call escalation and work within a global follow-the-sun operating model.

Preferred Qualifications

  • Experience with cloud platforms such as AWS, Azure, or GCP.
  • Experience building or scaling SRE teams and on-call programs.
  • Experience defining and managing SLOs, SLIs, and error budgets.
  • Prior experience in the healthcare or fintech industry.
  • Knowledge of security and compliance frameworks relevant to regulated environments.

Why Join NationsBenefits?

  • Competitive compensation and comprehensive benefits.
  • Unlimited PTO.
  • Fully remote work environment (US-based).
  • Opportunity to lead and grow a high-impact SRE organization.
  • Exposure to modern cloud-native technologies and large-scale reliability challenges.
  • Collaborative culture focused on innovation, learning, and continuous improvement.
  • Meaningful work that directly impacts healthcare technology and millions of members.

Ideal Candidate

We are looking for a technically strong SRE leader who enjoys building teams, improving operational maturity, and remaining hands-on during critical production events. The ideal candidate combines leadership, systems thinking, and automation expertise to help scale reliability practices across a fast-growing Healthcare FinTech organization.

NationsBenefits is an Equal Opportunity Employer.
Vacancy posted 20 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineering Manager in Remote vacancy
  •  ...Role Overview Help us ensure the reliability of Ajaib's fintech platform, serving millions of Indonesian investors. You'll lead...  ...We're Looking For - 5+ years in SRE/DevOps, with 2+ years managing engineers - Deep hands-on expertise in GCP and Kubernetes -... 
    Suggested
    Remote work

    Air

    United States
    4 days ago
  •  ...Site Reliability Engineering Manager Canonical is a leading provider of open-source software and operating systems for global enterprise and technology markets. Our platform, Ubuntu, is very widely used in breakthrough enterprise initiatives such as public cloud, data... 
    Suggested
    Work at office
    Local area
    Remote work
    Work from home
    Worldwide

    Canonical

    United States
    2 days ago
  • $155k - $180k

     ...Site Reliability Engineering Manager (Remote) Join to apply for the Site Reliability Engineering Manager (Remote) role at WebstaurantStore Base pay range: $155,000.00/yr - $180,000.00/yr Job Summary As the largest online distributor of restaurant supplies... 
    Suggested
    H1b
    Remote work
    Home office

    WebstaurantStore

    Lititz, PA
    1 day ago
  • $132.77k - $221.35k

     ...for our SRE function. In this role, you will mature the Site Reliability Engineering (SRE) function ensuring reliability and performance of organization...  ...Inc. (Nasdaq: LPLA) is among the fastest growing wealth management firms in the U.S. As a leader in the financial advisor-... 
    Suggested
    Full time
    Work from home

    LPL Financial

    Fort Mill, York County, SC
    3 days ago
  • $142.8k - $274.8k

     ...yearEmployment type: Full-TimeWork site: 0 days / week in-office -...  ...EngineeringDiscipline: Site Reliability EngineeringCompany:...  ...a Principal Site Reliability Engineer, you will set technical and operational...  ...eligibility requirements.For manager-level roles, a Tier 5 (T5)... 
    Suggested
    Ongoing contract
    Work at office
    Local area

    Microsoft

    Redmond, WA
    2 days ago
  •  ...sponsorship.Maintain and enhance the reliability, availability, and...  ...page of the Navy Federal Career Site.Protect Yourself from Job Scams...  ...degree in computer science, engineering, or the equivalent...  ...knowledge of incident response and managing production issues. Advanced communication... 
    Internship
    Monday to Friday

    Navy Federal Credit Union

    Vienna, VA
    14 hours ago
  •  .... Connecting. Growing together.We are seeking a Principal Site Reliability Engineer (SRE) to define and scale reliability practices across large...  ...in:Reliability engineering (SLOs, SLIs, incident management, observability)Distributed systems in cloud environments (... 
    Minimum wage
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work

    UnitedHealth Group

    Minnetonka, MN
    14 hours ago
  • $163.62k - $212.71k

     ...platforms, and processes that improve our engineering teams' productivity and streamline the...  ...seasoned and strategic Lead/Principal Site Reliability Engineer to drive the reliability,...  ...Operations (SRE Focus)Platform Design and Management: Architect, build, and maintain... 
    Full time
    Part time
    Work experience placement
    Work at office
    Local area
    Immediate start
    Remote work
    Work from home
    Flexible hours
    Shift work
    3 days per week
    1 day per week

    iSpot.tv

    Bellevue, WA
    2 days ago
  • $159k - $272k

     ...people thrive in an evolving world. As a premier global asset management organization with more than 85 years of experience, we...  ...that matter to you. Role SummaryIn this role as Principal Site Reliability Engineer, Infrastructure Observability you will help formulate, develop... 
    Full time
    Private practice
    Local area
    Remote work
    Work from home
    3 days per week

    T. Rowe Price

    Owings Mills, MD
    14 hours ago
  •  ...Principal Site Reliability Engineer location- Washington, DC -Onsite Remote- No 6+ Months Job Summary At Amtrak,...  ...or similar. Infrastructure Automation: Implement and manage IaC using tools like Terraform, AWS CloudFormation, or... 
    Remote work

    American IT Systems

    Washington DC
    1 day ago
  • $90k - $130k

     ...-on experience. We require 10+ years of experience in Site Reliability Engineering, Software Engineering, or Cloud Engineering. We need experience...  ...reliability engineering, including SLOs, SLIs, incident management, and observability. We require experience working with... 
    Full time
    Remote work

    UnitedHealth Group

    Minnetonka, MN
    3 days ago
  • $7,000 per month

     ...Principal Site Reliability Engineer Latin America The salary range for this role is negotiable, the range being $7000 - $12000 per month (Gross in USD) About Sezzle: With a mission to financially empower the next generation, Sezzle is revolutionizing the shopping... 
    Remote work
    Flexible hours

    Sezzle

    United States
    4 days ago
  •  ...Working remotely within the United States, the full-time Principal Site Reliability Engineer will lead project work to enhance platform reliability, mentor junior engineers, and engage in incident response while collaborating closely with product stakeholders and architects... 
    Full time
    Remote work

    Virtual Vocations Inc

    United States
    2 days ago
  • $175k - $220k

     ...across the U.S., Canada, and India. The Director, Site Reliability Engineering (SRE) will lead reliability, performance, and...  ...CD practices for assigned product families. Directors will manage multiple teams and collaborate with Product Development, Architecture... 
    Contract work
    Temporary work
    Work at office
    Work from home
    Flexible hours

    Vertafore

    Denver, CO
    2 days ago
  •  ...Site Reliability Engineering (SRE) Team Lead The Site Reliability Engineering (SRE) team is foundational to the growth and scale of our platform...  ...pipelines, using CI/CD systems and configuration management Why You'll Like It Here We are collaborative at... 
    Shift work

    Roberts Recruiting

    Boston, MA
    2 days ago
  • $200k - $250k

     ...progress, come build the future together. As a Principal Site Reliability Engineer, you'll shape the long-term strategy for the...  ...across critical infrastructure, including cluster lifecycle management, networking, identity and access management, observability... 
    Full time
    Immediate start
    Remote work

    DraftKings

    United States
    3 days ago
  •  ...Senior Principal Site Reliability Engineer Hong Kong SAR About Us Established in 2018, Bybit is one of the world's leading cryptocurrency...  ...a seamless ecosystem across trading, payments, wealth management, custody, institutional services, and Web3 — connecting users... 
    Remote work

    Bybit

    United States
    4 days ago
  •  ...Principal Site Reliability Engineer Deimos is a cloud-native developer and security operations technology services company. We help companies...  ...projects. You will report to a Site Reliability Engineering Manager. As a Principal Site Reliability Engineer you will be... 
    Currently hiring
    Remote work
    Work from home

    Deimos

    United States
    2 days ago
  •  ...Senior Manager, Site Reliability Engineering Remote - USA At Counterpart Health, we are transforming healthcare and improving patient care with our innovative primary care tool, Counterpart Assistant. By supporting Primary Care Physicians (PCPs), we deliver improved... 
    Work at office
    Remote work
    Flexible hours
    Shift work

    Clover Health

    United States
    1 day ago
  • $167.3k - $242.6k

     ...Principal Site Reliability Engineer Remote Your passion for uptime was forged from experience in production and refined through incident...  ...interview. We also have a goal that all Expletives have a great manager and have a voice in how their team is run and who runs it.... 
    Remote work
    Visa sponsorship
    Day shift

    Expel

    United States
    2 days ago
  •  ...software development teams to build reliable, scalable, secure, and cloud-...  ...architecture patterns across engineering teams, helping ensure systems...  ...of hands-on experience in Site Reliability Engineering,...  ...observability, monitoring, and incident management tools such as Honeycomb.io,... 
    Remote work

    ABC Fitness Solutions, LLC

    United States
    4 days ago
  • $230k - $255k

     ...Join Aya Healthcare, winner of multiple Top Workplace awards! We're looking for a highly experienced Manager, Site Reliability Engineering to lead the team behind one of healthcare's most relied-on workforce platforms. In this leadership role, you'll guide and... 
    Local area
    Remote work

    Aya Healthcare

    United States
    2 days ago
  • $160k - $180k

     ..., Canada, and India. We are seeking a Principal Site Reliability Engineer to define the strategic vision and own the enterprise-wide...  ...: Establish the governance models for defining and managing SLIs and SLOs across multiple product lines. ~ Delivery... 
    Contract work
    Temporary work
    Work at office
    Work from home
    Flexible hours

    Vertafore

    Denver, CO
    3 days ago
  • $84.9k - $209.5k

     ...spirit that promotes an upbeat and creative environment. We are unencumbered and will need your contribution to make it a special engineering center with the focus on excellence. Health Data Intelligence Platform has a rare opportunity to play a critical role in how... 
    Temporary work
    Immediate start
    Remote work
    Flexible hours

    Hackajob

    United States
    4 days ago
  •  ...trillion in crypto transactions. We are looking for a Head of Site Reliability Engineering (SRE) who will serve as the principal leader in...  ...join an exceptional crypto company and senior engineering management team at a time of high growth. This is a tremendous opportunity... 
    Full time
    Apprenticeship
    Work at office
    Remote work
    Worldwide
    Flexible hours

    Blockchain

    United Kingdom
    11 days ago
  • $189.59k - $220k

     ...Director, Site Reliability Engineering NBCUniversal is one of the world’s leading media and entertainment companies. We create world‑class content...  ...for hands‑on configuration and support as well as managing the work of other architects and engineers. ~ Work closely... 
    Full time
    For contractors
    Remote work

    NBCUniversal

    New York, NY
    1 day ago
  • Role Description Symmetrio is recruiting a Principal Site Reliability Engineer (SRE) for our customer, a rapidly growing healthcare technology...  ...Qualifications ~6+ years of hands-on experience supporting and managing AWS-based production environments ~4+ years of... 
    Full time

    Symmetrio

    Remote
    2 days ago
  • $151k - $297k

     ..., you will partner with SRE leaders and engineers to scale the platform that underpins all...  ...program execution, strengthen production reliability practices, and coordinate cross-...  ...criteria with SRE engineers and leaders. Manage dependencies across platform teams, keep... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    New York, NY
    14 hours ago
  •  ...Information Technology group delivers secure, reliable technology solutions that enable DTCC...  ...enterprise platforms.As a Principal Site Reliability Engineer (SRE), you will drive operational...  ...related monitoring platforms. Define and manage SLIs, SLOs, dashboards, alerts, and... 
    Remote work
    Flexible hours

    DTCC- The Depository Trust & Clearing Corporation

    Jersey City, NJ
    4 days ago
  • $132.77k - $221.35k

     ...teams to ensure products are designed with observability, reliability, and performance in mind, promoting operational...  ...with key stakeholders across the organization, including engineering, operations, and management. Create best‑in‑class reports and prepare presentations... 
    Work from home

    LPL Financial LLC

    Fort Mill, York County, SC
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineering Manager. Be the first to apply!