Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineering Manager

NationsBenefits, LLC

NationsBenefits is recognized as one of the fastest-growing companies in America and a Healthcare Fintech provider of supplemental benefits, flex cards, and member engagement solutions. We partner with managed care organizations to provide innovative healthcare solutions that drive growth, improve outcomes, reduce costs, and bring value to their members.Through our comprehensive suite of innovative supplemental benefits, fintech payment platforms, and member engagement solutions, we help health plans deliver high-quality benefits to their members that address the social determinants of health and improve member health outcomes and satisfaction.Our compliance-focused infrastructure, proprietary technology systems, and premier service delivery model allow our health plan partners to deliver high-quality, value-based care to millions of members.We offer a fulfilling work environment that attracts top talent and encourages all associates to contribute to delivering premier service to internal and external customers alike. Our goal is to transform the healthcare industry for the better! We provide career advancement opportunities from within the organization across multiple locations in the US, South America, and India.

Location: Remote (US-based candidates only)

Manager, Site Reliability Engineering (SRE)
Position Overview

We are seeking a Manager, Site Reliability Engineering (SRE) to lead our US-based SRE team and drive operational excellence across our production platforms.

This is a player-coach leadership role that combines people management with hands-on technical leadership. You will mentor and grow a team of Site Reliability Engineers while actively participating in major incident response, reliability initiatives, and operational reviews. The role is a key part of our global follow-the-sun support model and requires close collaboration with SRE leadership in India.
Key Responsibilities
Team Leadership & Development
  • Lead, mentor, and develop a US-based team of Site Reliability Engineers.
  • Conduct regular 1:1s, performance reviews, and career development discussions.
  • Own hiring, onboarding, and retention efforts as the team scales.
  • Foster a culture of ownership, blameless postmortems, and continuous improvement.
Operational Excellence & Incident Management
  • Lead day-to-day production operations and ensure timely incident triage, resolution, and escalation.
  • Serve as an escalation point and incident commander for major production incidents.
  • Drive problem management and root cause analysis processes.
  • Carry PagerDuty on-call escalation responsibilities for critical issues.
  • Track and report operational KPIs, SLAs, and SLOs, including availability, MTTR, and incident trends.
Reliability & Automation
  • Improve system reliability, observability, and resilience using Datadog and related tooling.
  • Drive automation, self-healing capabilities, and runbook maturity.
  • Partner with Development, DevOps, DevSecOps, and Engineering teams to embed reliability into the SDLC.
  • Contribute hands-on to tooling, automation, and technical reviews as needed.
Collaboration & Global Alignment
  • Coordinate closely with SRE leadership in India to ensure seamless follow-the-sun coverage.
  • Represent the US SRE organization in cross-functional planning and operational reviews.
  • Communicate effectively with both technical and non-technical stakeholders.
Documentation & Compliance
  • Maintain high-quality documentation for incidents, postmortems, runbooks, and operational procedures.
  • Ensure adherence to healthcare and fintech compliance standards, including HIPAA, PCI DSS, SOC 2, ISO 27001, and HITRUST.
Required Qualifications
  • 5-8 years of experience in Site Reliability Engineering, DevOps, Production Support, or Platform Engineering.
  • 1-2+ years of experience leading, mentoring, or managing engineers.
  • Demonstrated success operating in a player-coach leadership model.
  • Strong hands-on experience with production incident management and escalation processes.
  • Proficiency with Datadog or similar observability platforms.
  • Hands-on experience with Kubernetes and Docker in production environments.
  • Strong scripting or programming skills in PowerShell, Bash, Python, Java, or C#.
  • Experience with Helm, CI/CD pipelines, and deployment automation.
  • Working knowledge of ITIL processes and Agile methodologies.
  • Experience working with SQL, MySQL, or NoSQL databases.
  • Excellent communication and stakeholder management skills.
  • Willingness to participate in PagerDuty on-call escalation and work within a global follow-the-sun operating model.
Preferred Qualifications
  • Experience with cloud platforms such as AWS, Azure, or GCP.
  • Experience building or scaling SRE teams and on-call programs.
  • Experience defining and managing SLOs, SLIs, and error budgets.
  • Prior experience in the healthcare or fintech industry.
  • Knowledge of security and compliance frameworks relevant to regulated environments.
Why Join NationsBenefits?
  • Competitive compensation and comprehensive benefits.
  • Unlimited PTO.
  • Fully remote work environment (US-based).
  • Opportunity to lead and grow a high-impact SRE organization.
  • Exposure to modern cloud-native technologies and large-scale reliability challenges.
  • Collaborative culture focused on innovation, learning, and continuous improvement.
  • Meaningful work that directly impacts healthcare technology and millions of members.
Ideal Candidate

We are looking for a technically strong SRE leader who enjoys building teams, improving operational maturity, and remaining hands-on during critical production events. The ideal candidate combines leadership, systems thinking, and automation expertise to help scale reliability practices across a fast-growing Healthcare FinTech organization.

NationsBenefits is an Equal Opportunity Employer.
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineering Manager in United States vacancy
  •  ...Hands-on and detail-oriented, the full-time salaried Site Reliability Engineering Manager will lead a blended team of SRE and DevOps engineers in a remote setting, focusing on improving the availability and performance of Delinea's production environments while managing... 
    Suggested
    Full time
    For contractors
    Remote work

    Virtual Vocations Inc

    United States
    1 day ago
  • $135k - $140k

     ...Protein. Real People. Real Results. THAT is Link Snacks. Job Description JOB DESCRIPTION SUMMARY   The Site Reliability & Engineering Manager is the senior technical leader for the Minong facility, responsible for maintenance, reliability, facilities, utilities... 
    Suggested
    Permanent employment
    For contractors
    Work at office
    Relocation package
    Shift work

    Jack Link's Protein Snacks

    Minong, WI
    11 days ago
  •  ...Google Cloud in San Francisco, CA seeks a Manager, Software Engineer for Site Reliability Engineering to lead a team responsible for uptime, availability, and reliability at scale. This role blends hands-on software engineering with people leadership and strategic roadmapping... 
    Suggested

    Jobleads-US

    Kentucky
    2 days ago
  • $204k - $306k

     ...We're all in on this mission. If you are too, let's talk.Manager, Site Reliability EngineeringSan Francisco, CaliforniaSecure Every Identity,...  ...week in our San Francisco Office. The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes and provisions... 
    Suggested
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours
    2 days per week

    Okta

    Chicago, IL
    2 days ago
  • $84.9k - $209.5k

     ...strategic architecture with practical systems engineering, deployment, automation, patching,...  ...and compliance support. The Principal Site Reliability Engineer will work across Windows,...  ...improve deployment procedures. Experience managing complex or high-risk production... 
    Suggested
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    4 days ago
  • $139.7k - $232.9k

     ..., and continuously improving highly reliable, scalable, and resilient platform solutions...  ...as a subject matter expert (SME) in Site Reliability Engineering, driving reliability engineering...  ...requirements. • Lead incident management practices, including detection, response... 
    Full time
    Work experience placement

    M&T Bank

    Buffalo, NY
    2 hours ago
  •  .... Connecting. Growing together.We are seeking a Principal Site Reliability Engineer (SRE) to define and scale reliability practices across large...  ...in:Reliability engineering (SLOs, SLIs, incident management, observability)Distributed systems in cloud environments (... 
    Minimum wage
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work

    UnitedHealth Group

    Minnetonka, MN
    4 days ago
  • $102k - $234.6k

     ...architecting infrastructure and service for reliability and functionality. Provides day-to-day...  ...technology, execute improvements, build site reliability knowledge, and provide clear...  ...ResponsibilitiesCapacity Ingestion and Management:- Supports team members designing and... 
    Temporary work
    Immediate start
    Flexible hours

    Oracle Corporation

    Reston, VA
    3 days ago
  • $134.6k - $230.8k

     ...Connecting. Growing together.Are you passionate about reimagining operations through AI? Optum Financial is seeking a Principal Site Reliability Engineer to lead the evolution of our reliability platform by combining modern SRE practices with AI-assisted operations. You'll... 
    Minimum wage
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work

    UnitedHealth Group

    Eden Prairie, MN
    3 days ago
  • $248k - $396.75k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing, building, and operating large-scale production...  ...networking, cloud infrastructure, Kubernetes, databases, capacity management, continuous delivery, and observability. SRE at NVIDIA... 
    Full time

    Nvidia

    Santa Clara, CA
    2 hours ago
  •  ...Service (S3), and Auto Scaling Groups for dynamic resource management. Designs, develops, and executes performance tests using...  ...Delivery (CI/CD) pipelines and Kubernetes. Supports Site Reliability Engineering (SRE) functions by establishing Service Level Objectives... 
    Full time

    Fidelity Investments

    Durham, NC
    2 hours ago
  • $207k - $284.9k

     ...all in on this mission. If you are too, let's talk.Senior Manager, Site Reliability EngineeringSecure Every Identity, from AI to HumanIdentity...  ...mission. If you are too, let's talk.The Federal Operations Engineering GroupOkta's Federal Operations team supports government... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours
    Day shift

    Okta

    Washington DC
    4 days ago
  • $253k - $336k

     ...TEAM:CorpTech Platform is the internal engineering force multiplier behind Anduril's...  ...products.ABOUT THE JOB:The Director of Site Reliability Engineering owns the reliability system...  ...including hiring, coaching, performance management, succession planning, and development... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    3 days ago
  •  ...Infrastructure Code. Builds reliability into the ecosystem by applying...  ...practices in resiliency engineering and observability by developing...  ...engineering techniques with site reliability engineering...  ...business processes.Advises senior management on technical strategy and... 
    Full time

    Fidelity Investments

    Westlake, OH
    4 days ago
  •  ...thousands of companies. Join us as we help people all over the world thrive at work.Location: Salt Lake City, UTAs the Manager of Site Reliability Engineering, you will lead the strategy, execution, and evolution of reliability for our world-class employee recognition... 
    Full time
    Shift work

    O.C. Tanner

    Salt Lake City, UT
    2 hours ago
  • $182k - $250.8k

     ...at Okta is the backbone of our platform's reliability and operational excellence. We are a forward-thinking group of engineers and leaders who believe that great infrastructure...  ...for millions of users worldwide. As a Manager, Site Reliability Engineer, you'll lead this... 
    Permanent employment
    Local area
    Remote work
    Worldwide
    Flexible hours
    Weekend work
    Weekday work

    Okta

    Washington DC
    3 days ago
  • $142.8k - $274.8k

     ...yearEmployment type: Full-TimeWork site: 0 days / week in-office -...  ...EngineeringDiscipline: Site Reliability EngineeringCompany:...  ...a Principal Site Reliability Engineer, you will set technical and operational...  ...eligibility requirements.For manager-level roles, a Tier 5 (T5)... 
    Ongoing contract
    Work at office
    Local area

    Microsoft

    Redmond, WA
    1 day ago
  • $222k - $300.5k

     ...OverviewAbout the TeamIntuit's Infrastructure and Site Reliability organization owns the operational...  .... The Fintech Platform Systems Engineering team builds and operates the AWS-based...  ....The OpportunityWe're hiring a Senior Manager, Site Reliability Engineering to lead... 
    Worldwide
    Shift work

    Intuit

    Mountain View, CA
    2 hours ago
  • $121.5k - $264.1k

    Capacity Ingestion and Management:- Supports team members designing...  ...on practices and terms for reliability and functionality.- Supervises...  ...maintaining knowledge of site reliability trends and sharing...  ...years of experience in software engineering, infrastructure management,... 
    Temporary work
    Immediate start
    Flexible hours

    Oracle Corporation

    Reston, VA
    2 hours ago
  • $84.9k - $209.5k

     ...partner with customer support, service owners, and engineering teams around the globe to ensure high-quality...  ...Level - IC4Escalation points for junior site reliability engineers during complex or high-impact incidents.Manage and execute complex manual Change Management tickets... 
    Temporary work
    Monday to Friday
    Flexible hours
    Shift work
    Night shift

    Oracle Corporation

    Reston, VA
    3 days ago
  • $139.7k - $232.9k

    Manager, Site Reliability Engineering 62 M Overview Responsible for leading the Site Reliability Engineering Center of Excellence and the Forward Deployed SRE program supporting critical banking platforms, applications, and technology services. Manages an organization of... 
    Full time
    Work experience placement

    M&T Bank

    Buffalo, NY
    2 days ago
  •  ...perfect environment to have a major impact. As a Principal Site Reliability Engineer at JPMorgan Chase within the Corporate Technology Team,...  ...identify new opportunities to influence critical incident management and improve the end-to-end lifecycle of software development... 

    JP Morgan Chase

    Jersey City, NJ
    3 days ago
  • Job DescriptionJob Summary:The Director, Site Reliability Engineering (SRE) is a senior technical and people leader responsible for ensuring the...  ..., error budgets, observability, automation, and incident management excellence, while building a strong culture of ownership,... 
    Hourly pay
    Temporary work
    Work experience placement

    Walgreens Boots Alliance

    Deerfield, IL
    2 hours ago
  • $195k - $275k

     ...providing a wide range of investment banking, securities, investment management and wealth management services. The Firm's employees serve...  ...& Release Management, and the Chief Operating Office.The Reliability Operations (RO) within WMT is responsible for providing swift... 
    Temporary work
    Work at office
    Worldwide
    Night shift

    Morgan Stanley

    New York, NY
    2 hours ago
  • $159k - $272k

     ...people thrive in an evolving world. As a premier global asset management organization with more than 85 years of experience, we...  ...that matter to you. Role SummaryIn this role as Principal Site Reliability Engineer, Infrastructure Observability you will help formulate, develop... 
    Full time
    Private practice
    Local area
    Remote work
    Work from home
    3 days per week

    T. Rowe Price

    Owings Mills, MD
    4 days ago
  • $96.3k - $264.1k

     ...service, ensuring alignment with reliability and functionality standards....  ...and provides expertise in site reliability trends.Only...  ...drive the site reliability engineering strategy for large-scale, distributed...  ...deployment, configuration management, infrastructure provisioning... 
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    4 days ago
  • $262k - $364k

    Lead a team of software/systems engineers on projects for users and be directly responsible...  ...through quality technical execution.Manage on-call rotations across continents, using...  ...with machine learning infrastructure.Site Reliability Engineering (SRE) combines software and... 

    Google

    San Bruno, CA
    2 days ago
  • $207k - $300k

    Lead a team of Software/Systems Engineers on projects for users and be directly responsible...  ...through quality technical execution.Manage on-call rotations across continents, using...  ...year of people management experience. Site Reliability Engineering (SRE) combines software and... 

    Google

    Mountain View, CA
    1 day ago
  • $207k - $300k

    Manage a team of Software/Systems Engineers on projects for users and remain directly responsible for uptime.Own...  ...and establishing sustainable multi-site on-call rotations across...  ....Deep practical expertise in Site Reliability Engineering practices, including SLO... 

    Google

    San Jose, CA
    3 days ago
  • $122k - $207k

     ...services that help people, businesses and governments realize their greatest potential. Title and Summary Manager, Site Reliability Engineering Who is Mastercard? At Mastercard technology, we work to connect and power an inclusive, digital economy that... 
    Full time
    Part time
    Worldwide
    Flexible hours

    Mastercard

    O Fallon, MO
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineering Manager. Be the first to apply!