Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineering Manager

NationsBenefits, LLC

NationsBenefits is recognized as one of the fastest-growing companies in America and a Healthcare Fintech provider of supplemental benefits, flex cards, and member engagement solutions. We partner with managed care organizations to provide innovative healthcare solutions that drive growth, improve outcomes, reduce costs, and bring value to their members.Through our comprehensive suite of innovative supplemental benefits, fintech payment platforms, and member engagement solutions, we help health plans deliver high-quality benefits to their members that address the social determinants of health and improve member health outcomes and satisfaction.Our compliance-focused infrastructure, proprietary technology systems, and premier service delivery model allow our health plan partners to deliver high-quality, value-based care to millions of members.We offer a fulfilling work environment that attracts top talent and encourages all associates to contribute to delivering premier service to internal and external customers alike. Our goal is to transform the healthcare industry for the better! We provide career advancement opportunities from within the organization across multiple locations in the US, South America, and India.

Location: Remote (US-based candidates only)

Manager, Site Reliability Engineering (SRE)
Position Overview

We are seeking a Manager, Site Reliability Engineering (SRE) to lead our US-based SRE team and drive operational excellence across our production platforms.

This is a player-coach leadership role that combines people management with hands-on technical leadership. You will mentor and grow a team of Site Reliability Engineers while actively participating in major incident response, reliability initiatives, and operational reviews. The role is a key part of our global follow-the-sun support model and requires close collaboration with SRE leadership in India.
Key Responsibilities
Team Leadership & Development
  • Lead, mentor, and develop a US-based team of Site Reliability Engineers.
  • Conduct regular 1:1s, performance reviews, and career development discussions.
  • Own hiring, onboarding, and retention efforts as the team scales.
  • Foster a culture of ownership, blameless postmortems, and continuous improvement.
Operational Excellence & Incident Management
  • Lead day-to-day production operations and ensure timely incident triage, resolution, and escalation.
  • Serve as an escalation point and incident commander for major production incidents.
  • Drive problem management and root cause analysis processes.
  • Carry PagerDuty on-call escalation responsibilities for critical issues.
  • Track and report operational KPIs, SLAs, and SLOs, including availability, MTTR, and incident trends.
Reliability & Automation
  • Improve system reliability, observability, and resilience using Datadog and related tooling.
  • Drive automation, self-healing capabilities, and runbook maturity.
  • Partner with Development, DevOps, DevSecOps, and Engineering teams to embed reliability into the SDLC.
  • Contribute hands-on to tooling, automation, and technical reviews as needed.
Collaboration & Global Alignment
  • Coordinate closely with SRE leadership in India to ensure seamless follow-the-sun coverage.
  • Represent the US SRE organization in cross-functional planning and operational reviews.
  • Communicate effectively with both technical and non-technical stakeholders.
Documentation & Compliance
  • Maintain high-quality documentation for incidents, postmortems, runbooks, and operational procedures.
  • Ensure adherence to healthcare and fintech compliance standards, including HIPAA, PCI DSS, SOC 2, ISO 27001, and HITRUST.
Required Qualifications
  • 5-8 years of experience in Site Reliability Engineering, DevOps, Production Support, or Platform Engineering.
  • 1-2+ years of experience leading, mentoring, or managing engineers.
  • Demonstrated success operating in a player-coach leadership model.
  • Strong hands-on experience with production incident management and escalation processes.
  • Proficiency with Datadog or similar observability platforms.
  • Hands-on experience with Kubernetes and Docker in production environments.
  • Strong scripting or programming skills in PowerShell, Bash, Python, Java, or C#.
  • Experience with Helm, CI/CD pipelines, and deployment automation.
  • Working knowledge of ITIL processes and Agile methodologies.
  • Experience working with SQL, MySQL, or NoSQL databases.
  • Excellent communication and stakeholder management skills.
  • Willingness to participate in PagerDuty on-call escalation and work within a global follow-the-sun operating model.
Preferred Qualifications
  • Experience with cloud platforms such as AWS, Azure, or GCP.
  • Experience building or scaling SRE teams and on-call programs.
  • Experience defining and managing SLOs, SLIs, and error budgets.
  • Prior experience in the healthcare or fintech industry.
  • Knowledge of security and compliance frameworks relevant to regulated environments.
Why Join NationsBenefits?
  • Competitive compensation and comprehensive benefits.
  • Unlimited PTO.
  • Fully remote work environment (US-based).
  • Opportunity to lead and grow a high-impact SRE organization.
  • Exposure to modern cloud-native technologies and large-scale reliability challenges.
  • Collaborative culture focused on innovation, learning, and continuous improvement.
  • Meaningful work that directly impacts healthcare technology and millions of members.
Ideal Candidate

We are looking for a technically strong SRE leader who enjoys building teams, improving operational maturity, and remaining hands-on during critical production events. The ideal candidate combines leadership, systems thinking, and automation expertise to help scale reliability practices across a fast-growing Healthcare FinTech organization.

NationsBenefits is an Equal Opportunity Employer.
Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineering Manager in United States vacancy
  • $135k - $140k

     ...Job Description JOB DESCRIPTION SUMMARY The Site Reliability & Engineering Manager is the senior technical leader for the Minong facility, responsible for maintenance, reliability, facilities, utilities, and site engineering support. This role provides strategic... 
    Suggested
    Permanent employment
    For contractors
    Work at office
    Relocation package
    Shift work

    Jack Link's Beef Jerky

    Minong, WI
    1 day ago
  •  ...Hands-on and detail-oriented, the full-time salaried Site Reliability Engineering Manager will lead a blended team of SRE and DevOps engineers in a remote setting, focusing on improving the availability and performance of Delinea's production environments while managing... 
    Suggested
    Full time
    For contractors
    Remote work

    Virtual Vocations Inc

    United States
    4 days ago
  •  ...enterprise initiatives such as public cloud, data science, AI, engineering innovation and IoT. Our customers include the world’s...  ...led, profitable and growing. We are hiring a Site Reliability Engineering Manager aspiring for a world-class devops and gitops engineering... 
    Suggested
    Full time
    Work at office
    Local area
    Remote work
    Work from home
    Worldwide

    Canonical

    Remote
    a month ago
  •  ...Google Cloud in San Francisco, CA seeks a Manager, Software Engineer for Site Reliability Engineering to lead a team responsible for uptime, availability, and reliability at scale. This role blends hands-on software engineering with people leadership and strategic roadmapping... 
    Suggested

    Jobleads-US

    Kentucky
    14 hours ago
  •  ...Infrastructure Code. Builds reliability into the ecosystem by applying...  ...practices in resiliency engineering and observability by developing...  ...engineering techniques with site reliability engineering...  ...processes. Advises senior management on technical strategy and tools... 
    Suggested

    Fidelity Investments

    Roanoke, TX
    1 day ago
  •  ...Principal Site Reliability Engineer The Principal Site Reliability Engineer will be a critical technical leader responsible for driving the...  ...Site Reliability Engineering (SRE), including defining SLOs, managing error budgets, and leading incident response. You will... 

    Software Technology Inc

    Washington DC
    4 days ago
  • $194k - $237k

    ## Principal Site Reliability EngineerApplylocations: Scottsdaletime type: Full timeposted on...  ...Purpose**The Principal Site Reliability Engineer partners with development teams by...  ...Supports the company’s commitment to risk management and protecting the integrity and... 
    Hourly pay
    Work at office
    Immediate start
    Visa sponsorship
    Work visa
    Flexible hours

    Early Warning Services

    Eastern, KY
    14 hours ago
  • $205k - $305k

     ...Director Of Site Reliability Engineering Interested in working on cutting-edge blockchain technology and creating equitable access to the global...  ...and compute patterns, CI/CD, observability, secrets management, GitHub workflows, and infrastructure automation. Help... 
    Temporary work
    Work at office
    Local area
    Worldwide
    Flexible hours

    Stellar

    San Francisco, CA
    4 days ago
  •  ...Principal Site Reliability Engineer As a Principal member of the Site Reliability Engineering (SRE) team, you'll take ownership of highly available systems, influence service design, and work across teams to drive resiliency, automation, and operational excellence.... 
    Work at office
    Flexible hours
    3 days per week

    Oracle

    North Bloomfield, OH
    2 days ago
  • $248k - $396.75k

     ...US, CA, Santa Clara Full time JR2023973 Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing...  ...networking, cloud infrastructure, Kubernetes, databases, capacity management, continuous delivery, and observability. SRE at NVIDIA... 
    Full time

    NVIDIA

    Santa Clara, CA
    3 days ago
  •  ...About the job Director of Site Reliability Engineering About Stellar Stellar is a decentralised, public blockchain that gives developers...  ...and compute patterns, CI/CD, observability, secrets management, GitHub workflows, and infrastructure automation - Help... 

    TechChain Talent

    San Francisco, CA
    4 days ago
  • $151.6k - $245.3k

     ...infrastructure and is one of the largest GCP customers. As a Site Reliability Engineer, you will be part of a team supporting the services...  ...equivalent military experience ~ Expertise in configuration management with a framework such as Ansible, Terraform, Helm, Kubernetes... 
    Full time
    Work at office

    Palo Alto Networks

    Santa Clara, CA
    3 days ago
  • $207k - $284.9k

     ...This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Senior Manager, Site Reliability Engineering Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours
    Day shift

    Okta

    Washington DC
    5 days ago
  • $160k - $180k

     ...S., Canada, and India. We are seeking a Principal Site Reliability Engineer to define the strategic vision and own the enterprise-wide...  ...: Establish the governance models for defining and managing SLIs and SLOs across multiple product lines. ~ Delivery... 
    Contract work
    Work from home
    Flexible hours

    Vertafore

    Denver, CO
    18 days ago
  • $175k - $220k

     ...offices across the U.S., Canada, and India. The Director, Site Reliability Engineering (SRE) will lead reliability, performance, and...  ...CD practices for assigned product families. Directors will manage multiple teams and collaborate with Product Development, Architecture... 
    Contract work

    Vertafore

    Denver, CO
    18 days ago
  •  ...Principal Site Reliability Engineer About ShipperHQ: ShipperHQ is a trusted leader in the e-commerce shipping space, with over 15 years...  ...reliability standards, SLOs, SLIs, error budgets, and incident management best practices. Lead the design and implementation of... 
    Full time
    Work at office

    ShipperHQ

    Austin, TX
    more than 2 months ago
  • $204k - $306k

     ...excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Manager, Site Reliability Engineering San Francisco, California Secure Every Identity, from AI to Human Identity is the key to unlocking the... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours
    2 days per week

    Okta

    Washington DC
    5 days ago
  •  ...I.T. is actively seeking a Principal Engineer for an immediate full-time opportunity...  ...technology company is seeking experienced Site Reliability Engineers to take ownership of...  ...including SLI/SLO frameworks and error budget management Establish escalation protocols and... 
    Permanent employment
    Full time
    Temporary work
    Immediate start

    KēSTA I.T.

    Beverly Hills, CA
    18 days ago
  • $182k - $250.8k

     ...at Okta is the backbone of our platform's reliability and operational excellence. We are a forward-thinking group of engineers and leaders who believe that great infrastructure...  ...for millions of users worldwide. As a Manager, Site Reliability Engineer, you'll lead this... 
    Permanent employment
    Local area
    Remote work
    Worldwide
    Flexible hours
    Weekend work
    Weekday work

    Okta

    Washington DC
    5 days ago
  •  ...EIT) organization is expanding, and we are seeking a Senior Site Reliability Engineer to help drive a major architectural modernization. In this...  ...: • Everything as Code: Drive repository-led management across our public and private cloud environments to establish... 
    Permanent employment
    Full time
    H1b
    Local area
    Remote work
    Shift work

    Jack Henry & Associates

    New York, NY
    14 hours ago
  • $153k - $210k

     ...Senior Software Engineer, Site Reliability Engineering Reno, NV; San Ramon, CA; NYC - Hybrid Are you passionate about building resilient...  ...(SLOs), and error budget practices to proactively manage reliability. Identify capacity constraints and reliability... 
    Full time

    Ridgeline

    New York, NY
    14 hours ago
  • $195k - $275k

     ...providing a wide range of investment banking, securities, investment management and wealth management services. The Firm's employees serve...  ...Release Management, and the Chief Operating Office. The Reliability Operations (RO) within WMT is responsible for providing swift... 
    Full time
    Temporary work
    Work at office
    Worldwide
    Night shift

    Morgan Stanley

    New York, NY
    a month ago
  •  ...infrastructure and service for reliability and functionality. Provides...  ...execute improvements, build site reliability knowledge, and...  ...alignment with performance management processes, guidelines, and expectations...  ...of experience in software engineering, infrastructure management,... 

    Hackajob

    Reston, VA
    3 days ago
  • $140k - $230k

     ...Zoox is seeking a Site Reliability Engineer to help ensure the availability, performance, and resilience of the services that power the development...  ...or a similar role, with a strong, objective background in managing large-scale distributed systems. Cloud & Infrastructure... 
    Full time

    Zoox

    Remote
    14 hours ago
  •  ...Google Cloud is seeking a Manager, Software Engineer, Site Reliability Engineering in Sunnyvale, CA. You will lead a team focused on uptime, reliability, and scalable infrastructure, delivering automated solutions and architecting resilient systems. The role requires... 

    Jobleads-US

    Sunnyvale, CA
    14 hours ago
  •  ...only provider of enterprise-scale context engines capable of analyzing trillions of real-...  ...seeking a highly skilled and motivated Site Reliability Engineer (SRE) to join our growing team...  ...deployment, monitoring, and incident management to continuously improve overall system... 
    Full time

    Lovelace Ai

    Pittsburgh, PA
    14 hours ago
  • $213.1k - $300k

     ...Google Houston, TX, USA is seeking a Manager, Software Engineer in Site Reliability Engineering to lead a team of engineers focused on uptime, availability, and scalable infrastructure across global services. The role emphasizes ownership and decision making, with... 

    Jobleads-US

    Houston, TX
    1 day ago
  •  ...Sophos seeks an experienced Manager, Software Engineering (SRE) to lead a distributed team across the U.S. and Canada, focusing on reliability, scalability, and efficient cloud operations. You will guide AWS, Kubernetes/EKS, Terraform/IaC, and automation efforts while... 
    Remote job

    Jobleads-US

    Kentucky
    2 days ago
  • $151k - $297k

     ..., you will partner with SRE leaders and engineers to scale the platform that underpins all...  ...program execution, strengthen production reliability practices, and coordinate cross-...  ...criteria with SRE engineers and leaders. Manage dependencies across platform teams, keep... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Miami, FL
    4 days ago
  •  ...part of a company that values diversity, integrity, and growth. Role Overview PDI Technologies is looking for a Manager, Site Reliability Engineering to lead the SRE organization supporting Paylo, PDI’s payments, loyalty, and fuel-pricing product suite. This role... 

    PDI Technologies

    Dallas, TX
    17 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineering Manager. Be the first to apply!