Site Reliability Engineering Manager
NationsBenefits, LLC
NationsBenefits is recognized as one of the fastest-growing companies in America and a Healthcare Fintech provider of supplemental benefits, flex cards, and member engagement solutions. We partner with managed care organizations to provide innovative healthcare solutions that drive growth, improve outcomes, reduce costs, and bring value to their members.Through our comprehensive suite of innovative supplemental benefits, fintech payment platforms, and member engagement solutions, we help health plans deliver high-quality benefits to their members that address the social determinants of health and improve member health outcomes and satisfaction.Our compliance-focused infrastructure, proprietary technology systems, and premier service delivery model allow our health plan partners to deliver high-quality, value-based care to millions of members.We offer a fulfilling work environment that attracts top talent and encourages all associates to contribute to delivering premier service to internal and external customers alike. Our goal is to transform the healthcare industry for the better! We provide career advancement opportunities from within the organization across multiple locations in the US, South America, and India.
Location: Remote (US-based candidates only) Manager, Site Reliability Engineering (SRE)Position Overview
We are seeking a Manager, Site Reliability Engineering (SRE) to lead our US-based SRE team and drive operational excellence across our production platforms.
This is a player-coach leadership role that combines people management with hands-on technical leadership. You will mentor and grow a team of Site Reliability Engineers while actively participating in major incident response, reliability initiatives, and operational reviews. The role is a key part of our global follow-the-sun support model and requires close collaboration with SRE leadership in India.
Key Responsibilities
Team Leadership & Development
- Lead, mentor, and develop a US-based team of Site Reliability Engineers.
- Conduct regular 1:1s, performance reviews, and career development discussions.
- Own hiring, onboarding, and retention efforts as the team scales.
- Foster a culture of ownership, blameless postmortems, and continuous improvement.
Operational Excellence & Incident Management
- Lead day-to-day production operations and ensure timely incident triage, resolution, and escalation.
- Serve as an escalation point and incident commander for major production incidents.
- Drive problem management and root cause analysis processes.
- Carry PagerDuty on-call escalation responsibilities for critical issues.
- Track and report operational KPIs, SLAs, and SLOs, including availability, MTTR, and incident trends.
Reliability & Automation
- Improve system reliability, observability, and resilience using Datadog and related tooling.
- Drive automation, self-healing capabilities, and runbook maturity.
- Partner with Development, DevOps, DevSecOps, and Engineering teams to embed reliability into the SDLC.
- Contribute hands-on to tooling, automation, and technical reviews as needed.
Collaboration & Global Alignment
- Coordinate closely with SRE leadership in India to ensure seamless follow-the-sun coverage.
- Represent the US SRE organization in cross-functional planning and operational reviews.
- Communicate effectively with both technical and non-technical stakeholders.
Documentation & Compliance
- Maintain high-quality documentation for incidents, postmortems, runbooks, and operational procedures.
- Ensure adherence to healthcare and fintech compliance standards, including HIPAA, PCI DSS, SOC 2, ISO 27001, and HITRUST.
Required Qualifications
- 5–8 years of experience in Site Reliability Engineering, DevOps, Production Support, or Platform Engineering.
- 1–2+ years of experience leading, mentoring, or managing engineers.
- Demonstrated success operating in a player-coach leadership model.
- Strong hands-on experience with production incident management and escalation processes.
- Proficiency with Datadog or similar observability platforms.
- Hands-on experience with Kubernetes and Docker in production environments.
- Strong scripting or programming skills in PowerShell, Bash, Python, Java, or C#.
- Experience with Helm, CI/CD pipelines, and deployment automation.
- Working knowledge of ITIL processes and Agile methodologies.
- Experience working with SQL, MySQL, or NoSQL databases.
- Excellent communication and stakeholder management skills.
- Willingness to participate in PagerDuty on-call escalation and work within a global follow-the-sun operating model.
Preferred Qualifications
- Experience with cloud platforms such as AWS, Azure, or GCP.
- Experience building or scaling SRE teams and on-call programs.
- Experience defining and managing SLOs, SLIs, and error budgets.
- Prior experience in the healthcare or fintech industry.
- Knowledge of security and compliance frameworks relevant to regulated environments.
Why Join NationsBenefits?
- Competitive compensation and comprehensive benefits.
- Unlimited PTO.
- Fully remote work environment (US-based).
- Opportunity to lead and grow a high-impact SRE organization.
- Exposure to modern cloud-native technologies and large-scale reliability challenges.
- Collaborative culture focused on innovation, learning, and continuous improvement.
- Meaningful work that directly impacts healthcare technology and millions of members.
Ideal Candidate
We are looking for a technically strong SRE leader who enjoys building teams, improving operational maturity, and remaining hands-on during critical production events. The ideal candidate combines leadership, systems thinking, and automation expertise to help scale reliability practices across a fast-growing Healthcare FinTech organization.
NationsBenefits is an Equal Opportunity Employer.- ...Role Overview Help us ensure the reliability of Ajaib's fintech platform, serving millions of Indonesian investors. You'll lead... ...We're Looking For - 5+ years in SRE/DevOps, with 2+ years managing engineers - Deep hands-on expertise in GCP and Kubernetes -...SuggestedRemote work
- ...Site Reliability Engineering Manager Canonical is a leading provider of open-source software and operating systems for global enterprise and technology markets. Our platform, Ubuntu, is very widely used in breakthrough enterprise initiatives such as public cloud, data...SuggestedWork at officeLocal areaRemote workWork from homeWorldwide
$155k - $180k
...Site Reliability Engineering Manager (Remote) Join to apply for the Site Reliability Engineering Manager (Remote) role at WebstaurantStore Base pay range: $155,000.00/yr - $180,000.00/yr Job Summary As the largest online distributor of restaurant supplies...SuggestedH1bRemote workHome office$132.77k - $221.35k
...for our SRE function. In this role, you will mature the Site Reliability Engineering (SRE) function ensuring reliability and performance of organization... ...Inc. (Nasdaq: LPLA) is among the fastest growing wealth management firms in the U.S. As a leader in the financial advisor-...SuggestedFull timeWork from home$142.8k - $274.8k
...yearEmployment type: Full-TimeWork site: 0 days / week in-office -... ...EngineeringDiscipline: Site Reliability EngineeringCompany:... ...a Principal Site Reliability Engineer, you will set technical and operational... ...eligibility requirements.For manager-level roles, a Tier 5 (T5)...SuggestedOngoing contractWork at officeLocal area- ...sponsorship.Maintain and enhance the reliability, availability, and... ...page of the Navy Federal Career Site.Protect Yourself from Job Scams... ...degree in computer science, engineering, or the equivalent... ...knowledge of incident response and managing production issues. Advanced communication...InternshipMonday to Friday
- .... Connecting. Growing together.We are seeking a Principal Site Reliability Engineer (SRE) to define and scale reliability practices across large... ...in:Reliability engineering (SLOs, SLIs, incident management, observability)Distributed systems in cloud environments (...Minimum wageFull timeWork experience placementWork at officeLocal areaRemote work
$163.62k - $212.71k
...platforms, and processes that improve our engineering teams' productivity and streamline the... ...seasoned and strategic Lead/Principal Site Reliability Engineer to drive the reliability,... ...Operations (SRE Focus)Platform Design and Management: Architect, build, and maintain...Full timePart timeWork experience placementWork at officeLocal areaImmediate startRemote workWork from homeFlexible hoursShift work3 days per week1 day per week$159k - $272k
...people thrive in an evolving world. As a premier global asset management organization with more than 85 years of experience, we... ...that matter to you. Role SummaryIn this role as Principal Site Reliability Engineer, Infrastructure Observability you will help formulate, develop...Full timePrivate practiceLocal areaRemote workWork from home3 days per week- ...Principal Site Reliability Engineer location- Washington, DC -Onsite Remote- No 6+ Months Job Summary At Amtrak,... ...or similar. Infrastructure Automation: Implement and manage IaC using tools like Terraform, AWS CloudFormation, or...Remote work
$90k - $130k
...-on experience. We require 10+ years of experience in Site Reliability Engineering, Software Engineering, or Cloud Engineering. We need experience... ...reliability engineering, including SLOs, SLIs, incident management, and observability. We require experience working with...Full timeRemote work$7,000 per month
...Principal Site Reliability Engineer Latin America The salary range for this role is negotiable, the range being $7000 - $12000 per month (Gross in USD) About Sezzle: With a mission to financially empower the next generation, Sezzle is revolutionizing the shopping...Remote workFlexible hours- ...Working remotely within the United States, the full-time Principal Site Reliability Engineer will lead project work to enhance platform reliability, mentor junior engineers, and engage in incident response while collaborating closely with product stakeholders and architects...Full timeRemote work
$175k - $220k
...across the U.S., Canada, and India. The Director, Site Reliability Engineering (SRE) will lead reliability, performance, and... ...CD practices for assigned product families. Directors will manage multiple teams and collaborate with Product Development, Architecture...Contract workTemporary workWork at officeWork from homeFlexible hours- ...Site Reliability Engineering (SRE) Team Lead The Site Reliability Engineering (SRE) team is foundational to the growth and scale of our platform... ...pipelines, using CI/CD systems and configuration management Why You'll Like It Here We are collaborative at...Shift work
$200k - $250k
...progress, come build the future together. As a Principal Site Reliability Engineer, you'll shape the long-term strategy for the... ...across critical infrastructure, including cluster lifecycle management, networking, identity and access management, observability...Full timeImmediate startRemote work- ...Senior Principal Site Reliability Engineer Hong Kong SAR About Us Established in 2018, Bybit is one of the world's leading cryptocurrency... ...a seamless ecosystem across trading, payments, wealth management, custody, institutional services, and Web3 — connecting users...Remote work
- ...Principal Site Reliability Engineer Deimos is a cloud-native developer and security operations technology services company. We help companies... ...projects. You will report to a Site Reliability Engineering Manager. As a Principal Site Reliability Engineer you will be...Currently hiringRemote workWork from home
- ...Senior Manager, Site Reliability Engineering Remote - USA At Counterpart Health, we are transforming healthcare and improving patient care with our innovative primary care tool, Counterpart Assistant. By supporting Primary Care Physicians (PCPs), we deliver improved...Work at officeRemote workFlexible hoursShift work
$167.3k - $242.6k
...Principal Site Reliability Engineer Remote Your passion for uptime was forged from experience in production and refined through incident... ...interview. We also have a goal that all Expletives have a great manager and have a voice in how their team is run and who runs it....Remote workVisa sponsorshipDay shift- ...software development teams to build reliable, scalable, secure, and cloud-... ...architecture patterns across engineering teams, helping ensure systems... ...of hands-on experience in Site Reliability Engineering,... ...observability, monitoring, and incident management tools such as Honeycomb.io,...Remote work
$230k - $255k
...Join Aya Healthcare, winner of multiple Top Workplace awards! We're looking for a highly experienced Manager, Site Reliability Engineering to lead the team behind one of healthcare's most relied-on workforce platforms. In this leadership role, you'll guide and...Local areaRemote work$160k - $180k
..., Canada, and India. We are seeking a Principal Site Reliability Engineer to define the strategic vision and own the enterprise-wide... ...: Establish the governance models for defining and managing SLIs and SLOs across multiple product lines. ~ Delivery...Contract workTemporary workWork at officeWork from homeFlexible hours$84.9k - $209.5k
...spirit that promotes an upbeat and creative environment. We are unencumbered and will need your contribution to make it a special engineering center with the focus on excellence. Health Data Intelligence Platform has a rare opportunity to play a critical role in how...Temporary workImmediate startRemote workFlexible hours- ...trillion in crypto transactions. We are looking for a Head of Site Reliability Engineering (SRE) who will serve as the principal leader in... ...join an exceptional crypto company and senior engineering management team at a time of high growth. This is a tremendous opportunity...Full timeApprenticeshipWork at officeRemote workWorldwideFlexible hours
$189.59k - $220k
...Director, Site Reliability Engineering NBCUniversal is one of the world’s leading media and entertainment companies. We create world‑class content... ...for hands‑on configuration and support as well as managing the work of other architects and engineers. ~ Work closely...Full timeFor contractorsRemote work- Role Description Symmetrio is recruiting a Principal Site Reliability Engineer (SRE) for our customer, a rapidly growing healthcare technology... ...Qualifications ~6+ years of hands-on experience supporting and managing AWS-based production environments ~4+ years of...Full time
$151k - $297k
..., you will partner with SRE leaders and engineers to scale the platform that underpins all... ...program execution, strengthen production reliability practices, and coordinate cross-... ...criteria with SRE engineers and leaders. Manage dependencies across platform teams, keep...Local areaRemote workWorldwideFlexible hours- ...Information Technology group delivers secure, reliable technology solutions that enable DTCC... ...enterprise platforms.As a Principal Site Reliability Engineer (SRE), you will drive operational... ...related monitoring platforms. Define and manage SLIs, SLOs, dashboards, alerts, and...Remote workFlexible hours
$132.77k - $221.35k
...teams to ensure products are designed with observability, reliability, and performance in mind, promoting operational... ...with key stakeholders across the organization, including engineering, operations, and management. Create best‑in‑class reports and prepare presentations...Work from home
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineering Manager. Be the first to apply!



