Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Manager, Site Reliability Engineering

$230k - $255k

Aya Healthcare

Join Aya Healthcare, winner of multiple Top Workplace awards!


We're looking for a highly experienced Manager, Site Reliability Engineering to lead the team behind one of healthcare's most relied-on workforce platforms. In this leadership role, you'll guide and grow a team of engineers driving product and platform reliability - ensuring an exceptional experience for the clinicians, clients, and internal teams who depend on us every day. You'll shape our reliability architecture, lead complex operational initiatives, and drive the adoption of AI-native operations (AIOps) and automation to eliminate toil and advance performance - owning measurable business outcomes across uptime, customer trust, and platform efficiency, and leading with the radical ownership Aya expects of every leader.

Who We Are:

We're a $8+ billion, rapidly growing workforce solutions provider in the healthcare industry. We deliver tech-enabled services that help healthcare organizations meet and manage their contingent labor needs. We build and manage tech-enabled marketplaces for national and local healthcare talent and deliver contingent labor management solutions through our proprietary software platform.


At Aya, we're obsessed with creating exceptional experiences for our clients, clinicians, and employees. In fact, we put employee satisfaction above all else. Our team members are responsible for incomparable customer experience and we know that happy employees are critical to maintaining happy clients. We foster an entrepreneurial, high-energy, low-bureaucracy culture and value innovative thinking and creative problem-solving. We embrace diversity in thought and backgrounds unified by a commitment to high achievement. When you join Aya, you'll be surrounded by teammates who care about you as an individual and leaders who will help you grow both personally and professionally.


Responsibilities:
  • Lead and grow the SRE team
    • Lead, mentor, and grow a team of high-performing Site Reliability Engineers across hiring, performance management, career development, and on-call rotation health.
    • Set the operating cadence for the team - standups, incident reviews, SLO/error-budget reviews, post-incident learning, and capacity planning.
    • Build a culture of blameless learning, technical depth, customer empathy, and disciplined ownership.
    • Partner closely with DevSecOps, Security Engineering, DRE, Incident & Change Management, and product engineering leadership to remove cross-team friction.
  • Drive reliability, performance, and availability
    • Own the reliability strategy for customer-facing products and internal platforms - defining SLOs, SLIs, and error budgets in partnership with product and engineering leadership, and operationalizing them in the release process.
    • Lead major incident response as senior incident commander for severity-1 events; institutionalize blameless post-incident reviews and ensure systemic fixes ship.
    • Champion proactive reliability - chaos engineering, game days, failure-mode analysis, capacity and load testing - well before incidents force the conversation.
    • Manage software release support and 24/7 on-call escalation rotations across the platform surface area, with humane on-call load and clear escalation paths.
  • Operational intelligence and AI-native operations
    • Build the AIOps practice - anomaly detection, predictive alerting, intelligent correlation, and automated triage - to drive measurable reductions in MTTD and MTTR.
    • Operationalize AI-assisted workflows for incident summarization, runbook generation, log and trace analysis, change risk scoring, and post-incident narrative drafting.
    • Pilot and scale agentic remediation where appropriate, with strict guardrails, audit trails, and human-in-the-loop controls suitable for a HIPAA-regulated environment.
    • Evolve the observability platform (Datadog metrics, logs, traces, RUM, synthetics, CI Visibility) so engineering teams can operate their own services with confidence and clear ownership.
  • Platform efficiency and stakeholder trust
    • Treat reliability as a product with a roadmap, measurable outcomes, and an executive-credible narrative - not as overhead.
    • Drive platform unit economics by partnering with FinOps and platform leadership on cost-to-serve, right-sizing, capacity efficiency, and waste elimination.
    • Communicate outcomes to executive, product, and customer-facing stakeholders in plain language tied to clinician and client experience.
    • Uphold HIPAA, PHI, and security obligations across every reliability decision, change, and tool selection.
Required Qualifications:
  • 10+ years in a combination of Site Reliability Engineering, DevOps, Platform Engineering, or related production-operations roles.
  • 4+ years of direct people management experience - hiring, performance management, career development, and running remote on-call teams.
  • Demonstrated ownership of reliability outcomes for customer-facing SaaS at meaningful scale - defining and operationalizing SLOs/SLIs/error budgets and using them to drive engineering prioritization.
  • Deep Azure experience - 3+ years operating production workloads on Azure, with hands-on depth in AKS, networking, identity, and platform services. Equivalent depth in AWS or GCP will be considered.
  • Modern observability fluency - production-grade experience with Datadog (or equivalent: New Relic, Dynatrace, AppDynamics) across metrics, logs, traces, RUM, and synthetics.
  • AI in operations - hands-on experience integrating AI/LLM-assisted tooling into operational workflows (incident summarization, runbook generation, log analysis, anomaly triage, change risk scoring).
  • Incident command experience - proven ability to lead severity-1 incidents end-to-end, run blameless reviews, and convert lessons into systemic improvements.
  • Regulated-environment instinct - operates with HIPAA, PHI, SOC 2, or comparable compliance constraints as a default mindset, not an afterthought.
  • Executive-grade communication - translates reliability work into business outcomes for executive, product, and customer-facing audiences.
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field - or an equivalent combination of education, training, and experience.
Preferred Qualifications:
  • Cloudflare at the edge - production experience with Cloudflare CDN, WAF, Workers, Access (ZTNA), Tunnel, Turnstile, and certificate management.
  • IaC at scale - Terragrunt and Terraform in a multi-environment, policy-gated pipeline; experience evolving IaC from "it works" to "it scales safely."
  • CI/CD maturity - GitHub Actions with OIDC/workload identity federation, OPA/Conftest policy-as-code, progressive delivery, and DORA-metric instrumentation.
  • Container platform depth - Kubernetes/AKS in production, including Helm, ingress, service mesh, autoscaling, and node lifecycle.
  • ITSM integration - ServiceNow for change, incident, and problem management; experience tying observability and CI data into ITSM workflows.
  • Identity ecosystem - operating in an Okta / Entra ID / M365 identity environment, including PIM, conditional access, and service-principal hygiene.
  • Chaos and resilience engineering - running game days, fault injection, and resilience exercises as a routine practice.
  • FinOps fluency - cost-to-serve, right-sizing, capacity efficiency, and unit-economics work in cloud environments.
  • Agile delivery - Scrum/Kanban delivery with Jira; comfortable operating in a quarterly planning + continuous-delivery cadence.
What We Offer:
  • Free premium medical, dental, life and vision insurance
  • Generous 401(k) match
  • Aya also offers other benefits to those that are eligible and where required by applicable law, including reimbursements and discretionary bonuses
  • Aya provides paid sick leave in accordance with all applicable state, federal, and local laws. Aya's general sick leave policy is that employees accrue one hour of paid sick leave for every 30 hours worked. However, to the extent any provisions of the statement above conflict with any applicable paid sick leave laws, the applicable paid sick leave laws are controlling
  • Celebrations! We hit our goals and reward ourselves.
  • Company-sponsored virtual events, happy hours and team-building activities are always on the horizon - plus, you get a special treat on your birthday!
  • Unlimited DTO - we believe in time off!
  • Virtual yoga, meditation or boot camp classes offered daily
Compensation : Aya reasonably anticipates the pay scale for this position to be an annual salary of $230,000 to $255,000.


The pay scale for this position may vary if applicant possesses experience outside of what Aya reasonably anticipates for this position. Bonuses are subject to the role and your manager's discretion.


Aya is an Equal Opportunity Employer (EEO), including Disability / Vets, and welcomes all to apply. Please click here for our EEO policy
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Manager, Site Reliability Engineering in United States vacancy
  • $160k - $185k

     ...their fitness journey and revolutionized the industry along the way. And we’re just getting started!OverviewThe Sr. Manager, Site Reliability Engineering (SRE) leads the strategy, execution, and continuous improvement of reliability, availability, and performance across... 
    Suggested
    Work at office
    Local area
    Remote work
    Work from home

    Planet Fitness

    Hampton, NH
    17 hours ago
  • $112.7k - $193.2k

     ...next breakthrough? Join us to start Caring. Connecting. Growing together.We are seeking an experienced Senior Manager to lead enterprise Site Reliability Engineering (SRE), DevOps, IT Service Management (ITSM), and Operational Excellence initiatives across Optum bank. This... 
    Suggested
    Minimum wage
    Full time
    Work experience placement
    Local area
    Remote work

    UnitedHealth Group

    Basking Ridge, NJ
    17 hours ago
  • $151k - $297k

     ..., you will partner with SRE leaders and engineers to scale the platform that underpins all...  ...program execution, strengthen production reliability practices, and coordinate cross-...  ...criteria with SRE engineers and leaders. Manage dependencies across platform teams, keep... 
    Suggested
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    New York, NY
    17 hours ago
  • Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally...  ...yourself among the top echelon in site reliability.As a Senior Lead Site Reliability Engineer...  ...Experience with AWS platforms and managed data platforms such as Databricks, including... 
    Suggested
    Work at office

    JP Morgan Chase

    Jersey City, NJ
    17 hours ago
  • Manager - Production Operations & Site Reliability EngineeringAt Alcon, we are driven by the meaningful work we do to help people see brilliantly. We innovate...  ..., talented people to join Alcon. As a Principal Engineer you will provide technical leadership for the reliability... 
    Suggested
    Full time
    Temporary work

    Alcon Pharma

    Lake Forest, CA
    3 days ago
  •  ...Role Overview Help us ensure the reliability of Ajaib's fintech platform, serving millions of Indonesian investors. You'll lead...  ...We're Looking For - 5+ years in SRE/DevOps, with 2+ years managing engineers - Deep hands-on expertise in GCP and Kubernetes -... 
    Remote work

    Air

    United States
    4 days ago
  • $155k - $180k

     ...Site Reliability Engineering Manager (Remote) Join to apply for the Site Reliability Engineering Manager (Remote) role at WebstaurantStore Base pay range: $155,000.00/yr - $180,000.00/yr Job Summary As the largest online distributor of restaurant supplies... 
    H1b
    Remote work
    Home office

    WebstaurantStore

    Lititz, PA
    1 day ago
  •  ...enterprise initiatives such as public cloud, data science, AI, engineering innovation and IoT. Our customers include the world’s...  ...led, profitable and growing. We are hiring a Site Reliability Engineering Manager aspiring for a world-class devops and gitops engineering... 
    Full time
    Work at office
    Local area
    Remote work
    Work from home
    Worldwide

    Canonical Ltd.

    Remote
    16 days ago
  •  ...PNC is seeking a Software Engineering Manager—Site Reliability Engineering to lead a 24x7 production support team and drive reliability across mission-critical platforms powering PNC's digital experiences. You will manage incident response, RCA programs, and cross-functional... 

    Fairygodboss

    Lakewood, CO
    3 days ago
  • $135k - $138k

     .... Real People. Real Results. THAT is Link Snacks. Job Description JOB DESCRIPTION SUMMARY   The Site Reliability & Engineering Manager is the senior technical leader for the Minong facility, responsible for maintenance, reliability, facilities, utilities... 
    Permanent employment
    Full time
    For contractors
    Work at office
    Relocation package
    Shift work

    Jack Link's Protein Snacks

    Minong, WI
    3 days ago
  •  ...flex cards, and member engagement solutions. We partner with managed care organizations to provide innovative healthcare...  ...Location: Remote (US-based candidates only) Manager, Site Reliability Engineering (SRE) Position Overview We are seeking a Manager, Site... 
    Full time
    Remote work
    Flexible hours

    NationsBenefits, LLC

    Remote
    20 days ago
  •  ...future of our communities. This is a Lead Software Production Management & Reliability Engineering position at Director level which is part of the job...  ...Position Overview The Wealth Management Production Management Site Reliability Engineer position is a highly visible/... 
    Work at office

    Socket

    New York, NY
    3 days ago
  • PNC is seeking a Software Engineering Manager for Site Reliability Engineering (SRE) in Phoenix, AZ. You will lead a team to ensure reliability of critical platforms powering digital experiences, combining technical leadership with people management for 24x7 operations... 

    Fairygodboss

    Phoenix, AZ
    4 days ago
  • $140k - $230k

     ...Zoox is seeking a Site Reliability Engineer to help ensure the availability, performance, and resilience of the services that power the development...  ...or a similar role, with a strong, objective background in managing large-scale distributed systems. Cloud & Infrastructure... 
    Full time

    Zoox

    Remote
    17 hours ago
  • $175k - $250k

     ...developed by our expert team of lawyers, engineers and research scientists. We’ve found...  ...Overview As a Software Engineer on the Site Reliability team at Harvey, you will ensure the...  ...What You’ll Do Design, implement, and manage monitoring, alerting, and infrastructure... 
    Full time
    Relocation package

    Harvey

    Remote
    17 hours ago
  •  ...only provider of enterprise-scale context engines capable of analyzing trillions of real-...  ...seeking a highly skilled and motivated Site Reliability Engineer (SRE) to join our growing team...  ...deployment, monitoring, and incident management to continuously improve overall system... 
    Full time

    Lovelace Ai

    Pittsburgh, PA
    17 hours ago
  •  ...builds the platforms and tooling that help engineering teams develop, deploy, and operate...  ...default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll...  ...habits and tooling.Architect and manage the SLO and error-budget framework, empowering... 
    Permanent employment
    Work experience placement
    Work at office
    Local area

    Plaid Financial

    San Francisco, CA
    2 days ago
  • $104.43k - $156.65k

     ...Comcast prefers to have employees on-site collaborating unless the team has been...  ...businesses to solve their unique media management and video publishing requirements.Our...  ...Paramount+, and many others.Our Site Reliability Engineering (SRE) team is at the heart of our mission... 
    Permanent employment
    Full time
    Work at office
    Remote work
    Worldwide
    Flexible hours

    Comcast

    Centennial, CO
    17 hours ago
  • $100k - $120k

    OverviewThe Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and...  ...role.Strong knowledge of SRE best practices and incident management protocolsDeep experience using and/or configuring New Relic... 
    Full time
    Temporary work
    Work experience placement
    Flexible hours

    Origami Risk

    Atlanta, GA
    2 days ago
  • $105.6k - $145.2k

    Architect the Future as our Site Reliability Engineer!Are you ready to take your skills to the next level as a self-motivated and enthusiastic Site...  ..., logging, and alerting.Perform code deployments and manage CI/CD pipelines using Azure DevOps, GitHub, Terraform and related... 
    Ongoing contract
    Full time
    Work at office
    Local area
    Worldwide

    Trimble Navigation

    Westminster, CO
    4 days ago
  • $138.1k - $198.2k

     ...technology that simply works.  The SRE Engineering Enablement Team supports our CI...  ...environments, build tools, code review, artifact management, CI, education, and documentation. Our...  ...engineers at Cisco. Your Impact As a Site Reliability Engineer, you will be at the epicenter... 
    Permanent employment
    Full time
    Temporary work
    Work experience placement
    Local area
    Remote work
    Flexible hours

    CISCO Systems

    Chicago, IL
    2 days ago
  • $104.9k - $174.7k

     ...Verification, Fraud and Credit Risk mitigation and Customer Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability... 
    Full time
    Work at office
    Local area
    Remote work
    Work from home

    RELX Group

    Buford, GA
    2 days ago
  • $176k - $282k

     ...cloud environment.Design, deploy, and manage AWS infrastructure, including EC2,...  ...across operational excellence, security, reliability, performance efficiency, cost optimization, and sustainability.Perform Site Reliability Engineering (SRE) functions, including automation... 
    Contract work
    Shift work

    Peraton Corporation

    Annapolis Junction, MD
    3 days ago
  • $98.58k - $138.02k

     ...Silicon Valley Region / Denver, COProduct Engineering - DevOps /Full Time /HybridRestaurant3...  ..., TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for...  ...or CloudFormation. Work within change management protocols to provide maximum uptime for... 
    Full time
    Work at office

    Restaurant 365

    Austin, TX
    2 days ago
  •  ...collectively strive to build and maintain a rapid-feedback platform that enables our engineers to accomplish their own goals instead of creating friction.ResponsibilitiesEKS & Karpenter Management: Manage, upgrade, and autoscale EKS clusters across multiple environments (SIT,... 
    For contractors

    Varo Money

    San Francisco, CA
    12 hours ago
  • $112k - $137k

     ...will work at an MUFG office or client sites four days per week and work remotely one...  ...highly motivated Certified Sr. Cloud Site Reliability Engineer to build a robust, scalable, and...  ...through automated mechanisms.Develop and manage infrastructure as code using Terraform,... 
    Full time
    Work at office
    Local area
    Remote work

    MUFG

    Tempe, AZ
    3 days ago
  • $170k - $200k

    We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining...  ...maintain robust CI/CD pipelines for multiple environments.Manage datacenter infrastructure (Linux servers, network devices,... 
    Full time
    Worldwide

    Fortinet

    Sunnyvale, CA
    17 hours ago
  •  ...fully integrated solutions to manage everything from business...  ...what’s next.About the teamThe Engineering team at Airwallex is a diverse...  ...together to build scalable, reliable, and secure products that empower...  ....What you’ll doAs a Senior Site Reliability Engineer, you’ll... 
    Temporary work
    Local area
    Worldwide

    Airwallex

    San Francisco, CA
    3 days ago
  •  ...Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk.Design and implement processes that...  ...and reliability tooling.Actively participate in reliability engineering and resilience communities of practice, contributing to... 
    Full time

    Vanguard

    Wayne, PA
    1 day ago
  • $115.5k - $164.8k

     ...company where you matter.Your ImpactAs an engineer on the APX SRE CloudOps team, you will...  ...previously required human intervention with reliable, tested automation. You will also...  ...years of applicable experience.Experience managing cloud platforms such as Azure, AWS, or similar... 
    Work experience placement
    Work at office
    Remote work

    Axon

    Washington DC
    17 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Manager, Site Reliability Engineering. Be the first to apply!