Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Site Reliability Engineer

Full-time

Symmetrio

Role Description

Symmetrio is recruiting a Principal Site Reliability Engineer (SRE) for our customer, a rapidly growing healthcare technology organization focused on advanced healthcare technology solutions. This individual will play a critical role in ensuring the reliability, scalability, security, and performance of a mission-critical SaaS platform supporting healthcare providers across the United States.

The ideal candidate will possess a unique blend of cloud infrastructure expertise, application troubleshooting experience, production operations leadership, and customer-facing technical problem-solving skills. They will be equally comfortable:

  • Investigating application-level issues
  • Troubleshooting AWS networking and infrastructure
  • Leading production incident response efforts
  • Collaborating with development teams to improve operational excellence

Qualifications

  • 6+ years of hands-on experience supporting and managing AWS-based production environments
  • 4+ years of experience supporting web applications and backend services (Python/Django experience strongly preferred)
  • Experience with AWS networking technologies including VPCs, Site-to-Site VPNs, Transit Gateways, routing, NAT gateways, and security groups
  • Strong experience with Terraform and infrastructure-as-code deployment practices
  • Experience with containerized environments including ECS, Fargate, Kubernetes, or similar technologies
  • Experience building and supporting CI/CD pipelines and release automation processes
  • Familiarity with monitoring and observability platforms such as Datadog, CloudWatch, Sentry, Grafana, or similar tools
  • Experience leading production incidents, outage management, and root cause analysis initiatives
  • Exposure to Windows Server environments, Active Directory, Kerberos, and enterprise infrastructure concepts is preferred
  • Healthcare technology, healthcare SaaS, clinical software, or other regulated industry experience is highly preferred
  • Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related technical field preferred

Requirements

  • Serve as the primary technical owner for production reliability across U.S. customer environments
  • Investigate and resolve complex issues spanning web applications, APIs, backend services, data pipelines, cloud infrastructure, and customer integrations
  • Lead production incident response efforts, coordinating cross-functional teams to restore service and minimize customer impact
  • Perform root cause analysis and drive corrective actions that improve long-term system stability and resilience
  • Partner with software engineering and platform teams to identify recurring reliability risks and implement sustainable solutions
  • Design, configure, and validate secure customer connectivity solutions including Site-to-Site VPNs, Transit Gateway integrations, routing configurations, and secure network paths
  • Support customer onboarding initiatives by troubleshooting connectivity challenges and ensuring consistent implementation processes
  • Enhance platform observability through improvements in monitoring, logging, alerting, tracing, and operational dashboards
  • Contribute to CI/CD, infrastructure automation, and deployment processes that improve release safety and operational consistency
  • Develop operational tooling that supports incident response, troubleshooting, onboarding, and system monitoring activities
  • Collaborate with engineering leadership to improve cloud architecture, scalability, security, and operational readiness
  • Partner with customer-facing teams to communicate technical issues, remediation plans, and reliability improvements in a clear and effective manner
  • Support compliance, security, and risk management initiatives within highly regulated healthcare environments

Benefits

  • Health Care Plan (Medical, Dental & Vision)
  • Retirement Plan (401k, IRA)
  • Paid Time Off (Vacation, Sick & Public Holidays)
Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Principal Site Reliability Engineer in Remote vacancy
  • $142.8k - $274.8k

     ...yearEmployment type: Full-TimeWork site: 0 days / week in-office - remoteRole...  ...Software EngineeringDiscipline: Site Reliability EngineeringCompany:...  ...world’s most demanding workloads. As a Principal Site Reliability Engineer, you will set technical and operational... 
    Principal
    Ongoing contract
    Work at office
    Local area

    Microsoft

    Redmond, WA
    3 days ago
  •  ...lives. Ready to build the next breakthrough? Join us to start Caring. Connecting. Growing together.We are seeking a Principal Site Reliability Engineer (SRE) to define and scale reliability practices across large-scale cloud platforms.This is a senior individual contributor... 
    Principal
    Minimum wage
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work

    UnitedHealth Group

    Minnetonka, MN
    1 day ago
  • $120k

     ...Principal Site Reliability Engineer location- Washington DC Remote- No Salary- $120K/Y Tech M/ Amtrak Job Summary we're seeking a seasoned Principal Site Reliability Engineer with a strong focus on pipelines as code, CI... 
    Principal
    Remote work

    Yochana

    Washington DC
    1 day ago
  •  ...Seeking a full-time Principal Site Reliability Engineer to work remotely in the United States, responsible for architecting and scaling essential infrastructure, building a Kubernetes-based platform, and developing backend services to support autonomous systems. Key responsibilities... 
    Principal
    Full time
    Remote work

    Virtual Vocations Inc

    United States
    3 days ago
  •  ...Senior Principal Site Reliability Engineer Hong Kong SAR About Us Established in 2018, Bybit is one of the world's leading cryptocurrency exchanges and digital financial platforms, serving over 80 million users across more than 200 countries and regions. Powered... 
    Principal
    Remote work

    Bybit

    United States
    5 days ago
  • $159k - $272k

     ...generosity. Join us for the opportunity to grow and make a difference in ways that matter to you. Role SummaryIn this role as Principal Site Reliability Engineer, Infrastructure Observability you will help formulate, develop, and implement a team of Site Reliability Engineers (... 
    Principal
    Full time
    Private practice
    Local area
    Remote work
    Work from home
    3 days per week

    T. Rowe Price

    Owings Mills, MD
    1 day ago
  •  ...Your Impact The Site Reliability Engineering (SRE) team is foundational to the growth and scale of our platform. You and your team will help advance several initiatives tied to automation, SRE culture, and cloud architecture. You will ensure reliability and trust to over... 
    Principal
    Shift work

    Jobiqo

    Oklahoma City, OK
    2 days ago
  •  ...Principal Site Reliability Engineer Deimos is a cloud-native developer and security operations technology services company. We help companies of all sizes adopt the cloud for improved service delivery to their clients. We're a fully remote African-based team of engineers... 
    Principal
    Currently hiring
    Remote work
    Work from home

    Deimos

    United States
    3 days ago
  • $200k - $250k

     ...shaping it, one bold step at a time. To those who see AI as a driver of progress, come build the future together. As a Principal Site Reliability Engineer, you'll shape the long-term strategy for the infrastructure behind one of the most demanding platforms in sports... 
    Principal
    Full time
    Immediate start
    Remote work

    DraftKings

    United States
    4 days ago
  • $175.5k - $235.4k

     ...MyDisneyExperience and Hey, Disney! This role sits in the Commerce Site Reliability Engineering (SRE) specifically supporting Ecommerce , Consumer...  ...Technology teams from across the company. The Principal of DXT SRE will report to the Director of DXT Commerce SRE... 
    Principal
    Work experience placement
    Worldwide

    The Walt Disney Company

    Orlando, FL
    3 days ago
  • $84.9k - $209.5k

     ...spirit that promotes an upbeat and creative environment. We are unencumbered and will need your contribution to make it a special engineering center with the focus on excellence. Health Data Intelligence Platform has a rare opportunity to play a critical role in how... 
    Principal
    Temporary work
    Immediate start
    Remote work
    Flexible hours

    Hackajob

    United States
    5 days ago
  •  ...with software development teams to build reliable, scalable, secure, and cloud-native...  ...influence scalable architecture patterns across engineering teams, helping ensure systems are...  ...~8+ years of hands-on experience in Site Reliability Engineering, DevOps, cloud infrastructure... 
    Principal
    Remote work

    ABC Fitness Solutions, LLC

    United States
    5 days ago
  •  ...serve.The Information Technology group delivers secure, reliable technology solutions that enable DTCC to be the trusted...  ...scalability, and performance of enterprise platforms.As a Principal Site Reliability Engineer (SRE), you will drive operational excellence across... 
    Principal
    Remote work
    Flexible hours

    DTCC- The Depository Trust & Clearing Corporation

    Jersey City, NJ
    1 day ago
  •  ...: Meghana GorusuCompany: SRI Tech SolutionsJob Title: Senior Site Reliability EngineerLocation: Plano , TX (remote)Years of Experience: 8 to...  ...are seeking a highly skilled Senior Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate will have... 
    Remote work

    SRI Tech

    Plano, TX
    4 days ago
  • $174k - $252k

     ...systems by pushing for changes that improve reliability and velocity.Practice sustainable...  ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical...  ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you... 

    Google

    Raleigh, NC
    4 days ago
  • $210k - $230k

    GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation... 
    Currently hiring
    Remote work

    Govcio

    Arlington, VA
    1 day ago
  • $104.9k - $174.7k

     ...Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory... 
    Full time
    Work at office
    Local area
    Remote work
    Work from home

    LexisNexis Risk Solutions Group

    Alpharetta, GA
    3 days ago
  •  ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and...  ...and networking teams to improve service reliability and deployment workflowsDeploy and...  ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    1 day ago
  •  ...Lambda’s designated work from home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our cloud offering...  ...and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  • Recognized as the No. 1 site trusted by real estate professionals, Realtor.com has been at the forefront of online real estate...  ...confidence through expert guidance.We are seeking a Senior Site Reliability Engineer to join our newly formed Operations Excellence organization,... 
    Work at office
    Local area

    Realtor.com

    Austin, TX
    6 hours ago
  • $90k - $180k

     ...nutritionals and branded generic medicines. Our 115,000 colleagues serve people in more than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale, CA location in the Cardiac Rhythm Management Division.We are... 
    Remote work

    Abbott

    Sunnyvale, CA
    1 day ago
  • $15k

     ...packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage engineering skills... 
    Work at office
    Local area
    Remote work

    The Voleon Group

    Berkeley, CA
    2 days ago
  • $267k - $356k

     ...day is currently Tuesday.Lambda's Storage Engineering team is the backbone behind our world-...  ...workloads in the industry, which means reliability and performance aren't just goals—they're...  ...defined storage across new and existing sites using tools such as Ansible, Jenkins etc... 
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Chicago, IL
    1 day ago
  • Reliability Engineering Design, implement, and operate scalable, resilient, and highly available systems on Google Cloud Platform. Improve service...  ...Skills, and Abilities Three or more years of experience in Site Reliability Engineering, platform engineering, DevOps, cloud... 
    Remote work

    Patterson-UTI

    Houston, TX
    2 days ago
  • $158.5k - $172k

     ...exceptional value they deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will design, automate, and...  .... This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire technology... 
    Full time
    Temporary work
    Work at office
    Flexible hours
    3 days per week

    GrubHub

    Chicago, IL
    4 days ago
  • $86.9k - $198k

    Site Reliability Engineer, SeniorThe Opportunity: Engineering to make a system more resilient and efficient frees up time and money to build more capabilities. Whether you come from a background in network engineering, systems administration, or software development, if... 
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Aurora, CO
    1 day ago
  • $150k - $180k

     ...Umbra.About the JobWe are seeking an experienced SeniorSite Reliability Engineer to help design, build, operate, and scale the mission- and business...  ...impact across the organization.This position is based on-site in either our Arlington, VA office, Reston, VA office or... 
    Permanent employment
    Full time
    Work at office
    Local area
    Remote work
    Worldwide

    Umbra

    Arlington, VA
    4 days ago
  • $117k - $209.33k

    Job Requisition ID #26WD99276Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team supporting... 
    Full time
    For contractors
    Remote work

    Autodesk

    Plano, TX
    6 hours ago
  • Edmond, OKYouVersion - YouVersion Engineering /Full-Time/ Salary /On-siteThe YouVersion Senior Site Reliability Engineer is responsible for ensuring the integrity, performance, reliability, and cost-effectiveness of the cloud-based infrastructure and related systems supporting... 
    Full time
    Contract work
    Temporary work
    Work experience placement
    Casual work
    Internship
    Local area
    Worldwide

    Life.Church

    Edmond, OK
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Site Reliability Engineer. Be the first to apply!