Principal Site Reliability Engineer
Symmetrio
Role Description
Symmetrio is recruiting a Principal Site Reliability Engineer (SRE) for our customer, a rapidly growing healthcare technology organization focused on advanced healthcare technology solutions. This individual will play a critical role in ensuring the reliability, scalability, security, and performance of a mission-critical SaaS platform supporting healthcare providers across the United States.
The ideal candidate will possess a unique blend of cloud infrastructure expertise, application troubleshooting experience, production operations leadership, and customer-facing technical problem-solving skills. They will be equally comfortable:
- Investigating application-level issues
- Troubleshooting AWS networking and infrastructure
- Leading production incident response efforts
- Collaborating with development teams to improve operational excellence
Qualifications
- 6+ years of hands-on experience supporting and managing AWS-based production environments
- 4+ years of experience supporting web applications and backend services (Python/Django experience strongly preferred)
- Experience with AWS networking technologies including VPCs, Site-to-Site VPNs, Transit Gateways, routing, NAT gateways, and security groups
- Strong experience with Terraform and infrastructure-as-code deployment practices
- Experience with containerized environments including ECS, Fargate, Kubernetes, or similar technologies
- Experience building and supporting CI/CD pipelines and release automation processes
- Familiarity with monitoring and observability platforms such as Datadog, CloudWatch, Sentry, Grafana, or similar tools
- Experience leading production incidents, outage management, and root cause analysis initiatives
- Exposure to Windows Server environments, Active Directory, Kerberos, and enterprise infrastructure concepts is preferred
- Healthcare technology, healthcare SaaS, clinical software, or other regulated industry experience is highly preferred
- Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related technical field preferred
Requirements
- Serve as the primary technical owner for production reliability across U.S. customer environments
- Investigate and resolve complex issues spanning web applications, APIs, backend services, data pipelines, cloud infrastructure, and customer integrations
- Lead production incident response efforts, coordinating cross-functional teams to restore service and minimize customer impact
- Perform root cause analysis and drive corrective actions that improve long-term system stability and resilience
- Partner with software engineering and platform teams to identify recurring reliability risks and implement sustainable solutions
- Design, configure, and validate secure customer connectivity solutions including Site-to-Site VPNs, Transit Gateway integrations, routing configurations, and secure network paths
- Support customer onboarding initiatives by troubleshooting connectivity challenges and ensuring consistent implementation processes
- Enhance platform observability through improvements in monitoring, logging, alerting, tracing, and operational dashboards
- Contribute to CI/CD, infrastructure automation, and deployment processes that improve release safety and operational consistency
- Develop operational tooling that supports incident response, troubleshooting, onboarding, and system monitoring activities
- Collaborate with engineering leadership to improve cloud architecture, scalability, security, and operational readiness
- Partner with customer-facing teams to communicate technical issues, remediation plans, and reliability improvements in a clear and effective manner
- Support compliance, security, and risk management initiatives within highly regulated healthcare environments
Benefits
- Health Care Plan (Medical, Dental & Vision)
- Retirement Plan (401k, IRA)
- Paid Time Off (Vacation, Sick & Public Holidays)
- ...Site Reliability Engineering (SRE) Team Lead The Site Reliability Engineering (SRE) team is foundational to the growth and scale of our platform. You and your team will help advance several initiatives tied to automation, SRE culture, and cloud architecture. You will...PrincipalShift work
- ...The Platform Ops team within CloudOps is responsible for the reliability, scalability, and modernization of DigiCert’s cloud infrastructure... ...a Principle SRE, you will own the intersection of software engineering and operations—driving automation-first practices, reducing...PrincipalRemote workFlexible hours
$175.5k - $235.4k
...Principal Site Reliability Engineer We Power the Magic! That's our motto at Disney Experiences (DX). Our team creates world-class immersive digital experiences for the Company's premier vacation brands including Disney's Parks & Resorts worldwide, Disney Cruise Line...PrincipalWork experience placementRemote workWorldwide- ...Experience Technology. It works closely with other SREs, application delivery teams, and systems engineers from across the company. About The Role & Team The US Parks Site Reliability Organization is responsible for the operation of an assigned portfolio of applications...PrincipalWork experience placementLocal areaWorldwide
$163.62k - $212.71k
...Principal Site Reliability Engineer Bellevue, WA Immigration / Work Authorization Notice: Applicants must be currently authorized to work in the United States. iSpot is not able to sponsor or take over sponsorship of an employment visa for this position at this time...PrincipalFull timePart timeWork at officeLocal areaImmediate startRemote workWork from homeFlexible hoursShift work3 days per week1 day per week$160k - $180k
...global, with headquarters in Denver, Colorado, and offices across the U.S., Canada, and India. We are seeking a Principal Site Reliability Engineer to define the strategic vision and own the enterprise-wide reliability, scalability, and performance of our critical...PrincipalContract workTemporary workWork at officeWork from homeFlexible hours- ...ears, and hands on the ground at a government customer site, ensuring the reliability and performance of Twenty's mission-critical platform running... ...of deep technical ownership and customer-facing engineering: you'll define how we measure reliability, lead incident...Full timeWork at officeRemote workFlexible hours
$200k - $250k
Role Description As a Principal Site Reliability Engineer, you'll shape the long-term strategy for the infrastructure behind one of the most demanding platforms in sports betting and gaming. You'll drive the architectural direction of our cloud and on-premise platforms,...PrincipalFull time$185k - $278k
...transforming them into scalable solutions. • Debugging OS and engineering issues within our provided Linux environment. •... ...efficiently. The Impact You Will Have: • Enhancing the reliability and performance of our engineering environment. •...PrincipalRemote work- ...Job Description Job Description About the Role We're seeking an exceptional Principal Site Reliability Engineer to architect, design, and build our SRE foundation from the ground up at InfiniteChoice. This is a rare greenfield opportunity to establish SRE practices...PrincipalRemote work
$197k - $207k
...their requirements and build solutions, acting as a trusted technical advisor throughout the entire project lifecycle. Manage, engineer, deploy, monitor, and automate processes related to Red Hat community projects such as the GNOME Infrastructure. Champion Red Hat...PrincipalContract workWork experience placementWork at officeRemote workFlexible hours$150k - $170k
...Senior Site Reliability Engineer – Zip Co Join to apply for the Senior Site Reliability Engineer role at Zip Co At Zip, we build cloud‑native software applications that serve millions of customers and process billions of dollars in payments. We’re looking for a seasoned...Casual workWork at officeRemote workFlexible hours- ...of the world's most influential companies. As a Senior Principal Software Engineer at JPMorganChase within the CDAO AI/ML Data Platforms Team,... ...These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan,...Principal
- ...Seeking a full-time Senior Site Reliability Engineer with expertise in C# and .NET to ensure production reliability for customer-facing platforms and weather data services in a remote setting, focusing on high availability, incident response, and operational excellence...Full timeRemote work
$149.4k - $202k
...Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing...Remote work$178.13k - $205.4k
...telecommuting. Salary Range: $178,131 - $205,400 About You Basic Qualification Bachelor’s degree or foreign degree equivalent in Computer Engineering, Computer Science, Engineering, or related field plus five (5) years of progressive, post‑baccalaureate experience in job offered...Work at officeRemote workFlexible hours$200k - $220k
...pivots, and the challenges of building in a high‑growth startup, we’d love to talk. This is more than a job—it’s a journey. Site Reliability Engineers (SREs) are responsible for the overall performance and reliability of ASAPP's infrastructure and products. The team owns...Remote work- ...Site Reliability Engineers (SREs) are essential to PandaDoc's success, ensuring customers receive a reliable service with minimal downtime. The SRE team achieves this by: Owning the incident management processes and tools. Managing the observability stack and alerting...Full timeContract workRemote work
- ...Partner with software developers, platform engineers, and IT staff to improve system design,... ...requirements, service quality, reliability, security, and compliance needs. Drive... ...Required: ~8+ years of experience in Site Reliability Engineering, DevOps, Platform...Work at officeRemote work
$127k - $249k
...The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As...Work at officeLocal areaRemote workFlexible hours- ...The Role We're looking for a Senior Site Reliability Engineer to own the reliability, scalability, and operational excellence of the production systems that power Nectar's platform. We run high-volume data ingestion pipelines and real-time AI agents on top of a fast...Remote work
$86.9k - $198k
...Site Reliability Engineer, Senior The Opportunity Engineering to make a system more resilient and efficient frees up time and money to build more capabilities. Whether you come from a background in network engineering, systems administration, or software development, if...Full timeContract workPart timeLocal areaRemote work- ...headquartered in Atlanta. To learn more about QGenda, visit us at qgenda.com or follow us on Instagram or LinkedIn. As a Senior Site Reliability Engineer, you will work with our Infrastructure and Product Development Teams to increase the scalability, reliability, and...Permanent employmentWork at officeRemote workWork from home
- ...Senior Site Reliability Engineer We are seeking a full-time Senior Site Reliability Engineer at Garmin's U.S. headquarters in the Greater Kansas City area. In this role, you will be responsible for ensuring the integrity of Garmin's production environment is maintained...Full timeWork experience placementLocal areaRemote work
$160k - $240k
...passionate about building unified IT solutions that simplify the way IT organizations work. We are currently looking for a Senior Site Reliability Engineer to join our SRE team in the Platform Engineering organization and help us scale our products to millions of end-users. We...Permanent employmentFull timeRemote workWork from homeRelocationFlexible hours- ...To support a growing infrastructure team, the full-time Senior Site Reliability Engineer II - Infrastructure (AI Native) will design and maintain scalable platforms for over 200 backend services, utilizing AI tools to enhance operational efficiency in a remote work environment...Full timeRemote work
$135k - $150k
Senior Site Reliability Engineer Job number: 884 This is a remote position. Ad Hoc is a technology company that empowers organizations to deliver scalable, impactful digital services. Using modern, agile methods, our team creates products that meet people's...Remote workFlexible hours- ...with preference to candidates located in San Diego, CA, Norfolk, VA or Charleston, SC Position Overview: The Senior Site Reliability Engineer is a technical leader responsible for architecting the reliability strategy for large-scale, distributed government systems...Contract workRemote work
$195k - $240k
...Senior Site Reliability Engineer San Francisco (Hybrid) At You.com, we are building the AI Search Infrastructure that powers modern AI systems. Our goal is to create the trusted knowledge layer that agents, applications, and enterprises rely on to retrieve real-time...Full timeImmediate startRemote workWork from homeFlexible hours$117k - $209.33k
...Job Requisition ID # 26WD99276 Position Overview Want to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products. As part of a new...For contractorsRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Site Reliability Engineer. Be the first to apply!
- principal cloud engineer Remote
- senior principal engineer Remote
- assistant chief engineer Remote
- principal infrastructure engineer Remote
- general engineer Remote
- principal application developer Remote
- principal engineer Remote
- director of product engineering Remote
- director data engineering Remote
- data center chief engineer Remote



