Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Site Reliability Engineer

$165k - $185k

Tandem Diabetes

**GROW WITH US:** Tandem Diabetes Care creates new possibilities for people living with diabetes, their loved ones, and their healthcare providers through a positively different experience. We’d love for you to team up with us to “innovate every day,” put “people first,” and take the “no-shortcuts” approach that has propelled us to become a leader in the diabetes technology industry.**STAY AWESOME:** Tandem Diabetes Care is proud to manufacture and sell the Tandem Mobi system and t:slim X2 insulin pump with Control-IQ+ technology — an advanced predictive algorithm that automates insulin delivery. But we’re so much more than that. Our company’s human-centered approach to design, development, and support delivers innovative products and services for people who use insulin. Because many of our own team members live with diabetes, or have a loved one impacted by diabetes, the work is personal, and we are committed to the cause. Learn more at tandemdiabetes.com**A DAY IN THE LIFE:**The Principal Site Reliability Engineer (SRE) is responsible for the reliability, availability, and performance of the company's production systems. This role leads day-to-day production support and incident response and progressively replaces reactive firefighting with engineered SRE practice: SLOs, observability, on-call design, runbooks, and automation. It also advances infrastructure automation (Terraform/IaC), CI/CD reliability, disaster recovery, and security and compliance readiness in partnership with cross-functional teams. The role works with a distributed team that includes offshore consulting partners and is expected to raise their capability and independence; success is measured as much by what the team can do without the Principal SRE as by what the Principal SRE delivers personally.The Principal Site Reliability Engineer (SRE)'s at Tandem are also responsible for:**Production Support & Incident Management*** Leads day-to-day production support: intake, triage, prioritization, escalation, queue health, and change execution.* Establishes consistent support practices across a distributed team, including shift handoffs, ticket quality standards, and clear ownership of open issues.* Leads incident management end-to-end: incident command, stakeholder communication, and blameless postmortems with corrective actions tracked to closure.* Participates in and coordinates response activities for production security incidents, partnering with Security teams to contain threats, restore services, validate controls, and implement corrective actions.* Owns on-call strategy: rotation design, escalation paths, alert tuning, and tooling (e.g., PagerDuty, New Relic), with explicit goals of reducing alert fatigue and building coverage that works across time zones. Participates as a senior escalation tier for high-severity incidents.* Builds runbooks that standardize response to common failure modes and enable first-line resolution by engineers who did not build the system.**Reliability Engineering & Observability*** Defines and owns SLIs and SLOs for critical services, using them to guide monitoring, alerting, and reliability priorities.* Drives systemic reduction of MTTD and MTTR through better instrumentation, alerting, diagnostics, and automation.* Converts recurring support burden into permanent fixes, automation, or documentation rather than absorbing it as ongoing manual work.* Establishes and maintains technology currency and lifecycle management practices for production platforms, ensuring cloud services, Kubernetes clusters, operating systems, runtimes, and infrastructure components remain supported, secure, and aligned with organizational standards. Proactively identifies and mitigates End-of-Life (EOL), End-of-Support (EOS), and technology obsolescence risks.* Owns business continuity and disaster recovery readiness for production platforms, including backup and recovery strategies, recovery testing, failover capabilities, recovery runbooks, and adherence to defined RTO/RPO objectives.**Infrastructure Automation & CI/CD*** Leads infrastructure automation with Terraform, ensuring infrastructure is version-controlled, modular, reusable, and auditable, working within and improving existing patterns where they are sound.* Eliminates toil through automation, reducing manual, repetitive operational work across the team.* Adds reliability guardrails to CI/CD pipelines (automated rollback, change-risk checks, progressive delivery) with the teams that own them.**Operational Governance and Compliance*** Maintains production systems in accordance with applicable regulatory and compliance requirements, ensuring audit readiness and the ongoing effectiveness of access management, change management, logging, vulnerability remediation, patch management, and other operational controls.* Maintains documentation and evidence supporting business continuity and disaster recovery controls, including recovery testing results, remediation plans, and audit artifacts.* Partners with Security, Quality, and Compliance teams to maintain regulatory compliance, support internal and external audits, and drive timely resolution of audit findings and remediation activities.**Leadership, Mentorship & Collaboration*** Actively grows the capability of SRE and DevOps engineers, including consulting-partner engineers, through pairing, design and code review, and incident debriefs.* Creates conditions in which junior and contract engineers raise concerns and disagree openly, recognizing that silent agreement is a production risk.* Communicates in a documentation-first manner suited to a distributed, multi-time-zone team, so context is durable rather than held by one person.* Partners with software engineering, QA, and architecture to embed reliability and operability into the development lifecycle rather than only after incidents.**Contributing Responsibilities** (partners with other owners):* Informs capacity planning and scaling strategy with the Test team (who own load testing and modeling) and introduces proactive resilience testing such as game days once baseline support and observability are stable.* Partners with application, architecture, and business stakeholders to align business continuity and disaster recovery capabilities with application requirements and recovery objectives.* Supports cloud cost optimization: rightsizing, reserved capacity, and observability spend governance.* Ensures work complies with company policies, including Privacy/HIPAA and other regulatory, legal, and safety requirements. Other duties as assigned.**WHEN & WHERE YOU’LL WORK:**Remote: This position is fully remote and open to candidates within the United States. Equipment for the role will be provided and training will occur virtually.**WHAT YOU’LL NEED:*** We are looking for a forward-thinking professional who actively **leverages AI** to enhance, automate, and optimize internal tooling and system integrations. In this role, you won't just use existing systems—you will use AI to actively build smarter workflows and bridge gaps between our platforms.* Demonstrated experience leading production support and incident management for production systems, including incident command during high-severity events.* Strong grounding in SRE principles: SLIs/SLOs, blameless postmortems, toil reduction, and treating reliability as an engineering discipline rather than a purely operational function.* Demonstrated experience owning on-call strategy, including rotation design, alert tuning, and escalation.* Expertise with Terraform or comparable IaC at scale: module design, state management, and policy-as-code guardrails.* Hands-on experience building CI/CD pipelines with reliability guardrails using tools such as GitHub Actions, Octopus Deploy, or Azure DevOps.* Deep experience with at least one major cloud platform (AWS, Azure, or GCP; [ preferred]) and with containerization and orchestration (Docker/Kubernetes).* Working knowledge of observability tooling (e.g., Prometheus, Grafana, Datadog, CloudWatch, ELK/OpenSearch) and using it to drive SLO-based alerting.* Experience designing and testing disaster recovery, including backup/restore, failover, and RTO/RPO validation.* Working knowledge of cloud security and compliance practices (IAM, network segmentation, encryption, vulnerability and patch management) and cloud cost optimization.* Proficiency in at least one scripting or programming language (e.g., Python, Go, Bash) for automation and tooling.* Experience in FDA and ISO regulated industries and with agile methodologies preferred.**EXTRA AWESOME:*** B.S. in Computer Science or an equivalent combination of education and applicable job experience, including technical school training and certifications in networks, servers, and cloud infrastructure; demonstrated production experience weighs more heavily than degree.* Relevant cloud certifications (e.g., AWS/Azure/GCP Professional or Architect level) preferred.* 10+ years in Site Reliability Engineering, DevOps, or infrastructure engineering, including: + Leading production support, incident management, and postmortem practice. + Managing production cloud infrastructure as code (Terraform or equivalent) at scale. + Operating and improving CI/CD pipelines in a production environment. + Designing, executing, and testing disaster recovery plans. + Partnering with security teams on production security and compliance posture.* 2+ years mentoring or technically leading other engineers, including engineers who are remote, offshore, or contracted through a partner organization.**COMPENSATION & BENEFITS:**The starting base pay range for this position is **$165,000 to $185,000** annually. Base pay will vary based on job-related knowledge, skills, experience and may also fluctuate depending on candidate’s location and the overall job market. In addition to base pay, Tandem offers a competitive compensation package that includes bonus and a robust benefits package.Tandem offers health care benefits such as medical, dental, vision available your first day, as well as health savings accounts and flexible saving accounts. You’ll also receive 11 paid holidays per year, a minimum of 20 days of paid time off (with accrual starting on day 1) and you will have access to a 401k plan with company match as well as an Employee Stock Purchase plan. Learn more about Tandem’s benefits here!**YOU SHOULD KNOW:**Qualified applicants with arrest or conviction records will be considered for employment in accordance with applicable state and local Fair Chance laws and regulations. A conditional offer of employment from Tandem is contingent upon successful completion of a pre-employment screening process comprised of a drug test (excluding marijuana) and background check, which includes a review of criminal history information.Tandem has good cause to conduct a review of criminal history information of candidates for this position, as this role may involve access to proprietary, sensitive and/or confidential information, including customer protected health information. This review is required to ensure that individuals in such roles uphold high standards of trust and integrity so as to protect the interests of our customers, employees, and stakeholders.**WHY YOU’LL LOVE WORKING HERE:** At Tandem, we believe joy fuels excellence. That's why we've built a workplace that celebrates your achievements and supports your well-being. Our team thrives on pushing boundaries and fostering growth, all while maintaining a spirit of fun and camaraderie. This is just one of the ways we stay awesome! Explore the benefits and reasons to love Tandem at YOU, WITH US!**We embrace the value that every single one of us brings to the table. But sometimes we forget that when we don’t meet 100% of a job description’s criteria – maybe you’re feeling that way right now? We encourage you to apply anyway. Because we want you to be you, with us.Tandem is firmly committed to being an equal opportunity employer and does not discriminate on the basis of age, disability, sex, race, religion or belief, gender identity or expression, marriage/civil partnership, pregnancy/maternity, or sexual orientation. We are an inclusive organization, and we welcome applications from a wide range of candidates. Selection for roles will be based on individual merit alone.**REFERRALS:** We love a good referral! If you know someone who would be a great fit for this position, please share!**APPLICATION DEADLINE:** The position will be posted until a final candidate is selected for the requisition or the requisition has a sufficient number of applications.Make a move that matters. Join Tandem Diabetes Care, where we're turning challenges into triumphs every day and where your talents will help shape a healthier, happier tomorrow. #J-18808-Ljbffr

Vacancy posted 4 hours ago
Similar jobs that could be interesting for youBased on the Principal Site Reliability Engineer in Eastern, KY vacancy
  •  ...Onshape is hiring a Principal Software Engineer (SRE) in a hybrid role based in Boston, MA. You will lead reliability initiatives, shape strategy, and act as a technical authority to ensure the platform is fast, resilient, and scalable for customers. You will drive... 
    Principal

    Onshape

    Eastern, KY
    4 hours ago
  •  ...role will be supporting products and services, including Gen Insurance and Engine by Gen, delivering trusted digital and financial experiences to consumers at scale. As a Principal Site Reliability Engineer, you will define and lead the reliability, scalability, and... 
    Principal

    Apply

    Eastern, KY
    1 day ago
  • $169.3k - $304.7k

     ...building and maintaining fast, efficient, scalable, and reliable routing software and infrastructure that is responsible...  ...the growth and stability of our global platform. As a Principal Site Reliability Engineer - Network, you will be responsible for: Architecting... 
    Principal
    Work experience placement
    Work at office
    Remote work

    Akamai

    Eastern, KY
    10 hours ago
  • $194k - $237k

    ## Principal Site Reliability EngineerApplylocations: Scottsdaletime type: Full timeposted on: Posted 5 Days Agojob requisition id: REQ2026426At...  ....**Overall Purpose**The Principal Site Reliability Engineer partners with development teams by designing availability... 
    Principal
    Hourly pay
    Work at office
    Immediate start
    Visa sponsorship
    Work visa
    Flexible hours

    Early Warning Services

    Eastern, KY
    3 days ago
  • $160k - $180k

     ...global, with headquarters in Denver, Colorado, and offices across the U.S., Canada, and India. We are seeking a Principal Site Reliability Engineer to define the strategic vision and own the enterprise-wide reliability, scalability, and performance of our critical... 
    Principal
    Flexible hours

    Hirebridge

    Eastern, KY
    4 hours ago
  •  ...Discover exciting DevOps job opportunities and connect with 28,396 DevOps professionals. The Senior Site Reliability Engineer role at Jobicy is designed for experienced professionals who are passionate about enhancing system reliability and operational efficiency. The... 
    Remote work
    Flexible hours

    DevOpsChat

    Eastern, KY
    4 hours ago
  • $182.8k - $247.3k

     ...to develop education for our half a billion (and growing!) learners around the world. About the role... As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed systems... 
    Work experience placement

    Socket

    Eastern, KY
    3 days ago
  •  ...profitable developer-tooling company whose product is used by engineering teams at thousands of software companies for application...  ...well-resourced group of nine. As Senior SRE you will lead reliability initiatives across the platform — from defining and driving SLOs... 

    Kovoro

    Eastern, KY
    4 hours ago
  •  ...% uptime. You'll own SLOs, incident response, and production reliability for a system that processes millions of identity verifications...  ...Sentry error tracking, structured logging Implement chaos engineering practices to proactively identify failure modes Optimize... 
    Remote work

    Xident B.V.

    Eastern, KY
    4 hours ago
  • $180k - $230k

     ...Acceleration Job Description We're looking for a Senior SRE to own the reliability, scalability, and observability of our production systems. You'll work closely with platform and data engineering to keep high-throughput, data-intensive services running at the... 
    Work at office
    Local area
    Immediate start
    Remote work
    3 days per week

    GridCARE, Inc.

    Eastern, KY
    5 days ago
  •  ...As a Senior Site Reliability Engineer on our cloud engineering team, you'll keep our production environment healthy, secure, and running smoothly. This is an operations-focused role: you'll own the day-to-day administration of our AWS accounts and databases, backup posture... 
    Work experience placement

    Apply

    Eastern, KY
    4 hours ago
  •  ...Cloudflare, GitHub Actions, PostgreSQL, Redis/BullMQ, Node.js/NestJS, Datadog, TypeScript, React, SQL Position: Senior Site Reliability Engineer Engagement period: Ongoing Interview timeline: ASAP Interview process: 1) CV review 2) Interview with our CTO 3)... 
    Contract work
    Immediate start

    Devspace

    Eastern, KY
    4 hours ago
  • $110k - $145k

     ...global, with headquarters in Denver, Colorado, and offices across the U.S., Canada, and India. We are seeking a Senior Site Reliability Engineer to own the reliability, scalability, performance, and operational integrity of critical production services. This role is... 
    Flexible hours

    Hirebridge

    Eastern, KY
    4 hours ago
  • $150.4k - $277.6k

     ...Services The Media Platforms SRE team under the Apple Service Engineering division is one of the most exciting examples of Apple’s long...  ...field with 4+ years experience At least 6 years in a Reliability Engineering, DevOps or infrastructure focused role Advanced... 
    Relocation
    Day shift

    Apple

    Eastern, KY
    1 day ago
  •  ...Quarterhill is seeking a Senior Site Reliability Engineer (SRE) to join our growing team. This role is an exciting opportunity to contribute to the reliability and performance of smart transportation systems, including a next-generation, cloud-native tolling platform that... 
    Local area

    Electronic Transaction Consultants

    Eastern, KY
    4 hours ago
  • $150k - $220k

     ...Senior Site Reliability EngineerJob detailsDepartment / EngineeringRemoteFull-time$150,000 USD - $220,000 USD## About UsMetaRouter is a customer...  ...architecture.## About The RoleAs a Senior Site Reliability Engineer, you own significant pieces of our infrastructure and... 
    Full time
    Remote work

    PLP Group

    Eastern, KY
    4 hours ago
  •  ...match. The role We're looking for a Senior SRE to own the reliability, scalability, and operational posture of Satsuma's multi-...  ...using AI-assisted development workflows Partner closely with engineering on reliability reviews and architecture decisions ~5-8... 

    Satsuma AI, Inc.

    Eastern, KY
    4 hours ago
  • $148.5k - $223.9k

     ...Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations, this organization provides a global team of engineers monitoring cloud service... 
    Worldwide
    Weekend work

    Salesforce.Com Inc

    Eastern, KY
    4 hours ago
  • $180k - $200k

     ...Come join tastytrade, part of IG Group, as we build the reliability practice behind the brokerage platform that active options...  ...equities traders rely on every market day. As our first Senior Site Reliability Engineer, you'll define what reliability means at tastytrade, from... 
    Work at office
    3 days per week

    tastyworks

    Eastern, KY
    4 hours ago
  • $104k - $178k

    ## Sr. Site Reliability Engineer IApply: Hybrid: NYC Global HQ: Full time: Posted 12 Days Ago: JR00000779# ****Who We Are****DV is the leader in digital performance solutions, helping our advertiser and agency partners Verify the quality of their digital campaigns, Optimise... 
    Full time

    DoubleVerify

    Eastern, KY
    4 hours ago
  •  ...to grow, adapt and use your skills consistently. Our customers rely on us in the moments that matter. Engineering delivers on that promise.The Senior Site Reliability Engineer is responsible for ensuring our SaaS products are fast, stable and optimized for our customers... 
    Work experience placement
    Flexible hours

    Donnelley Financial Solutions

    Eastern, KY
    4 hours ago
  • $135.2k - $181.2k

     ...enhance electrical, mechanical, and sensor-based systems to ensure reliability and performance. Configure, calibrate, and validate...  ...professional development, including an interest in emerging data engineering tools and methodologies. Preferred Qualifications: ~8+... 
    Worldwide

    Disney Cruise Line

    Eastern, KY
    4 hours ago
  •  ...Engineering, Product, Design, and Marketing Engineering Compensation ~ Zone 1 Base Pay: $214K – $260K Superhuman offers...  ...role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them... 
    Worldwide
    Home office
    Flexible hours

    Superhuman

    Eastern, KY
    2 days ago
  •  ...A senior Site Reliability Engineer will join an established infrastructure function responsible for highly available, security-conscious cloud systems supporting complex business-critical workloads. You’ll take significant ownership of reliability, scalability, and... 
    Full time
    Remote work

    Motion Recruitment Partners LLC

    Eastern, KY
    4 hours ago
  • $114k - $148k

     ...Total compensation is based on experience, skills, and location using objective, job-related criteria. Summary As a Site Reliability Engineer, you will focus on ensuring the platform and services customers rely on are reliable, performant, and highly available. If... 
    Work experience placement

    OneStream Software

    Eastern, KY
    4 hours ago
  • $107.9k - $195.05k

     ...The Digital Sector at Leidos currently has an opening for a Site Reliability Engineer (SRE) / Senior Cloud Engineer to work in our Baltimore, Maryland office. This is an exciting opportunity to use your experience helping the Center for Medicare and Medicaid Services (... 
    Contract work
    Work at office

    Koitecc Solutions

    Eastern, KY
    4 hours ago
  • $110k - $145k

     ...operations. You will liaise with product and engineering teams to ensure applications and...  ...feedback loop for platform and product reliability. The ideal candidate is a solutions-oriented...  ...experience as a platform engineer, site reliability engineer, systems engineer... 
    Work experience placement

    CoSM

    Eastern, KY
    4 hours ago
  • $115.5k - $164.8k

     ...matters at a company where you matter. Your Impact As an engineer on the APX SRE CloudOps team, you will spend a significant portion...  ...that replace what previously required human intervention with reliable, tested automation. You will also participate in on‑call... 
    Work experience placement
    Work at office
    Remote work

    Koitecc Solutions

    Eastern, KY
    4 hours ago
  • $140k - $195k

     ...Improve reliability, observability, service health, incident response, and operational readiness. CodeVertex works across data...  ..., secure systems, and operational clarity matter. The Site Reliability Engineer role helps turn business needs into reliable execution, whether... 
    Remote work

    Codevertex Innovations

    Eastern, KY
    4 hours ago
  •  ...personalized care faster. We are building AI agents to support the full arc of the patient journey. The Opportunity: Machine Learning Engineer Patients count on our platform 24/7. You'll build and maintain the tooling, alerts and incident-response playbooks that keep... 

    Tala Health

    Eastern, KY
    4 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Site Reliability Engineer. Be the first to apply!