Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Site Reliability Engineer

MetaRouter

Principal Site Reliability Engineer

As a Principal Site Reliability Engineer, you set the reliability strategy for the platform. You will define how we build, deploy, observe, and operate a distributed system that runs both in our own cloud and inside customer-controlled environments — and you will hold the organization to that standard. We run dedicated, isolated environments per customer, which makes repeatability and automation the central engineering problem rather than an afterthought. Depth of judgment about reliability engineering matters far more here than experience with any particular cloud, orchestrator, or observability vendor. This is an individual contributor role with organization-level influence.

Core Responsibilities:

  • Own the reliability architecture of the platform: deployment topology, failure domains, capacity strategy, and the automation that makes environments reproducible.
  • Define service level objectives with product and engineering leadership, and drive the work needed to meet them.
  • Set the standard for observability — dashboards, logs, metrics, tracing, and alerting — so that issues are detected before customers report them.
  • Lead major incidents, run blameless postmortems, and make sure the corrective work actually lands.
  • Contribute to design and architecture across infrastructure and applications, with automation, performance, reliability, and security as first-class concerns.
  • Drive infrastructure lifecycle at scale: provisioning, upgrades, and decommissioning across many isolated environments.
  • Ensure infrastructure and applications meet or exceed enterprise compliance requirements, and design identity and access controls across platforms and services.
  • Partner with enterprise customers on custom infrastructure requirements, translating their constraints into repeatable patterns rather than one-off work.
  • Raise the bar through code and design review, and mentor SREs and product engineers on reliability practice.
  • Improve and maintain infrastructure and process documentation.
  • Participate in and help evolve the on-call rotation, including how the team balances operational load against project work.

Qualifications and Experience:

  • 10+ years in infrastructure, SRE, or platform engineering, including deep experience operating large-scale distributed systems in production.
  • Expertise designing, analyzing, and troubleshooting distributed systems, with a track record of reliability decisions that held up under growth.
  • Deep experience with at least one major public cloud provider, and the ability to reason across providers rather than within one.
  • Strong command of container orchestration: cluster operation, workload scheduling, networking, and the failure modes that come with them.
  • Fluency with infrastructure as code, configuration management, and CI/CD pipeline design.
  • Strong scripting and automation ability, and comfort reading and debugging application code in the languages your services are written in.
  • Experience defining observability strategy — instrumentation, query languages, dashboards, and alert design that minimizes noise.
  • Demonstrated ability to influence without authority and align multiple teams behind a technical direction.
  • Experience operating under enterprise security and compliance frameworks.
  • Solid understanding of Unix/Linux operating systems and networking fundamentals.
Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Principal Site Reliability Engineer in United States vacancy
  •  ...Principal Site Reliability Engineer Location: New York, NY (Onsite) Job type: Contract Job Description: Job Requirements Must Have: - Site Reliability Engineering and system reliability optimization - GitLab CI/CD and HashiCorp Vault secrets management -... 
    Principal
    Contract work

    Argyle Infotech

    New York, NY
    4 days ago
  •  ...Setting the reliability strategy for the platform, the full-time Principal Site Reliability Engineer will define deployment and operational standards for distributed systems, ensuring reliability and automation across customer environments while working remotely. Key... 
    Principal
    Full time
    Remote work

    Virtual Vocations Inc

    United States
    3 days ago
  • $139.7k - $232.9k

     ...designing, implementing, and continuously improving highly reliable, scalable, and resilient platform solutions across the enterprise. Operates as a subject matter expert (SME) in Site Reliability Engineering, driving reliability engineering practices, operational excellence... 
    Principal
    Full time
    Work experience placement

    M&T Bank

    Buffalo, NY
    2 days ago
  • $142.8k - $274.8k

     ...yearEmployment type: Full-TimeWork site: 0 days / week in-office - remoteRole...  ...Software EngineeringDiscipline: Site Reliability EngineeringCompany:...  ...world’s most demanding workloads. As a Principal Site Reliability Engineer, you will set technical and operational... 
    Principal
    Ongoing contract
    Work at office
    Local area

    Microsoft

    Redmond, WA
    3 days ago
  •  ...lives. Ready to build the next breakthrough? Join us to start Caring. Connecting. Growing together.We are seeking a Principal Site Reliability Engineer (SRE) to define and scale reliability practices across large-scale cloud platforms.This is a senior individual contributor... 
    Principal
    Minimum wage
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work

    UnitedHealth Group

    Minnetonka, MN
    1 day ago
  • $134.6k - $230.8k

     ...Connecting. Growing together.Are you passionate about reimagining operations through AI? Optum Financial is seeking a Principal Site Reliability Engineer to lead the evolution of our reliability platform by combining modern SRE practices with AI-assisted operations. You'... 
    Principal
    Minimum wage
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work

    UnitedHealth Group

    Eden Prairie, MN
    5 days ago
  •  ...role will be supporting products and services, including Gen Insurance and Engine by Gen, delivering trusted digital and financial experiences to consumers at scale. As a Principal Site Reliability Engineer, you will define and lead the reliability, scalability, and... 
    Principal

    Gen Digital Inc.

    Tempe, AZ
    2 days ago
  • $248k - $396.75k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing, building, and operating large-scale production...  ...strengthen the reliability of production environments.As a Principal SRE, you will shape the technical direction of NVIDIA’s... 
    Principal
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...Continuous Delivery (CI/CD) pipelines and Kubernetes. Supports Site Reliability Engineering (SRE) functions by establishing Service Level Objectives...  ...equivalent) and five (5) years of experience as a Principal Site Reliability Engineer (or closely related occupation)... 
    Principal
    Full time

    Fidelity Investments

    Durham, NC
    2 days ago
  • $190k - $220k

     ...anticipating our clients’ needs and exceeding their expectations. About the RoleWe are looking for a skilled and motivated Principal Site Reliability Engineer to join our team. In this role, you will be responsible for ensuring the reliability, scalability, and performance of... 
    Principal
    Full time
    Worldwide
    Flexible hours

    FactSet Research Systems

    Norwalk, CT
    22 hours ago
  • $84.9k - $209.5k

    Designs and architects infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity needs...  ...new tools and develops and maintains advanced knowledge of site reliability trends.Only Oracle brings together the data, infrastructure... 
    Principal
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle Corporation

    North Carolina
    3 days ago
  •  ...Principal Site Reliability Engineer The Principal Site Reliability Engineer will be a critical technical leader responsible for driving the operational excellence, resilience, and security of our core systems for a key Randstad client in the Washington D.C. area. This... 
    Principal

    Software Technology Inc

    Washington DC
    3 days ago
  • $84.9k - $209.5k

    This role combines strategic architecture with practical systems engineering, deployment, automation, patching, troubleshooting, incident response, and compliance support. The Principal Site Reliability Engineer will work across Windows, Linux, Oracle Cloud Infrastructure... 
    Principal
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    1 day ago
  • $200k - $250k

     ...build the future together. The Crown Is Yours As a Principal Site Reliability Engie r , you'll shape the long-term strategy for the infrastructure...  ...direction of our cloud and on-premise platforms, helping engineering teams build, deploy, and operate highly reliable systems... 
    Principal
    Full time
    Immediate start

    DraftKings

    Boston, MA
    1 day ago
  • $169.3k - $304.7k

     ...building and maintaining fast, efficient, scalable, and reliable routing software and infrastructure that is responsible...  ...the growth and stability of our global platform. As a Principal Site Reliability Engineer - Network, you will be responsible for: Architecting... 
    Principal
    Work experience placement
    Work at office
    Remote work

    Akamai

    Eastern, KY
    2 days ago
  •  ...Principal Site Reliability Engineer Are you passionate about enabling next-generation server hardware and Linux platforms while building reliable, high-performance infrastructure for a global cloud platform? Are you excited about solving complex server reliability... 
    Principal
    Work at office
    Remote work

    Akamai

    United States
    4 days ago
  • $84.9k - $209.5k

     .... You’ll partner with customer support, service owners, and engineering teams around the globe to ensure high-quality service for customers...  ...posted.Career Level - IC4Escalation points for junior site reliability engineers during complex or high-impact incidents.Manage and... 
    Principal
    Temporary work
    Monday to Friday
    Flexible hours
    Shift work
    Night shift

    Oracle Corporation

    Reston, VA
    5 days ago
  • $198.24k - $272.58k

    We’re looking for a Principal Site Reliability Engineer to join Procore’s Compute Division to work on our FedRAMP initiative. In this role, you’ll help build Procore’s next-generation construction compute platform for others to build upon, including Procore developers,... 
    Principal
    Full time
    Contract work
    Work at office
    Local area
    Immediate start

    Procore Technologies

    Austin, TX
    22 hours ago
  •  ...Infrastructure Code. Builds reliability into the ecosystem by applying...  ...practices in resiliency engineering and observability by developing...  ...engineering techniques with site reliability engineering...  ...5) years of experience as a Principal Site Reliability Engineer (or... 
    Principal
    Full time

    Fidelity Investments

    Westlake, OH
    3 days ago
  • $159k - $272k

     ...generosity. Join us for the opportunity to grow and make a difference in ways that matter to you. Role SummaryIn this role as Principal Site Reliability Engineer, Infrastructure Observability you will help formulate, develop, and implement a team of Site Reliability Engineers (... 
    Principal
    Full time
    Private practice
    Local area
    Remote work
    Work from home
    3 days per week

    T. Rowe Price

    Owings Mills, MD
    1 day ago
  •  ...Principal Site Reliability Engineer Deimos is a cloud-native developer and security operations technology services company. We help companies of all sizes adopt the cloud for improved service delivery to their clients. We're a fully remote African-based team of engineers... 
    Principal
    Currently hiring
    Remote work
    Work from home

    Deimos

    United States
    2 days ago
  • $194k - $237k

    ## Principal Site Reliability EngineerApplylocations: Scottsdaletime type: Full timeposted on: Posted 5 Days Agojob requisition id: REQ2026426At...  ....**Overall Purpose**The Principal Site Reliability Engineer partners with development teams by designing availability... 
    Principal
    Hourly pay
    Work at office
    Immediate start
    Visa sponsorship
    Work visa
    Flexible hours

    Early Warning Services

    Eastern, KY
    5 days ago
  • $96.3k - $264.1k

     ...infrastructure and service, ensuring alignment with reliability and functionality standards. Takes full...  ...tools and provides expertise in site reliability trends.Only Oracle brings...  ...LeadershipDefine and drive the site reliability engineering strategy for large-scale, distributed,... 
    Principal
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    1 day ago
  • $165k - $185k

     ...the work is personal, and we are committed to the cause. Learn more at tandemdiabetes.com A DAY IN THE LIFE: The Principal Site Reliability Engineer (SRE) is responsible for the reliability, availability, and performance of the company's production systems. This... 
    Principal
    Permanent employment
    Contract work
    Local area
    Remote work
    Flexible hours
    Shift work

    Tandem Diabetes

    United States
    2 days ago
  •  ...Principal Site Reliability Engineer LivePerson transforms customer care from voice calls to mobile messaging. Our cloud-based software platform, LiveEngage, allows brands with millions of customers and tens of thousands of care agents to deliver digital experiences... 
    Principal
    Local area
    Remote work

    LivePerson

    United States
    5 days ago
  • $132.6k - $214.5k

     ...As part of this role, you will collaborate closely with our engineering teams to develop innovative solutions that provide clear and...  ...team to influence the operability of the product and ensure the reliability and availability of our services. Qualifications DevOps... 
    Principal
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks

    Santa Clara, CA
    2 days ago
  • $272k - $431.25k

    NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning and Process) Cloud Infrastructure Team. IPP is a global organization within NVIDIA. This group works with various other groups within NVIDIA such as Graphics... 
    Principal
    Full time
    Work experience placement
    Worldwide

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...Principal Site Reliability Engineer Your work will help clinicians and healthcare staff access applications consistently, reduce disruption from platform changes, and accelerate recovery when problems occur. You will help the team build a platform that is easier to... 
    Principal
    Temporary work
    Remote work
    Flexible hours

    hackajob

    United States
    4 days ago
  •  ...with software development teams to build reliable, scalable, secure, and cloud-native...  ...influence scalable architecture patterns across engineering teams, helping ensure systems are...  ...~8+ years of hands-on experience in Site Reliability Engineering, DevOps, cloud infrastructure... 
    Principal
    Remote work

    ABC Fitness Solutions, LLC

    United States
    5 days ago
  • $152.3k - $246.4k

     ...Palo Alto Networks runs a large infrastructure and is one of the largest Google Cloud Platform customers. As a Principal Site Reliability Engineer for the ADEM (Autonomous Digital Experience Management) team, you will be part of a team supporting the services that... 
    Principal
    Full time
    Work at office
    Visa sponsorship
    Work visa

    PaloAlto Networks

    California
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Site Reliability Engineer. Be the first to apply!