Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineering Team Lead (Principal SRE)

Full-time

jobgether

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineering Team Lead (Principal SRE) based in the United States.

This Principal-level role owns the reliability, availability, and operational health of a global cloud-native AI platform.
You will combine deep hands-on SRE expertise with technical leadership, shaping reliability strategy across critical production services.
The role encompasses SLI/SLO/SLA governance, observability, automation, incident response, production readiness, and high-risk change management.
You will work closely with engineering, DevOps, platform, architecture, and operations teams to embed reliability throughout the software development lifecycle.
As a technical leader without direct reports, you will influence through expertise, mentorship, standards, and informed decision-making.
The environment is distributed and highly technical, with complex systems requiring strong availability, scalability, and operational resilience.
This opportunity is ideal for an experienced SRE professional who enjoys solving challenging infrastructure problems while building sustainable reliability practices.

Accountabilities

  • Provide technical leadership across the Site Reliability Engineering function, helping select, mentor, and develop engineers across multiple locations.

  • Establish technical direction, priorities, engineering standards, and reliability practices while contributing performance and growth feedback to team managers.

  • Own and execute a reliability roadmap covering a 2–3 quarter planning horizon.

  • Define and govern SLI, SLO, and SLA frameworks supporting contracted availability targets of up to 99.95%.

  • Design and maintain a sustainable on-call model while monitoring operational workload, page volume, and team health.

  • Serve as a Tier 2 technical escalation point for major production incidents and collaborate with incident management and operations teams.

  • Promote a blameless postmortem culture and ensure incident reviews result in actionable systemic improvements.

  • Lead Production Readiness and non-functional requirements reviews with development teams.

  • Contribute to root cause analysis and drive reliability improvements resulting from production incidents.

  • Act as an approval authority for high-risk and out-of-window production changes.

  • Define strategic direction for metrics, dashboards, alerting, SLI/SLO monitoring, escalation, and automation.

  • Drive CI/CD automation for service deployments, rollbacks, and operational processes.

  • Partner with DevOps and platform teams to evolve shared infrastructure and reliability capabilities.

  • Work with engineering managers and architects to incorporate reliability principles into the SDLC by default.

  • Participate in architecture reviews and reliability consulting while clearly communicating technical risks and reliability posture to technical and non-technical stakeholders.

Requirements

  • 8+ years of hands-on experience in Site Reliability Engineering, DevOps, cloud platforms, or closely related roles, including experience leading a team or owning a technical function.

  • Demonstrated ability to establish technical direction, maintain engineering standards, and influence teams through technical authority, with or without formal management responsibility.

  • Hands-on experience with container orchestration and service technologies such as Kubernetes, Docker, and Istio.

  • Strong experience with public cloud platforms, particularly Azure, with exposure to AWS and Google Cloud.

  • Experience with observability technologies covering metrics, dashboards, and alerting, such as Zabbix, Prometheus, and Grafana.

  • Experience designing and operating CI/CD pipelines and infrastructure-as-code solutions, including technologies such as Terraform and Flux.

  • Proficiency in at least one scripting or programming language, such as Python, Go, or Shell.

  • Strong UNIX/Linux expertise, including system configuration, performance troubleshooting, and networking fundamentals such as Layer 4/5, DNS, and TLS.

  • Strong understanding of high-availability architecture, including redundancy, failover strategies, and blast-radius management.

  • Excellent written and verbal communication skills in English, with the ability to explain complex technical concepts clearly.

  • Previous SRE leadership experience and experience managing or influencing distributed technical teams are preferred.

  • Experience with log aggregation and analytics platforms such as Loki or Thanos is preferred.

  • Familiarity with ITSM and project management tools such as Jira and Confluence is beneficial.

  • Experience in automotive, embedded systems, or other latency-sensitive production environments is advantageous.

  • Strong collaborative mindset, sound judgment under pressure, and the ability to operate effectively in ambiguous and technically complex environments.

Benefits

  • Competitive compensation and benefits package.

  • Annual bonus opportunity.

  • Medical, dental, and vision insurance coverage.

  • Life and disability insurance.

  • Paid time off and paid holidays.

  • Company contribution to an RRSP retirement savings plan.

  • Equity awards for eligible positions and levels.

  • Remote and/or hybrid work options depending on the position and location.

  • Opportunity to work on large-scale cloud-native AI and connected technology platforms.

  • Exposure to distributed engineering teams and complex global production environments.

  • Opportunities to influence technical strategy, reliability standards, and engineering practices at a Principal level.

  • A collaborative environment focused on innovation, technical growth, and continuous improvement.

How Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineering Team Lead (Principal SRE) in United States vacancy
  •  ...provider of enterprise-scale context engines capable of analyzing trillions of...  ...a highly skilled and motivated Site Reliability Engineer (SRE) to join our growing team. As an SRE at Lovelace AI, you...  ...before they impact end-users. Lead troubleshooting efforts for complex... 
    Suggested
    Full time

    Lovelace Ai

    Pittsburgh, PA
    14 hours ago
  • $175k - $215k

     ...these exciting experiences. Sr. Manager, Site Reliability Engineer provides strategic leadership across multiple SRE teams and their managers, ensuring alignment with organizational...  ...Engineering.  What You’ll Do   ~ Lead & Inspire: Provide strategic leadership for... 
    Suggested

    Disney Experiences Careers

    Orlando, FL
    4 days ago
  •  ...looking for a Senior Site Reliability Engineer to help us mature and...  ...establish the standards, and lead the transformation....  ...ID, and service principals Identify and...  ...equivalent) that give the team real signal Manage...  ...~​​6+ years in SRE, DevOps, or infrastructure... 
    Suggested
    Remote work
    Flexible hours

    Dental Intelligence

    United States
    3 days ago
  •  ...Senior Site Reliability Engineer (SRE) We are looking for a highly experienced and driven Senior Site...  ...thinking cloud development and operations team. In this role, you will contribute to...  ...on cutting-edge hardware from leading vendors. This role focuses mainly on... 
    Suggested
    Remote work

    Mirantis

    United States
    3 days ago
  • Description The Digital - Principal SRE (AI Engineer) role is a position that blends...  ..., machine learning, and reliability engineering. This professional...  ..., DevOps, and operations teams to deliver robust,...  ...Visit Huntington's Career Web Site for more details. Agency Recruiter... 
    Principal
    Work at office
    Remote work
    Work from home
    Flexible hours

    Huntington National Bank

    Columbus, OH
    2 days ago
  •  ...Job Title:  Site Reliability Engineer (Azure Government & Infrastructure) Pay Type : SALARIED EXEMPT...  ...Responsibilities The Site Reliability Engineer (SRE) for Azure Government & Infrastructure...  ...with security operations and DevSecOps teams. Proactively monitor and optimize... 
    Full time
    Remote work
    Monday to Friday

    Quzara LLC

    United States
    4 days ago
  • $100k - $180k

     ...Site Reliability Engineer (SRE) - Remote Bright Vision Technologies is a technology consulting and software...  ...and prioritization decisions. Lead incident response and resolution for production...  ...closely with application development teams to embed reliability practices early... 
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    United States
    4 days ago
  • $100k - $180k

     ...Site Reliability Engineer (SRE) Bright Vision Technologies is a technology consulting and software development...  ...and prioritization decisions. Lead incident response and resolution for...  ...closely with application development teams to embed reliability practices early in... 
    Full time
    H1b
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    United States
    3 days ago
  •  ...Setting the reliability strategy for the platform, the full-time Principal Site Reliability Engineer will define deployment and operational standards...  ...and engineering leadership Lead major incidents and...  ...0+ years in infrastructure, SRE, or platform engineering with... 
    Principal
    Full time
    Remote work

    Virtual Vocations Inc

    United States
    1 day ago
  •  ...deploy, and operate highly reliable cloud systems supporting mission...  ...centered on DevSecOps and site reliability engineering, with a strong emphasis on...  ...experience as an SRE, DevOps, reliability, infrastructure...  ...to be part of a small team with a large direct impact... 
    Permanent employment
    Remote work

    Quindar

    United States
    1 day ago
  •  ...Site Reliability Engineer (SRE) New York (Remote) We are seeking an experienced Site Reliability Engineer (SRE) with strong expertise in Dynatrace to join our growing engineering team. The ideal candidate will be responsible for ensuring the reliability, scalability... 
    Remote work

    Staffing the Universe

    United States
    14 hours ago
  • $165k - $225k

     ...existing data center. Our team of AI infrastructure...  ...with enterprise-grade reliability and compliance. Your...  ...closely with our systems engineers, network engineers, and...  ..., and ELK stack. Lead incident response, conduct...  ...Experience: 5+ years in SRE, DevOps, or infrastructure... 
    Remote work
    Flexible hours

    Moonlite

    Chicago, IL
    17 days ago
  • $60 - $80 per hour

     ...Job Title: Senior Site Reliability Engineer (SRE) - Hybrid Duration (Contract): 6 Months Client...  ...Reliability Engineer (SRE) , you will lead reliability, observability, automation...  ...engineering, operations, and product teams to improve system availability and performance... 
    Hourly pay
    Contract work

    Smart IMS Inc

    Austin, TX
    2 days ago
  • $101k - $161k

     ...prestigious awards, such as Best Engineering Team, Best Company for Diversity,...  ...Work WithWe’re looking for Site Reliability Engineers to join our...  ...as-a-Service (CVaaS) global SRE team. SREs at Arista combine...  ...chance to be drive, develop, and lead projects in any of the... 

    Arista Networks

    Santa Clara, CA
    a month ago
  •  ...Job Description Job Title: Senior AWS Site Reliability Engineer (SRE) Location: Birmingham, Alabama Type...  ...and operational readiness. Lead technical discussions and coordinate resolution...  ..., performance, and production support teams. Requirements ~8+ years of hands... 
    Contract work
    Local area

    System One

    Birmingham, AL
    4 days ago
  •  ...Overview We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud...  ..., security, and application teams.   Key Responsibilities Operate...  ...detect and prevent service degradation. Lead incident response for infrastructure-related... 
    Temporary work

    Long Finch Technologies

    Jersey City, NJ
    a month ago
  •  ...systems. Define and implement SRE practices, standards, and reliability engineering strategies . Establish and manage...  ...databases, and cloud services. Lead incident response, root-cause...  ...engineering. Partner with development teams to improve application... 

    Staffing Spot, Inc.

    Charlotte, NC
    3 days ago
  • $80 per hour

     ...Overview Essnova Solutions, Inc. is seeking an experienced Site Reliability Engineer (SRE) to support the National Energy Research Scientific...  ...service management activities. Collaborate across technical teams to identify and resolve operational bottlenecks and... 
    Hourly pay
    Full time
    Work at office
    Local area
    Shift work
    Night shift

    Essnova Solutions, Inc.

    Berkeley, CA
    16 days ago
  •  ...Job Description Job Description Site Reliability Engineer (SRE) – Application Support Diversified Services Network, Inc. (DSN) is seeking...  ...Reliability Engineer (SRE) – Application Support to join our team in their choice of our Chicago, IL or Peoria, IL office... 
    Full time
    Work at office
    Shift work
    Weekend work

    Diversified Services Network, Inc.

    Peoria, IL
    4 days ago
  •  ...tasks using scripting and tools Python Bash etc Collaborate with development infrastructure and support teams to improve system reliability Drive adoption of SRE practices like SLIs SLOs and error budgets Ensure performance optimisation capacity planning and... 
    Permanent employment
    Temporary work
    Work experience placement
    Plano, TX
    27 days ago
  • $80k - $95k

     ...curious individual to join our dynamic team supporting the company’s users,...  ...product offerings. In this role, the Site Reliability Engineer (SRE) will play a key role in maintaining resources...  ...it possible to grow, contribute, and lead—while living a life that works for you... 
    Remote work
    Visa sponsorship
    Work visa

    RANE Network

    New York, NY
    a month ago
  •  ...are seeking a highly motivated Systems Reliability Engineer (SRE) to lead the design and implementation of...  ...with the product owner(s), developer teams, and security operations teams.  What...  ...sensitive and cleared workforces. The Site Reliability Engineer (SRE) - SecOps... 
    For contractors
    Work at office
    Flexible hours

    Arkenstone Defense

    Menlo Park, CA
    9 days ago
  •  ...development, cloud infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,...  ...workflows for accuracy and reliability. Work with AWS, Azure, GCP, Kubernetes...  ...DevOps Cloud Infrastructure Site Reliability Engineering (SRE) Platform... 
    For contractors
    Remote work

    YO AI Labs

    Dallas, TX
    24 days ago
  • $70.8k - $131.4k

     ...Reuters is strengthening its Site Reliability Engineering capability to help engineering and operations teams build, operate, and improve reliable...  ...Support and maintain SRE operational tooling, including...  ...tools and knowledge to grow, lead, and thrive in an AI-enabled future... 
    Full time
    Work at office
    Local area
    Flexible hours

    Thomson Reuters

    Eagan, MN
    3 days ago
  •  ...We are seeking an experienced Site Reliability Engineer (SRE) to support and maintain production systems hosted on AWS. The role focuses on production support, incident management, monitoring, observability, troubleshooting, and improving system reliability and availability... 
    Temporary work

    2T Consulting

    Atlanta, GA
    17 days ago
  •  ...Netherlands. On behalf of Feeld , GT is looking for a Site Reliability Engineer (SRE) to join a fast-growing consumer mobile product in the online...  ...50 people distributed across Europe and the US. The team works in small, autonomous product squads, each... 
    Full time
    Remote work

    GT

    Remote
    20 days ago
  • $87.72k - $109.65k

     ...organization, apply now. We are currently seeking a Site Reliability Engineering (SRE) - Maryland, US to join our team in Baltimore, Maryland (US-MD), United States (US...  ..., and problem-solving skills. ~ Ability to lead initiatives and mentor engineers. ~ Must be... 
    Temporary work
    Work at office
    Remote work
    Flexible hours

    NTT DATA, Inc.

    Baltimore, MD
    a month ago
  • $120k - $180k

     .... You will be the first dedicated Site Reliability Engineer and own critical infrastructure end to...  ...on-premises infrastructure as the sole SRE. Architect migrations from AWS into on-...  ...infrastructure ownership at a startup or small team is important. Experience with on-... 
    Permanent employment
    Full time
    Relocation package

    Raydar

    New York, NY
    25 days ago
  • $248k - $396.75k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing...  ...environments.As a Principal SRE, you will shape the technical...  ...s AI Platform Runtime and lead reliability engineering...  ..., and influence how teams design and operate critical... 
    Principal
    Full time

    NVIDIA

    Santa Clara, CA
    20 hours ago
  • $100k - $200k

    OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about... 
    Full time

    OPPO

    Palo Alto, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineering Team Lead (Principal SRE). Be the first to apply!