Site Reliability Engineering Team Lead (Principal SRE)
jobgether
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineering Team Lead (Principal SRE) based in the United States.
This Principal-level role owns the reliability, availability, and operational health of a global cloud-native AI platform.
You will combine deep hands-on SRE expertise with technical leadership, shaping reliability strategy across critical production services.
The role encompasses SLI/SLO/SLA governance, observability, automation, incident response, production readiness, and high-risk change management.
You will work closely with engineering, DevOps, platform, architecture, and operations teams to embed reliability throughout the software development lifecycle.
As a technical leader without direct reports, you will influence through expertise, mentorship, standards, and informed decision-making.
The environment is distributed and highly technical, with complex systems requiring strong availability, scalability, and operational resilience.
This opportunity is ideal for an experienced SRE professional who enjoys solving challenging infrastructure problems while building sustainable reliability practices.
Accountabilities
Provide technical leadership across the Site Reliability Engineering function, helping select, mentor, and develop engineers across multiple locations.
Establish technical direction, priorities, engineering standards, and reliability practices while contributing performance and growth feedback to team managers.
Own and execute a reliability roadmap covering a 2–3 quarter planning horizon.
Define and govern SLI, SLO, and SLA frameworks supporting contracted availability targets of up to 99.95%.
Design and maintain a sustainable on-call model while monitoring operational workload, page volume, and team health.
Serve as a Tier 2 technical escalation point for major production incidents and collaborate with incident management and operations teams.
Promote a blameless postmortem culture and ensure incident reviews result in actionable systemic improvements.
Lead Production Readiness and non-functional requirements reviews with development teams.
Contribute to root cause analysis and drive reliability improvements resulting from production incidents.
Act as an approval authority for high-risk and out-of-window production changes.
Define strategic direction for metrics, dashboards, alerting, SLI/SLO monitoring, escalation, and automation.
Drive CI/CD automation for service deployments, rollbacks, and operational processes.
Partner with DevOps and platform teams to evolve shared infrastructure and reliability capabilities.
Work with engineering managers and architects to incorporate reliability principles into the SDLC by default.
Participate in architecture reviews and reliability consulting while clearly communicating technical risks and reliability posture to technical and non-technical stakeholders.
Requirements
8+ years of hands-on experience in Site Reliability Engineering, DevOps, cloud platforms, or closely related roles, including experience leading a team or owning a technical function.
Demonstrated ability to establish technical direction, maintain engineering standards, and influence teams through technical authority, with or without formal management responsibility.
Hands-on experience with container orchestration and service technologies such as Kubernetes, Docker, and Istio.
Strong experience with public cloud platforms, particularly Azure, with exposure to AWS and Google Cloud.
Experience with observability technologies covering metrics, dashboards, and alerting, such as Zabbix, Prometheus, and Grafana.
Experience designing and operating CI/CD pipelines and infrastructure-as-code solutions, including technologies such as Terraform and Flux.
Proficiency in at least one scripting or programming language, such as Python, Go, or Shell.
Strong UNIX/Linux expertise, including system configuration, performance troubleshooting, and networking fundamentals such as Layer 4/5, DNS, and TLS.
Strong understanding of high-availability architecture, including redundancy, failover strategies, and blast-radius management.
Excellent written and verbal communication skills in English, with the ability to explain complex technical concepts clearly.
Previous SRE leadership experience and experience managing or influencing distributed technical teams are preferred.
Experience with log aggregation and analytics platforms such as Loki or Thanos is preferred.
Familiarity with ITSM and project management tools such as Jira and Confluence is beneficial.
Experience in automotive, embedded systems, or other latency-sensitive production environments is advantageous.
Strong collaborative mindset, sound judgment under pressure, and the ability to operate effectively in ambiguous and technically complex environments.
Benefits
Competitive compensation and benefits package.
Annual bonus opportunity.
Medical, dental, and vision insurance coverage.
Life and disability insurance.
Paid time off and paid holidays.
Company contribution to an RRSP retirement savings plan.
Equity awards for eligible positions and levels.
Remote and/or hybrid work options depending on the position and location.
Opportunity to work on large-scale cloud-native AI and connected technology platforms.
Exposure to distributed engineering teams and complex global production environments.
Opportunities to influence technical strategy, reliability standards, and engineering practices at a Principal level.
A collaborative environment focused on innovation, technical growth, and continuous improvement.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
- ...provider of enterprise-scale context engines capable of analyzing trillions of... ...a highly skilled and motivated Site Reliability Engineer (SRE) to join our growing team. As an SRE at Lovelace AI, you... ...before they impact end-users. Lead troubleshooting efforts for complex...SuggestedFull time
$175k - $215k
...these exciting experiences. Sr. Manager, Site Reliability Engineer provides strategic leadership across multiple SRE teams and their managers, ensuring alignment with organizational... ...Engineering. What You’ll Do ~ Lead & Inspire: Provide strategic leadership for...Suggested- ...looking for a Senior Site Reliability Engineer to help us mature and... ...establish the standards, and lead the transformation.... ...ID, and service principals Identify and... ...equivalent) that give the team real signal Manage... ...~6+ years in SRE, DevOps, or infrastructure...SuggestedRemote workFlexible hours
- ...Senior Site Reliability Engineer (SRE) We are looking for a highly experienced and driven Senior Site... ...thinking cloud development and operations team. In this role, you will contribute to... ...on cutting-edge hardware from leading vendors. This role focuses mainly on...SuggestedRemote work
- Description The Digital - Principal SRE (AI Engineer) role is a position that blends... ..., machine learning, and reliability engineering. This professional... ..., DevOps, and operations teams to deliver robust,... ...Visit Huntington's Career Web Site for more details. Agency Recruiter...PrincipalWork at officeRemote workWork from homeFlexible hours
- ...Job Title: Site Reliability Engineer (Azure Government & Infrastructure) Pay Type : SALARIED EXEMPT... ...Responsibilities The Site Reliability Engineer (SRE) for Azure Government & Infrastructure... ...with security operations and DevSecOps teams. Proactively monitor and optimize...Full timeRemote workMonday to Friday
$100k - $180k
...Site Reliability Engineer (SRE) - Remote Bright Vision Technologies is a technology consulting and software... ...and prioritization decisions. Lead incident response and resolution for production... ...closely with application development teams to embed reliability practices early...Full timeH1bLocal areaImmediate startRemote workVisa sponsorship$100k - $180k
...Site Reliability Engineer (SRE) Bright Vision Technologies is a technology consulting and software development... ...and prioritization decisions. Lead incident response and resolution for... ...closely with application development teams to embed reliability practices early in...Full timeH1bImmediate startRemote workVisa sponsorship- ...Setting the reliability strategy for the platform, the full-time Principal Site Reliability Engineer will define deployment and operational standards... ...and engineering leadership Lead major incidents and... ...0+ years in infrastructure, SRE, or platform engineering with...PrincipalFull timeRemote work
- ...deploy, and operate highly reliable cloud systems supporting mission... ...centered on DevSecOps and site reliability engineering, with a strong emphasis on... ...experience as an SRE, DevOps, reliability, infrastructure... ...to be part of a small team with a large direct impact...Permanent employmentRemote work
- ...Site Reliability Engineer (SRE) New York (Remote) We are seeking an experienced Site Reliability Engineer (SRE) with strong expertise in Dynatrace to join our growing engineering team. The ideal candidate will be responsible for ensuring the reliability, scalability...Remote work
$165k - $225k
...existing data center. Our team of AI infrastructure... ...with enterprise-grade reliability and compliance. Your... ...closely with our systems engineers, network engineers, and... ..., and ELK stack. Lead incident response, conduct... ...Experience: 5+ years in SRE, DevOps, or infrastructure...Remote workFlexible hours$60 - $80 per hour
...Job Title: Senior Site Reliability Engineer (SRE) - Hybrid Duration (Contract): 6 Months Client... ...Reliability Engineer (SRE) , you will lead reliability, observability, automation... ...engineering, operations, and product teams to improve system availability and performance...Hourly payContract work$101k - $161k
...prestigious awards, such as Best Engineering Team, Best Company for Diversity,... ...Work WithWe’re looking for Site Reliability Engineers to join our... ...as-a-Service (CVaaS) global SRE team. SREs at Arista combine... ...chance to be drive, develop, and lead projects in any of the...- ...Job Description Job Title: Senior AWS Site Reliability Engineer (SRE) Location: Birmingham, Alabama Type... ...and operational readiness. Lead technical discussions and coordinate resolution... ..., performance, and production support teams. Requirements ~8+ years of hands...Contract workLocal area
- ...Overview We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud... ..., security, and application teams. Key Responsibilities Operate... ...detect and prevent service degradation. Lead incident response for infrastructure-related...Temporary work
- ...systems. Define and implement SRE practices, standards, and reliability engineering strategies . Establish and manage... ...databases, and cloud services. Lead incident response, root-cause... ...engineering. Partner with development teams to improve application...
$80 per hour
...Overview Essnova Solutions, Inc. is seeking an experienced Site Reliability Engineer (SRE) to support the National Energy Research Scientific... ...service management activities. Collaborate across technical teams to identify and resolve operational bottlenecks and...Hourly payFull timeWork at officeLocal areaShift workNight shift- ...Job Description Job Description Site Reliability Engineer (SRE) – Application Support Diversified Services Network, Inc. (DSN) is seeking... ...Reliability Engineer (SRE) – Application Support to join our team in their choice of our Chicago, IL or Peoria, IL office...Full timeWork at officeShift workWeekend work
- ...tasks using scripting and tools Python Bash etc Collaborate with development infrastructure and support teams to improve system reliability Drive adoption of SRE practices like SLIs SLOs and error budgets Ensure performance optimisation capacity planning and...Permanent employmentTemporary workWork experience placement
$80k - $95k
...curious individual to join our dynamic team supporting the company’s users,... ...product offerings. In this role, the Site Reliability Engineer (SRE) will play a key role in maintaining resources... ...it possible to grow, contribute, and lead—while living a life that works for you...Remote workVisa sponsorshipWork visa- ...are seeking a highly motivated Systems Reliability Engineer (SRE) to lead the design and implementation of... ...with the product owner(s), developer teams, and security operations teams. What... ...sensitive and cleared workforces. The Site Reliability Engineer (SRE) - SecOps...For contractorsWork at officeFlexible hours
- ...development, cloud infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,... ...workflows for accuracy and reliability. Work with AWS, Azure, GCP, Kubernetes... ...DevOps Cloud Infrastructure Site Reliability Engineering (SRE) Platform...For contractorsRemote work
$70.8k - $131.4k
...Reuters is strengthening its Site Reliability Engineering capability to help engineering and operations teams build, operate, and improve reliable... ...Support and maintain SRE operational tooling, including... ...tools and knowledge to grow, lead, and thrive in an AI-enabled future...Full timeWork at officeLocal areaFlexible hours- ...We are seeking an experienced Site Reliability Engineer (SRE) to support and maintain production systems hosted on AWS. The role focuses on production support, incident management, monitoring, observability, troubleshooting, and improving system reliability and availability...Temporary work
- ...Netherlands. On behalf of Feeld , GT is looking for a Site Reliability Engineer (SRE) to join a fast-growing consumer mobile product in the online... ...50 people distributed across Europe and the US. The team works in small, autonomous product squads, each...Full timeRemote work
$87.72k - $109.65k
...organization, apply now. We are currently seeking a Site Reliability Engineering (SRE) - Maryland, US to join our team in Baltimore, Maryland (US-MD), United States (US... ..., and problem-solving skills. ~ Ability to lead initiatives and mentor engineers. ~ Must be...Temporary workWork at officeRemote workFlexible hours$120k - $180k
.... You will be the first dedicated Site Reliability Engineer and own critical infrastructure end to... ...on-premises infrastructure as the sole SRE. Architect migrations from AWS into on-... ...infrastructure ownership at a startup or small team is important. Experience with on-...Permanent employmentFull timeRelocation package$248k - $396.75k
Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing... ...environments.As a Principal SRE, you will shape the technical... ...s AI Platform Runtime and lead reliability engineering... ..., and influence how teams design and operate critical...PrincipalFull time$100k - $200k
OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineering Team Lead (Principal SRE). Be the first to apply!
- site reliability engineer sre United States
- site reliability engineering manager United States
- site reliability engineer United States
- site reliability engineer remote United States
- engineering leader United States
- engineering team lead United States
- site recruiter United States
- site services specialist United States
- junior website developer United States
- official site United States



