Principal Site Reliability Engineer
Arizona Staffing
Principal Site Reliability Engineer
As a Principal Site Reliability Engineer (IC4), you will be responsible for designing, building, and operating highly available, scalable, secure, and resilient cloud services. You will combine software engineering with infrastructure expertise to improve service reliability, operational efficiency, and developer productivity across large-scale distributed systems.
You will lead complex reliability initiatives, drive automation-first operational practices, and develop software solutions that eliminate manual toil. You will partner closely with software engineering, cloud infrastructure, security, and product teams to architect resilient platforms that meet aggressive availability, scalability, and performance objectives.
Success in this role requires deep expertise in distributed systems, cloud infrastructure, coding, automation, observability, incident management, and operational excellence. You will leverage modern AI technologies, machine learning, and intelligent automation to streamline operations, accelerate incident response, improve troubleshooting, and enable autonomous system management.
You are expected to be a technical leader who influences architecture, establishes engineering best practices, mentors other engineers, and drives continuous improvements across multiple services and organizations.
Responsibilities
Reliability Engineering & Service Ownership
+ Design, build, and operate highly available, scalable, and fault-tolerant cloud services that meet defined Service Level Objectives (SLOs) and Service Level Agreements (SLAs).
+ Lead architecture reviews to improve resiliency, scalability, observability, and operational readiness.
+ Forecast infrastructure growth, capacity requirements, and resource utilization while proactively mitigating operational risks.
+ Continuously improve platform reliability through engineering solutions rather than manual operational processes.
+ Define reliability standards, operational best practices, and service readiness criteria across multiple engineering teams.
Software Engineering, Automation & AI
+ Design and develop production-quality software, automation frameworks, and internal platforms using Python, Java, Go, or similar programming languages.
+ Build scalable automation to eliminate repetitive operational work, reduce manual intervention, and improve engineering productivity.
+ Develop APIs, microservices, and tooling that simplify infrastructure management, deployment, monitoring, and operational workflows.
+ Leverage Generative AI, Large Language Models (LLMs), AI agents, and intelligent automation to:
+ Automate routine operational tasks and runbooks.
+ Accelerate incident triage and root cause analysis.
+ Improve log analysis and anomaly detection.
+ Generate operational insights and recommendations.
+ Automate knowledge management and operational documentation.
+ Enhance developer productivity and self-service capabilities.
+ Identify opportunities to incorporate AI-driven operational intelligence into existing systems to improve efficiency, reliability, and scalability.
Infrastructure Engineering
+ Design and optimize cloud infrastructure supporting distributed services across multiple regions and availability domains.
+ Improve system resiliency through redundancy, automation, and infrastructure-as-code.
+ Build and maintain deployment pipelines, provisioning frameworks, and configuration management solutions.
+ Drive infrastructure standardization and platform modernization initiatives.
Observability & Operational Excellence
+ Design comprehensive monitoring, logging, tracing, and alerting strategies.
+ Build meaningful dashboards, health reporting, and service performance metrics.
+ Improve alert quality, reduce operational noise, and enhance system visibility.
+ Define and measure Service Level Indicators (SLIs), SLOs, and error budgets.
+ Continuously optimize operational processes using data-driven insights.
Incident Management & Reliability
+ Lead critical production incident response and act as a senior escalation point during major service events.
+ Drive root cause analysis, corrective actions, and post-incident reviews with a focus on long-term engineering improvements.
+ Develop automated remediation solutions to reduce Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR).
+ Continuously improve operational readiness through disaster recovery testing, game days, and failure injection exercises.
Performance & Scalability
+ Identify performance bottlenecks across applications, infrastructure, storage, databases, and networking.
+ Drive optimization initiatives to improve system throughput, latency, efficiency, and cost.
+ Perform capacity planning and predictive scaling using historical trends and telemetry data.
+ Optimize resource utilization while maintaining service reliability and customer experience.
Technical Leadership
+ Serve as the technical leader for reliability engineering initiatives spanning multiple services and organizations.
+ Establish engineering standards, architectural patterns, and operational best practices.
+ Review system designs and influence technical decisions across engineering organizations.
+ Mentor engineers on software engineering, distributed systems, cloud technologies, operational excellence, and automation.
+ Drive engineering excellence through code reviews, design reviews, technical guidance, and knowledge sharing.
Innovation & Continuous Improvement
+ Evaluate emerging cloud technologies, AI capabilities, automation frameworks, and observability platforms.
+ Drive adoption of modern engineering practices including GitOps, Infrastructure as Code, continuous delivery, and policy-as-code.
+ Identify opportunities to simplify architecture, eliminate operational complexity, and improve platform reliability.
+ Champion engineering initiatives that increase service scalability, operational efficiency, and customer satisfaction.
Core Competencies
Planning & Execution
+ Lead complex, cross-functional technical initiatives from design through production deployment.
+ Balance reliability, scalability, performance, cost, and delivery priorities across multiple concurrent projects.
+ Drive execution with minimal direction while effectively managing technical risk and dependencies.
Collaboration & Influence
+ Partner with software engineering, cloud infrastructure, security, networking, and product organizations to deliver reliable services.
+ Influence technical direction across organizations through strong communication and technical leadership.
+ Build consensus among stakeholders with differing priorities and technical perspectives.
Problem Solving
+ Solve highly ambiguous and complex technical problems involving distributed systems at cloud scale.
+ Apply systematic debugging and data-driven analysis to identify root causes and implement long-term engineering solutions.
+ Make sound technical decisions using engineering judgment and operational experience.
Continuous Learning
+ Stay current with advances in cloud computing, distributed systems, AI, software engineering, cybersecurity, and Site Reliability Engineering.
+ Evaluate emerging technologies and drive adoption where they provide measurable operational value.
+ Foster a culture of continuous learning, experimentation, and technical excellence.
Talent Development
+ Mentor engineers and contribute to their technical growth through coaching and knowledge sharing.
+ Participate in hiring, interviewing, and technical assessments to build high-performing engineering teams.
+ Lead by example through engineering excellence, ownership, and customer-focused decision making.
$86.25k - $158.13k
...our employees feel respected, valued and have an opportunity to contribute to the company’s success. As a Software Engineer Lead within PNC's Site Reliability Engineering Center (SRC), you will be based in one of PNC's IT Hubs: Cleveland, Ohio; Pittsburgh, Pennsylvania;...SuggestedFull timeTemporary workPart timeWork experience placementWork at office$141k - $208k
...be a part of our journey! About the role We are committed to providing our customers with reliable and secure services so we are expanding our central Site Reliability Engineering team. You will be responsible for building and leading processes to ensure the reliability...SuggestedLocal areaRemote workHome officeFlexible hours$170k - $240.75k
...that reflect who they uniquely are.We are seeking a Senior Principal Software Engineer to lead the design and development of next-generation Agentic... ...to Diversity, Equity, and Inclusion on our Career Site.The compensation package for this role is based on multiple...PrincipalRemote work- ...Site Reliability Engineer Hybrid Onsite Worker is required to work onsite 2-3 days per week in Phoenix, AZ OR Plano, TX Main Responsibilities ~ Experience in leading observability initiatives as lead engineer. ~ Development and implementation of build release...SuggestedWork experience placement2 days per week3 days per week
- ...scale, we invite you to bring your talents to Zscaler to help shape the future of cybersecurity. Role We are looking for a Site Reliability Engineer-SkillBridge Intern (San JosA Ca or Bellevue WA) to join our Zero Trust Exchange team. This is a remote role based in San...SuggestedInternshipWork at officeLocal areaRemote workWorldwide
$16 per hour
...top American bank holding company and financial services corporation to run a Paid 13 week boot-camp to learn how to be a Site Reliability Engineer! What is an SRE? Location: Phoenix, AZ OR Pittsburgh, PA (MUST BE LOCAL) Duration or contract: 12 Month contract...Contract workLocal areaRemote work- ...customers globally. As part of the ongoing investment in the reliability and modernization of these systems, client is making a multi-... ...delivery of change without impact to reliability. Integration Engineering, ensuring that client is consuming the right infrastructure...Work experience placement3 days per week
- ...National Pavement Partners, Inc. is seeking an Operations Manager based in Phoenix, Arizona. You will be responsible for overseeing on-site activities for construction projects, managing daily operations and ensuring projects are delivered on time and within budget. The...
- Job Title:Principal Engineer I - Senior DevOps EngineerLocation:Block 23What you'll do:As a Principal Engineer I you'll provide SME expertise... ...domains to ensure solutions are safe, secure, compliant and reliable. You'll identify development and support needs as well as...PrincipalFull time
- Job Title:Principal Engineer II - Digital Banking PlatformLocation:OH - ColumbusWhat you'll do:As a Principal Engineer, you will serve as... ...strong observability, operational support, and production reliability practices.Strong knowledge of system integration patterns,...PrincipalFull time
$65k - $187.2k
...where all of our employees feel respected, valued and have an opportunity to contribute to the company’s success. As a(n) Software Engineer Principle within PNC's Retail Management Information Systems Technology and Product Management organization, you will be based in...PrincipalFull timeTemporary workPart timeWork experience placementWork at office$112k - $249.6k
...opportunity to contribute to the company’s success. As a Software Engineer Principal within PNC's Technology organization, you will be based in... ...exposure across enterprise source code.You will own the reliability, support, and continued development of the platform,...PrincipalFull timeTemporary workPart timeWork experience placementWork at office- Job Title:Principal Security Platform EngineerLocation:Block 23What you'll do:As a Principal Engineer I you'll provide SME expertise in your respective domain as well as adjacent... ...are safe, secure, compliant, and reliable. You'll identify development and support needs...PrincipalFull timeWork at officeFlexible hoursShift work
$114.6k - $234.6k
...creates durable fixes and preventive controls. Designs performance, reliability, and fault-tolerance improvements for drivers, services, and... ...development lifecycle; provides guidance and coaching to engineers to drive improvements. Utilizes advanced knowledge to develop...PrincipalTemporary workFlexible hoursShift work$250.6k - $362.6k
...comprehensive security outcomes, as a Principal Engineer. The team delivers secure, scalable capabilities... ...networking, with a strong emphasis on reliability, interoperability, and long-term... ...insurance. Please see the Cisco careers site to discover more benefits and perks....PrincipalFull timeTemporary workLocal areaRemote workFlexible hours- ...relentlessly on delivering unrivaled digital products that drive a more reliable and profitable airline.The Software domain refers to the area... ...impactDelivers high quality work and coaches more junior engineers on technical craftsmanshipConducts root cause analysis to...PrincipalFlexible hours
$231.4k - $331.8k
...visibility and intelligence across diverse environments. As a Principal Engineer, you will be at the forefront of advancing Tetragon’s... ...coverage, and basic life insurance. Please see the Cisco careers site to discover more benefits and perks. Employees may be eligible...PrincipalFull timeTemporary workLocal areaFlexible hours- ...users across our industrial-scale Software-as-a-Service observability platform. We are seeking experienced engineers who can craft intuitive, fast, and reliable user experiences empowering software engineers to diagnose and resolve incidents with confidence. You will collaborate...PrincipalFull timeRemote workFlexible hours
- Job Title:Principal Engineer I - Azure Cloud EngineerLocation:OH - ColumbusWhat you'll do:The Cloud Principal Engineer is a senior technical... ...Cloud Operations and service owners to improve platform reliability, observability, recoverability, and performance, while keeping...PrincipalFull timeImmediate start
- ...Principal Systems Safety Engineer Location: Phoenix, AZ Onsite Day 1. As a Principal Systems Safety Engineer here at Honeywell, you will play a critical role in ensuring the safety and reliability of our systems. You will be responsible for developing and implementing...Principal
- ...Senior Director, Principal Gifts About the Company Philanthropic organization supporting Indigenous culture & individuals Industry Non-Profit Organization Management Type Non Profit Founded 2017 Employees 11-50 Categories...Principal
- Participate in mock regulatory examinations and compliance program reviews. Assist with policy and procedure reviews. Draft reports and manuals for compliance. Provide compliance assistance to clients including private funds and registered products. Take a leading role...Principal
- Job Title:Principal Engineer I - AI Full Stack EngineerLocation:Block 23What you'll do:As a Principal Engineer I, you will serve as a senior... ...controls.Contribute to observability, performance, and reliability practices to ensure Lending and Card systems meet operational...PrincipalFull time
$60 - $65 per hour
...Principal Systems Engineer Immediate need for a talented Principal Systems Engineer. This is a 06 months contract opportunity and is located in Phoenix, AZ (Hybrid). Pay Range: $60/hr - $65/hour. Employee benefits include, but are not limited to, health insurance...PrincipalContract workLocal areaImmediate start$114.6k - $234.6k
...automation and AIOps systems. Lead end-to-end system design for scalability, reliability, and observability. Stay hands-on with coding, debugging, and production delivery. Drive engineering excellence through code reviews and best practices. Mentor engineers and...PrincipalFull timeTemporary workRemote workFlexible hours- ...Charlotte, NC, Phoenix, AZ and Remote is not an option. The Principal Software Engineer serves as a technical authority and innovation leader,... ...States. For an overview of our benefits, visit our Careers site - . Some job boards have started using jobseeker-reported...PrincipalLocal areaRemote workFlexible hours
$163.85k - $185k
...perspectives because we believe people power progress. Join us as a Principal DevOps Engineer and help us do what we do best: propelling business... ..., test, and release workflows Optimize build performance, reliability, and scalability Security & Compliance Integrate security...PrincipalWork at officeLocal areaRemote workWork from homeRelocationHome officeFlexible hours$168k - $240k
...realize their desired business outcomes. We accelerate the growth of more impactful work and the evolution of Slalom.The Role: M&A Principal/Senior PrincipalWhat You’ll Do:* Delivery areas include:* Executing operational due diligence* Creating integration strategies,...PrincipalTemporary workWork at officeLocal areaImmediate start- ...Clarkston seeks a Supply Chain Planning Principal – Life Sciences to drive end-to-end planning transformations for global Life Sciences organizations. You will partner with directors and executives to shape strategy and deliver S&OP, IBP, and related processes. You will...Principal
$250.6k - $362.6k
...comprehensive security outcomes, as a Principal Engineer. The team delivers secure, scalable capabilities... ...networking, with a strong emphasis on reliability, interoperability, and long-term... ...insurance. Please see the Cisco careers site to discover more benefits and perks....PrincipalFull timeTemporary workPart timeLocal areaRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Site Reliability Engineer. Be the first to apply!
- chief engineer Phoenix, AZ
- engineering director Phoenix, AZ
- director of electrical engineering Phoenix, AZ
- senior director engineering Phoenix, AZ
- director data engineering Phoenix, AZ
- senior chief engineer Phoenix, AZ
- hotel chief engineer Phoenix, AZ
- principal developer Phoenix, AZ
- civil engineer project manager Phoenix, AZ
- principal engineer Phoenix, AZ


