Principal Site Reliability Engineer
Prosum
Principal Site Reliability Engineer
Position Overview
We are seeking an experienced Principal Site Reliability Engineer (SRE) to provide technical leadership across highly available, large-scale production environments. This role combines software engineering, systems engineering, cloud infrastructure, automation, DevOps, and production reliability to improve the resilience, scalability, performance, observability, and operational health of critical services.
The Principal SRE will partner closely with Software Engineering, Platform Engineering, Cloud Infrastructure, DevOps, and other technology teams to ensure reliability and operational readiness are incorporated throughout the software development lifecycle.
This is a senior individual contributor position with enterprise-level influence. The successful candidate will identify systemic reliability risks, establish technical direction, influence architecture and engineering practices, and help improve reliability capabilities across multiple engineering teams.
Key Responsibilities
- Apply Site Reliability Engineering (SRE), software engineering, automation, and DevOps principles to improve how production services are built, tested, deployed, monitored, operated, and recovered.
- Establish and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), error budgets, availability metrics, and service-health measurements.
- Design and enhance observability capabilities using metrics, logging, distributed tracing, monitoring, alerting, dashboards, and service-health instrumentation.
- Drive continuous improvement across CI/CD pipelines, Infrastructure as Code (IaC), cloud infrastructure, deployment practices, automation, testing, incident management, capacity planning, resilience, disaster recovery, and operational readiness.
- Analyze production environments to identify systemic reliability risks, performance bottlenecks, recurring incidents, and opportunities for automation.
- Translate production and operational experience into improvements in application code, architecture, infrastructure, tooling, automation, and engineering standards.
- Partner with Software Engineering teams to incorporate reliability, resiliency, scalability, performance, observability, recoverability, and operational readiness throughout the development lifecycle.
- Lead or participate in production incident response, troubleshooting, root cause analysis, service restoration, and blameless post-incident reviews.
- Provide technical leadership during critical production incidents and help improve incident response, escalation procedures, service restoration, and sustainable on-call practices.
- Reduce operational toil and manual intervention through software development, scripting, automation, reusable tooling, platforms, and engineering patterns.
- Apply data-driven analysis, experimentation, and engineering principles to validate assumptions and guide technical decisions.
- Establish and influence enterprise-level SRE, DevOps, cloud, reliability, and operational engineering standards and best practices.
- Mentor engineers and technical leaders while promoting knowledge sharing and sustainable engineering capabilities across the organization.
- Operate independently across complex, business-critical reliability and infrastructure challenges.
Required Qualifications
- 15+ years of relevant professional experience in one or more of the following areas:
- Site Reliability Engineering (SRE)
- Software Engineering
- Systems Engineering
- Cloud Engineering
- Platform Engineering
- DevOps Engineering
- Infrastructure Engineering
- Systems Architecture
- Strong experience with software development and/or scripting using one or more modern programming languages.
- Advanced understanding of software engineering principles, distributed systems, production environments, troubleshooting, automation, and observability.
- Experience designing, supporting, or improving highly available, scalable production systems and distributed applications.
- Experience with public cloud platforms and cloud-native architectures, preferably AWS.
- Strong knowledge of Linux/Unix systems, networking, infrastructure, application architecture, and production operations.
- Demonstrated experience diagnosing complex production issues and implementing sustainable technical solutions.
- Strong analytical, troubleshooting, problem-solving, communication, and cross-functional collaboration skills.
- Ability to provide technical direction and influence engineering practices across multiple teams and organizational boundaries.
Preferred Qualifications
- Extensive hands-on experience with Amazon Web Services (AWS) or another major cloud platform, including Microsoft Azure, Google Cloud Platform (GCP), or Oracle Cloud Infrastructure (OCI).
- Experience with CI/CD pipelines and software delivery automation.
- Experience with Infrastructure as Code (IaC) technologies and practices.
- Experience with containers and container orchestration technologies.
- Strong experience with monitoring, logging, distributed tracing, dashboards, alerting, and observability platforms.
- Experience defining and managing SLIs, SLOs, error budgets, availability targets, and reliability metrics.
- Experience with incident management, root cause analysis, performance engineering, capacity planning, resilience testing, disaster recovery, and operational readiness.
- Experience creating reusable automation, tooling, platforms, frameworks, engineering patterns, or standards that improve engineering productivity and system reliability.
- Experience influencing architecture and technical strategy for large-scale or business-critical production systems.
- Demonstrated ability to mentor senior engineers and improve technical capabilities across engineering organizations.
- Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, Information Systems, or a related technical discipline, or equivalent practical experience.
Key Technical Skills / ATS Keywords
Site Reliability Engineering (SRE), AWS, Cloud Computing, DevOps, Software Engineering, Systems Engineering, Platform Engineering, Distributed Systems, Production Reliability, High Availability, Scalability, Resilience, Observability, Infrastructure as Code (IaC), CI/CD, Automation, Linux, Unix, Networking, Containers, Container Orchestration, Monitoring, Logging, Distributed Tracing, Alerting, SLIs, SLOs, Error Budgets, Incident Management, Root Cause Analysis, Production Support, Performance Engineering, Capacity Planning, Disaster Recovery, Resilience Testing, Operational Readiness, Cloud Architecture, Production Operations, Software Development, Scripting, Troubleshooting, Technical Leadership.
$194k - $237k
...employer, at the date of hire. This position is ineligible for employment Visa sponsorship. Overall Purpose The Principal Site Reliability Engineer partners with development teams by designing availability and resiliency patterns in applications and infrastructure....PrincipalHourly payWork at officeImmediate startVisa sponsorshipWork visaFlexible hours$106k - $130k
...for any employer, at the date of hire. This position is ineligible for employment Visa sponsorship.Role Summary The Senior Site Reliability Engineer applies software engineering and systems engineering practices to improve the reliability, resilience, scalability, and...SuggestedHourly payFull timeImmediate startVisa sponsorshipWork visaFlexible hours$140k - $160k
...Overview Senior Site Reliability Engineer Type: Full-time | Hybrid | Scottsdale Base pay range: $140,000.00/yr - $160,000.00/yr Salary: $140,000 - $150,000 + 10% bonus About Our Client Our client is a rapidly growing technology company at the forefront...SuggestedFull time- ...Early Warning Services, LLC is seeking a Principal Site Reliability Engineer who partners with development teams to design availability and resiliency patterns for applications and infrastructure, ensuring scalable, secure production systems. You will lead cross‑functional...Suggested
- Job Title Good understanding of Production Support, Tools & Automation with 5+ years of Experience Requires knowledge using AppDynamics and APM Solutions to monitor application performance & infrastructure and aide in troubleshooting Experience on GCP, Microservices...Suggested
- ...Purple Drive Site Reliability Engineer Share Contractual Scottsdale, AZ Overview: Site Reliability Engineer Experience: ~3–5 years in Service Reliability/Operations managing large-scale, high-performance hybrid applications (on-prem + cloud). ~2–4 years...
- ...healthcare technology company in Scottsdale, Arizona, is seeking a Principal Software Engineer . AdviNOW builds an AI-driven medical automation... ...patient intake, diagnosis and documentation. This is an on-site, full-time role at the company's Scottsdale office. For...PrincipalFull timeWork at office
$140k - $210k
...Our Mission As the world’s number 1 job site*, our mission is to help people get jobs. We strive to cultivate an... ...Comscore, Total Visits, March 2026) Day to Day As an Engineering Manager in Site Reliability Engineering at Indeed, you will manage and grow a team that...Work experience placementLocal area- ...the standard for how arrivia builds with AI agents, and stay hands-on enough to build it yourself.About the RoleAs a Principal Agentic Software Engineer, you are a hands-on, product-minded, customer-focused full-stack engineer who sets technical direction across multiple...PrincipalFull timeTemporary workRemote work
$184k - $230k
...any employer, at the date of hire. This position is ineligible for employment Visa sponsorship.Overall Purpose The Principal Platform Security Engineer is a hands-on enterprise technical leader responsible for defining the long-term technical vision, secure target states...PrincipalHourly payFull timeImmediate startVisa sponsorshipWork visaFlexible hours$182.14k - $202.06k
...degree in Electrical or Computer Engineering, or a related Science,... ...this Position As a Senior Principal ASIC FPGA Engineer, on the Space... ...of specifications and reliability• May be a book boss on a medium... ..., flex or onsite.While on-site, you will be a part of the Scottsdale...PrincipalFor subcontractorWork at officeRemote workRelocation packageFlexible hours- ...Collaborate withinERM’smultidisciplinaryWesternRockiesteam,includingenvironmentalplanners,biologists,wetlandspecialists,GISspecialists,engineers,andothertechnicalprofessionals.Represent ERMinmeetingsandcommunicationswithclients,agencyrepresentatives,teamingpartners,...PrincipalFull timeFor subcontractor
- ...Year of ServiceFertility Assistance ProgramFour-Week Company-Paid Sabbatical Eligibility After Five Years of ServiceRyan is seeking Principal level talent in our Sales and Use Tax Consulting Practice in the western U.S. We do not have a location preference. Any major city...PrincipalFull time
- West Pharmaceutical Services seeks a Principal Specialist, Regulatory Affairs (Medical Devices) to drive regulatory strategies and submissions worldwide. The role partners with QA, R&D and global teams to ensure compliant, timely approvals across US, EU and other markets...PrincipalWorldwide3 days per week
- ...Principal Data Architect The Principal Data Architect will be part of a dynamic team that builds and implements data products serving... ...for a given domain, providing leadership and oversight to engineers and application architects who are intimately familiar with the...Principal
- ...Release Engineer We are looking for an experienced and passionate Release Engineer to join our team. As a Release Engineer, you will be responsible for ensuring products can effortlessly be delivered to users and customers using different distribution mechanisms and...
- Principal Systems Architect, Communications Hardware (Remote Eligible | Relocation Assistance Available)United StatesJoin Axon and be... ...will define and drive the technical standards, architecture, and engineering processes that shape how our mission-critical hardware is...PrincipalRemote jobFull timeContract workWork experience placementRelocationRelocation package
- ...at a company where you matter.Your Impact As a Senior Software Engineer on the Enterprise New Markets team, you will lead the design... ...design decisions, accelerate development, and improve system reliability. As a senior engineer, you are expected to raise the technical...Work at officeRemote work
$108.2k - $144.8k
...Transition period to the new role should be discussed per role and based on regional differencesERM is seeking a highly motivated Principal Technical Consultant, Paleontologist, to join our global consulting firm as part of our Cultural and Paleontological Resources Services...PrincipalFull timeFixed term contractCasual workWork at officeFlexible hours$123.5k - $183.7k
...each other, and our communities.Job Summary:As a Senior Software Engineer (CL6) on the Credit Platform, you will design and operate Java... ...of credit assetsOptimize BigQuery performance, cost, and reliability using partitioning, clustering, and query tuningBuild and operate...Full timeWork at officeLocal areaImmediate startFlexible hours$124k - $174k
...Bachelor's degree in a STEM field, such as Chemistry, Chemical Engineering, or a related discipline.8+ years of functional experience in developing... ...this request form should you require accommodation.For additional Colgate terms and conditions, please click here.#LI-On-sitePrincipalHourly payLocal areaRelocation- ...Senior Software Engineer Scottsdale, Arizona, United States Position Summary: As a Software Engineer, you will work on the back-end services and APIs of our products in a challenging, fast-paced environment. You will be helping to bridge engineering best practices...
- ...our N. Scottsdale office ** Who are we looking for? Choice Hotels has an exciting new opportunity as our Senior Software Engineer in the SkyTouch Technology division . SkyTouch Technology, is an independently operated division of Choice Hotels that provides...Work at officeWorldwide
$117.38k - $168.29k
...infrastructure is at a turning point. As a Principal Consultant, Environmental FERC Project... ...) in Environmental Studies, Planning, Engineering, Geology, or related field.Prior consulting... ...Proven ability to manage complex, multi-site projects on time and within budget....PrincipalFull timeFixed term contractCasual workLocal areaFlexible hours- ...Senior Reliability Engineer Onsemi is seeking a senior reliability engineer to work in a fast-paced environment and be a key team member to... ...reliability, and effective communication skills. This is an on-site role based in Scottsdale, Arizona. Responsibilities...Full timeLocal area
- ...Role Overview We are seeking a Senior Embedded Software Engineer for Small UAV Flight Controls to architect, develop, and test... ...autonomous operations, human-machine teaming, and mission-critical applications where reliability, performance, and speed matter....Full timeImmediate startRemote work
- ...Role Overview The Platform Engineer will play a key role in managing and evolving the organization’s enterprise data platform. This position... ...of AI capabilities and ensure the platform remains secure, reliable, and optimized for analytics, business intelligence, and AI...Full time
$70 - $80 per hour
...within 60 miles of Scottsdale, AZ, 6 Month Contract to Hire Pay: $70 - $80 per hour W-2 Only - NO C2C As a Sernior Software Engineer in a cross-functional software development team with minimum of 8-10 years of professional experience , you will work closely with...Hourly payFull timeContract workH1bLocal area- ...Osaic Careers Customer Service Opportunity in Financial Services Principal Location(s): Remote Role Type: FT/Non-Exempt Salary: 70,000.00 per year + annual performance-based bonus Actual compensation offered will be determined individually, based on a number of job-related...PrincipalRemote work
$70 - $89 per hour
Scottsdale, ArizonaHybridContract$70/hr - $89/hrA financial services organization is currently looking to hire a Senior Cloud Engineer on a contract basis in the Phoenix, AZ area. This role follows a hybrid schedule with three onsite days per week and focuses on AWS cloud...Full timeContract workTemporary workCurrently hiringFlexible hours3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Site Reliability Engineer. Be the first to apply!
- chief engineer Scottsdale, AZ
- principal developer Scottsdale, AZ
- general engineer Scottsdale, AZ
- director software engineering Scottsdale, AZ
- engineering director Scottsdale, AZ
- hotel chief engineer Scottsdale, AZ
- senior civil engineer project manager Scottsdale, AZ
- principal engineer Scottsdale, AZ
- data center chief engineer Scottsdale, AZ
- senior principal cloud computing engineer Scottsdale, AZ



