Senior Site Reliability Engineer
$140k - $165kColorado Public Employees' Retirement Association
Senior Site Reliability Engineer
Penn Center - Denver, CO
Salary Range $140,000.00 - $165,000.00 Salary Level Experienced Position Type Full Time Job Shift Day
Senior Site Reliability Engineer
Summary of Job Responsibilities The Senior Site Reliability Engineer (SRE) is Colorado PERA's technical owner for production reliability, observability, and platform automation across a hybrid AWS, Azure and on-premises environment. This position leads the design, implementation, and ongoing operation of PERA's enterprise observability platform, drives incident management practices that reduce time to resolution, scales CI/CD delivery pipelines, and partners with Infrastructure, Application Development, Investment Technical Services and Product teams to achieve uptime for critical services. The role is accountable for turning monitoring data into measurable improvements in reliability, delivery speed, and cloud cost efficiency and serves as the senior technical voice for reliability engineering practices across the organization.
Ideal Candidate Statement The ideal candidate is a senior, hands-on reliability engineer who thinks in systems, automates and treats every incident as a chance to remove future value leakage. They are fluent across AWS and Azure cloud, Kubernetes, and IaC from build to configuration management utilizing AI tooling to speed MTTR and improve uptime by consolidating metrics, logs, and distributed trace collection into observability tools with well-defined dashboards and actionable integrations into ITSM Service Now management in an environment primarily on premises and rapidly moving to cloud platform and infrastructure services.
Essential Duties and Responsibilities Employees are held accountable for all duties of the job. Individuals must be able to perform these duties with or without reasonable accommodation.
- Observability Platform & Automation Tooling
- Lead the evaluation, selection, and enterprise rollout of PERA's observability platform, including licensing, architecture, and integration strategy.
- Design and maintain monitoring, logging, and alerting standards across AWS, Azure and on-premises infrastructure, Kubernetes workloads, APIs, services and microservices.
- Tune alerting to reduce noise and false positives while improving detection of real degradation.
- Build automation with AI-driven integrations to eliminate manual, repetitive operational work, including self-healing and auto-remediation driving improvement of SLO/KPI metrics where appropriate.
- Develop dashboards and reporting that translate technical telemetry into business-relevant reliability and cost metrics for IT leadership.
- Incident Management & Reliability Engineering
- Serve as a senior technical responder and escalation point for Priority 1/Priority 2 production incidents, driving time-to-resolution down through structured incident command practices.
- Own PERA's blameless postmortem process: facilitate root-cause reviews, document findings, and track corrective actions to closure.
- Define and maintain Service Level Objectives (SLOs) and error budgets for critical production services.
- Build and maintain runbooks, escalation paths, and on-call procedures that reduce reliance on tribal knowledge.
- Analyze incident trends to identify systemic reliability risks and prioritize remediation work.
- Architecture & Scaling
- Partner with Infrastructure, Application Development and Investment Technical Services to design for scalability, resiliency, and fault tolerance across hybrid Kubernetes cluster orchestration.
- Conduct capacity planning and load/performance analysis to proactively identify scaling risks before they affect uptime.
- Contribute reliability and observability requirements into architecture reviews for new services and major changes.
- CI/CD & Release Engineering
- Assess and scale PERA's CI/CD pipelines to support faster, safer, and more frequent deployments.
- Implement methodologies to safeguard modern hybrid DevSecOps, and Kubernetes environments through security controls in CI/CD pipelines.
- Integrate observability and automated testing gates into the CI/CD pipeline so reliability issues are caught before production.
- Cost Optimization & Vendor/Tooling Governance
- Own the business case, budget, and ongoing vendor relationship for the selected observability platform.
- Identify and implement cloud cost optimization opportunities (rightsizing, autoscaling, reserved capacity) surfaced through observability data.
- Evaluate emerging SRE/observability tooling and recommend investments aligned with PERA's hybrid technology roadmap.
Job Qualifications
- 6+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or a closely related discipline.
- Bachelor's degree in Computer Science, Information Technology, or related field preferred, or an equivalent combination of education and experience.
- Relevant certifications preferred: AWS Certified DevOps Engineer or Solutions Architect, Certified Kubernetes Administrator (CKA), HashiCorp Terraform Associate.
- Production experience with AWS and Azure across compute, networking, and managed services.
- Hands-on experience operating and troubleshooting Kubernetes in a hybrid (on-premises and cloud) environment.
- Strong Python, Powershell and Linux scripting/automation skills; ability to build tooling, not just run it.
- Experience designing and operating an enterprise observability/monitoring platforms.
- Demonstrated experience reducing MTTR/MTTD through tooling, automation, and process not headcount.
- Experience with CI/CD tooling and modern release practices.
- Working knowledge of microservices architecture and the reliability challenges specific to distributed systems.
- Experience defining and reporting on SLOs, error budgets, and uptime commitments (99.9%+ environments).
- Experience participating in an on-call rotation and leading through live production incidents.
Preferred Leadership & Strategic Experience
- Experience leading a vendor evaluation and selection process for enterprise tooling, including business case development.
- Experience establishing reliability standards or practices adopted across multiple engineering teams without direct reporting authority.
- Experience mentoring engineers on reliability, automation, or observability practices.
- Experience presenting reliability and cost metrics to IT leadership or executive audiences.
Working Conditions The physical demands described here are representative of those that must be met by an employee to successfully perform the essential functions of this job. Reasonable accommodation may be made to enable individuals with disabilities to perform the essential functions.
- Standard office environment with frequent computer operation and use of collaboration/communication tools.
- Participation in an on-call rotation, including evenings, weekends, and holidays as needed to support 99.99% uptime commitments.
- Ability to remain calm, clear, and decisive under pressure during live production incidents.
- Ability to sit for prolonged periods of time and operate standard PC equipment.
- Ability to handle stress associated with production incidents, tight deadlines, and competing priorities.
Hybrid Work Option Opportunity to work from home up to three days per week. Eligibility dependent upon factors detailed in PERA's Work from Home Policy and on-call coverage needs.
$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....SeniorWork at officeLocal areaRemote workWorldwideFlexible hours$168k - $200k
...that is passionate about creating transformative change in healthcare. What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable,...Senior- ...to be part of this journey? At UiPath's Site Reliability team, we build the platforms and... ..., powered by AI. This is a software engineering role. You will not be the person who identifies... ...What you will be doing at UiPath As a Senior Software Engineer, you will Design,...SeniorWork at officeImmediate startRemote work
$192.4k - $275.8k
...the world's most demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service Engineering disciplines... ...this is the team for you Your ImpactYou will be the most senior technical individual contributor on the team — setting the...SeniorFull timeTemporary workLocal areaFlexible hours$55k - $187k
...Not Applicable Specialism IFS - Internal Firm Services - Other Management Level Senior Associate Job Description & Summary The Opportunity As a Site Reliability Engineer - Senior Associate, you will play a pivotal role in enhancing the reliability,...SeniorFull timeH1b- ...Position: Sr. Site Reliability Engineer Length: 12+ Month contract Location: 6380 S Fiddlers Green Cir, Greenwood Village, CO... ...streaming video, broadband internet, and mobile services. As a Senior SRE Engineer in the applied AI and data science program,...SeniorContract work
- ...Principal Site Reliability Engineer Location: Denver, CO - On-Site Required Duration: 6 Months Responsible for providing the primary management, administration, support, and ongoing maintenance of production platforms within a 24x7x365 environment and data center...Night shift
$110k - $155k
...global, with headquarters in Denver, Colorado, and offices across the U.S., Canada, and India. We are seeking a Senior Site Reliability Engineer to own the reliability, scalability, performance, and operational integrity of critical production services. This role...SeniorContract workWork at officeWork from homeFlexible hours$105.6k - $145.2k
Architect the Future as our Site Reliability Engineer!Are you ready to take your skills to the next level as a self-motivated and enthusiastic Site Reliability Engineer with hands-on experience supporting multiple connected Cloud-based products? Trimble is a global technology...Ongoing contractFull timeWork at officeLocal areaWorldwide$160k - $180k
...with headquarters in Denver, Colorado, and offices across the U.S., Canada, and India. We are seeking a Principal Site Reliability Engineer to define the strategic vision and own the enterprise-wide reliability, scalability, and performance of our critical...Contract workTemporary workWork at officeWork from homeFlexible hours$135k - $155k
...tap into global manufacturing capacity.Xometry is seeking a Site Reliability Engineer II to join our Site Reliability Engineering (SRE)... ...statements and drive them to completion with guidance from senior engineers.Write clean, efficient, and well-documented code...Flexible hours- ...We are seeking an experienced Site Reliability Engineer (SRE) to join the Applied AI and Data Science program. This role focuses on deploying, monitoring, and optimizing cloud-based applications and infrastructure to ensure high availability and performance. The...
- ...Site Reliability Engineer - Greenwood Village, CO (Hybrid) About the Role: Join a forward-thinking engineering team as a Site Reliability Engineer, specializing in enterprise-scale experimentation and configuration management platforms. In this operations-focused...Contract work
- ...focusing on private cloud systems supporting 5G wireless systems. This position will focus on platform monitoring, logging, and reliability aspects supporting the Mobile Core team. A critical goal is to gather metrics of the platform during stress and load events to ensure...
$160k - $190k
...Site Reliability Engineer (Classified Deployments) Location: Southern California or Washington, D.C. Clearance: Active Secret required; TS/SCI strongly preferred Work Mode: Hybrid/On-site with government customers Citizenship: U.S. Citizen Compensation:...$98.58k - $138.02k
...Site Reliability Engineer II Restaurant365 is a SaaS company disrupting the restaurant industry! Our cloud-based platform provides a unique, centralized solution for accounting and back-office operations for restaurants. Restaurant365's culture is focused on empowering...Work at office$95k - $134k
...build. For more information, visit Job Application Deadline: 10/31/2026 The Opportunity DAT is looking for a Site Reliability Engineer to join our SRE platform team. This position will work hybrid in Denver, CO Candidate profile DAT is seeking an...Temporary workFor contractorsWork experience placementWork at officeLocal areaImmediate startFlexible hours- ...Job Description Site Reliability Engineer to be part of the Video Streaming team to build automation solutions to deploy and maintain applications, focusing on scalability, reliability and resiliency. We are looking for a dynamic team player who is an owner, accountable...
$110k - $125k
...The Site Reliability Engineer maintains and improves the availability, performance, resilience, and operational recoverability of enterprise identity... ...support repeatable operations and recovery. Working with senior engineers, identity engineers, and operations teams, the...Contract workWork at office- ...Site Reliability Engineer Location: Greenwood Village, Colorado (Hybrid) Role Overview As a Senior Reliability and Optimization Engineer, you will hold responsibility for the reliability of custom, enterprise-scale experimentation and configuration management...
$146k - $234k
..., get crucial goods where they need to go, and make mobility more efficient and accessible for all. We’re searching for a Senior Software Engineer - Camera Systems. In this role, you will Develop software for hardware calibration, benchtop testing, and sensor data...SeniorFull time$120k - $160k
...Manager, Site Reliability Engineering Ready to Help Shape the Future of Legal Tech?! At Litera, we don't just build software, we transform how the world's top law firms operate. Every day, we Raise The Bar™ for what's possible through AI, innovation, and solutions...Work experience placementWork at officeWorldwide3 days per week$175k - $220k
...is global, with headquarters in Denver, Colorado, and offices across the U.S., Canada, and India. The Director, Site Reliability Engineering (SRE) will lead reliability, performance, and observability initiatives for a portfolio of Vertafore products. This role...Contract workTemporary workWork at officeWork from homeFlexible hours$150k - $200k
...Are you looking to build tools that engineers actually keep in their workflow? Do you get excited about AI agents that fix real vulnerabilities... ...matter most have not been made yet. We are looking for a senior engineer to own that loop. You will spend most of your time on...SeniorFull timeLive inImmediate startRemote workHome officeWeekend work- ...The Senior Software Developer must have working knowledge and an understanding of software engineering practices. In addition, must have experience collaborating with architects, team leads, project managers and product owners to help deliver high quality software features...SeniorFull time
- ...to consider the job opening with US Tech Solutions that fits your expertise and skillset. Job Description Job Title: Senior Software Engineer Location: Denver, CO Description Of Services Experience managing and provisioning computing infrastructure -...SeniorFull time
$120.54k - $140k
...Job Title: Senior Software Engineer Employer: Procare Software, LLC Job Location: Denver, CO Salary: $120,536 - $140,000 Job Duties: Create and support enterprise software solutions both web and mobile applications by maintaining and supporting existing...SeniorFull timeRemote work$165k - $216.56k
...redefine the future of how work gets done.We are looking for a Senior Solution Engineer who is accustomed to solving customer’s most complex... ...States, please visit the job posting on the Snowflake Careers Site for salary and benefits information: careers.snowflake.comCompensation...SeniorRemote work$120k - $175k
...progress, and guiding our clients to success. At Focused, our Engineers are both practitioners and experts in their craft. This role... ...Practice TDD and continuous refactoring to ensure code is reliable, maintainable, and understandable. Guide architecture decisions...SeniorFull timeWork at officeWork visa3 days per week$70 - $80 per hour
...Company Description This position is on-site in Englewood CO. Our Client delivers innovative products and services that power... .... Job Description Summary: Our client is seeking a Senior Software Developer for our Englewood, CO office. Work with a small...SeniorFull timeContract workWork at officeShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- site reliability engineer sre Denver, CO
- site reliability engineer Denver, CO
- senior operations technician Denver, CO
- senior cloud service delivery manager Denver, CO
- senior it service manager Denver, CO
- senior project engineer Denver, CO
- senior chief engineer Denver, CO
- sr operations manager Denver, CO
- senior account director Denver, CO
- senior director clinical development Denver, CO



