Manager, Site Reliability Engineering
O.C. Tanner
O.C. Tanner is the global leader in software and services that improve workplace culture through meaningful employee experiences. Our Culture Cloud is a suite of apps designed to enhance the employee experience with strategic recognition, service awards, wellbeing, leadership, and events that help people thrive at work. Our Culture by Design approach provides expert services to organizations looking to create great workplaces.Our global team of 1,500 people hail from 58 countries and speak 62 languages. As programmers, researchers, designers, client professionals and craftspeople we create the tech, tools and awards that connect employees to purpose at thousands of companies. Join us as we help people all over the world thrive at work.Location: Salt Lake City, UTAs the Manager of Site Reliability Engineering, you will lead the strategy, execution, and evolution of reliability for our world-class employee recognition platform. You will build, mentor, and empower a team of Site Reliability Engineers while partnering closely with Engineering, Product, and Support organizations to deliver highly available, scalable, and resilient services that serve millions of users. We are seeking a leader who is passionate about operational excellence, continuous improvement, and fostering a reliability-first culture through automation, observability, and shared ownership. In this role, you will champion the development of self-healing platforms, drive incident and operational maturity, and enable engineering teams to innovate faster while delivering exceptional customer experiences.Key Responsibilities: Lead, mentor, and develop a team of Site Reliability Engineers, fostering a culture of reliability, accountability, operational excellence, and continuous improvement.Define and execute the organization's reliability strategy, improving availability, scalability, performance, and resilience through automation and engineering best practices.Establish team priorities, goals, and success metrics aligned with business objectives, customer needs, and platform health.Partner with Engineering, Product, and Support leaders to drive shared ownership of production services and embed reliability, observability, and operational excellence throughout the software development lifecycle.Build and evolve observability capabilities using OpenTelemetry, Datadog, Coralogix, or similar tools, establishing enterprise standards for metrics, logs, traces, alerting, and Service Level Objectives (SLOs).Oversee production triage, incident response, and escalation processes, ensuring timely service restoration, effective root cause analysis, and blameless post-incident reviews.Champion a reliability-first engineering culture focused on automation, proactive risk reduction, operational readiness, shift-left quality practices, and continuous improvement.Collaborate with global engineering teams in a follow-the-sun support model, ensuring seamless 24x7 coverage, effective operational handoffs, and consistent service ownership.Own on-call programs, incident management practices, and operational health metrics, driving improvements in alert quality, operational efficiency, and toil reduction.Manage team capacity, hiring, performance management, career development, budgeting, and workforce planning to ensure effective support of business-critical services.Provide regular reporting to engineering and executive leadership on reliability trends, incidents, risks, performance metrics, and strategic initiatives.Required Qualifications5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related disciplines, including 2+ years in a technical leadership or people management role.Proven experience leading teams responsible for production operations, reliability engineering, incident management, and operational excellence.Experience designing and implementing SRE practices, reliability programs, or operational maturity initiatives within growing engineering organizations.Experience operating large-scale, customer-facing SaaS platforms with high availability, performance, and scalability requirements.Strong understanding of modern software engineering practices and partnering with development teams to build reliable, resilient systems.Hands-on experience with observability platforms such as OpenTelemetry, Datadog, Coralogix, or similar technologies.Strong knowledge of AWS and Kubernetes in production environments.Deep understanding of monitoring, logging, distributed tracing, SLIs, SLOs, error budgets, and reliability engineering principles.Demonstrated ability to lead cross-functional initiatives and influence stakeholders across Engineering, Product, and Support organizations.Experience developing engineering roadmaps, defining team objectives, aligning reliability investments with business priorities, and driving continuous operational improvement through incident learning and post-incident reviews.Preferred QualificationsExperience leading distributed or globally dispersed engineering teams.Experience with multiple cloud providers or cloud-agnostic platform architectures.Familiarity with security, compliance, governance, and operational risk management frameworks.Proficiency with modern Infrastructure-as-Code and technologies such as Terraform, Golang, Python, Playwright, and Performance Monitoring tools.Experience with relational and distributed data technologies such as PostgreSQL, OpenSearch, Redis/ElastiCache, or Aurora.Experience with messaging and streaming platforms such as Kafka, ActiveMQ, SNS/SQS, or similar event-driven technologies.Strong understanding of cost optimization, platform sustainability, and engineering efficiency metrics.SummaryLocation: USA - Utah-Salt Lake City-HeadquartersType: Full time
- ...world thrive at work.Location: Salt Lake City, UTAs a Senior Site Reliability Engineer, you will help define the future of reliability for our... ...diagnosing and resolving service disruptions. Drive incident management, root cause analysis, and blameless post incident reviews...SuggestedFull timeShift work
- A recruitment agency for technical graduates seeks candidates for a program linking recent graduates to leading global employers in technology. Applicants will undergo rigorous training and support while working in production support roles, gaining valuable experience ...Suggested
$130k - $160k
...About the Role The Site Reliability Engineering team at iCapital is fundamental to ensuring our platform delivers consistent, reliable service... ...patterns). ~ Familiarity with common data stores and managed services (e.g., Postgres, MongoDB, DynamoDB) and how they...SuggestedFull timeWork at officeRemote work$95k - $171k
...infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's... ...inference platform. As an Site Reliability Engineer II, you will be responsible for:... ...integrating into Akamai's existing incident management processes Contributing to SLO tracking...SuggestedPermanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours$121.4k - $218.6k
...for ensuring best-in-class uptime and reliability of our AI hardware infrastructure... ...when they are breached. As a Senior Site Reliability Engineer, you will be responsible for:... ...rotations, spearheading real-time incident management, and managing high-severity service...SuggestedWork experience placementWork at office$124k - $280k
...Description & SummaryAt PwC, our people in data and analytics engineering focus on leveraging advanced technologies and techniques to design... ...data storage solutions using cloud services- Designing and managing data warehouses and data lakes- Implementing IAM roles and...Full timeH1b$99k - $232k
...SectorNot ApplicableSpecialismSAPManagement LevelManagerJob Description & SummaryThe OpportunityAs a SAP Order to Cash Consultant, Manager, you will lead our clients in their customer transformation journey by reimagining exceptional experiences for their customers and...Full timeH1b$99k - $232k
...for coaching, leveraging team member’s unique strengths, and managing performance to deliver on client expectations. With your growing... ...requirements.The OpportunityAs part of the Data and Analytics Engineering team, you will serve as both a technical leader and a trusted...Full timeH1b$73.5k - $212.28k
Industry/SectorNot ApplicableSpecialismIFS - Information Technology (IT)Management LevelManagerJob Description & SummaryAt PwC, our people in data and analytics engineering focus on leveraging advanced technologies and techniques to design and develop robust data solutions...Full timeH1b- ...product teams to move faster and operate reliably. Success in this role requires building strong partnerships across engineering while delivering secure, scalable platform... ...and continuous growth.Partner with Product Management, Quality Engineering, and cross-functional...Full time
$73.5k - $212.28k
...At PwC, our people in data and analytics engineering focus on leveraging advanced technologies and techniques to design and develop robust... ...for coaching, leveraging team member's unique strengths, and managing performance to deliver on client expectations. With your...H1b$77k - $202k
...requirements and performs as expected.Focused on relationships, you are building meaningful client connections, and learning how to manage and inspire others. Navigating increasingly complex situations, you are growing your personal brand, deepening technical expertise...Full timeH1b$124k - $280k
...Description & SummaryAt PwC, our people in data and analytics engineering focus on leveraging advanced technologies and techniques to design... ...complex needs of health system and health plans. As a Senior Manager, you will drive use case development across clinical decision...Full timeH1b$71.3k - $140.6k
...solutionsBuild, test, and validate integration components to support reliable data exchange across systemsTroubleshoot integration issues,... ...relationshipsAbility to lead projects or workstreamsAbility to manage and prioritize multiple tasks in a fast-paced and dynamic...Local area$114.1k - $268.18k
...in Advisory.KPMG is currently seeking a Manager, SailPoint Identity Governance Technical... ...example:SailPoint Certified IdentityNow Engineer) are preferredProvenapplication onboarding... ...towards the bottom of our KPMG US Careers site at Benefits & How We Work.Follow this link...H1bLocal area$115k - $150k
...team and be responsible for supporting Deloitte’s Application Management Services (AMS) engagement providing ongoing functional... ...preferably in Computer Science, Information Technology, Computer Engineering, or related IT discipline; or equivalent experienceLimited immigration...Local areaRemote workVisa sponsorship$113k - $141.53k
...global energy. Senior Solutions Engineer - Systems Integration serves... ...in the field, ensuring safe, reliable, and performant operation... ...local jurisdiction discussions, managing integration risks, and... ...travel to factories and project sites (25%).Preferred QualificationsMaster...Full timeFor contractorsLocal areaWorldwideFlexible hours$73.8k - $110.7k
...mission to expand access to high-quality, affordable education. Our engineering teams build the platforms, services, and tools that support... ...systems, working with cloud-native technologies, and building reliable software solutions that support learners and employees. What...Full timeInternshipFlexible hours$160.8k - $214.1k
...of modern infrastructure. The Customer Reliability Engineering team is the deep technical escalation... ...cases escalated by Cisco TAC, applying Site Reliability Engineering practices... ...on-premises Kubernetes controller that manages the security policy enforced on it. The...Full timeTemporary workLocal areaRemote workFlexible hours- ...opportunity. We're looking for a talented and motivated Software Engineer II to join our Web Platform Team and help build the... ...infrastructureMonitor platform health and continuously improve reliability, scalability, and developer productivityContribute to an Agile...Full timeFlexible hours
$103.71k - $138.28k
...and experience in system architecture and engineering disciplines. Specific technical... ...Supports due diligence activities including site surveys, design, design review, bill of... ...experience to include indexing, clustering, managing, and troubleshooting. -5+ years with automation...Temporary workRemote work- ...solutions connecting the space, air, land, sea and cyber domains in the interest of national security. Job Title: Senior Manager, System Engineering Job Code: 39284 Job Location: Salt Lake City, UT Job Schedule: 9/80- employees work 9 out of 14 days-...
$136k - $184k
Platform Infrastructure Engineer IV or V, DOEHybrid (Office 3 days/wk - Onsite-Flex) within... ...in cloud, networking, automation, and reliability engineering—tackling complex technical... ...as Code and configuration management (preferred: Terraform and/or Ansible) with...Full timeWork at officeImmediate startWork from homeFlexible hours- ...DescriptionWe're looking for a Senior Software Engineer - Salesforce to join the Student... ...with CI/CD pipeline design and management (Copado or equivalent) Preferred: Salesforce... ...systemic platform issue (performance, reliability, or architectural) Becoming the go-to person...Full timeWork at officeFlexible hours
$163.2k - $220.8k
Wilson Sonsini is the premier legal advisor to technology, life sciences, and other growth enterprises worldwide. We represent companies at every stage of development, from entrepreneurial start-ups to multibillion-dollar global corporations, as well as the venture firms...Full timeRemote workWorldwide$152.5k - $205k
...be responsible for:The Senior Software Engineer is responsible for extending Circle's in... ...microservices that are responsible for reliable and secure APIs that transfer value and... ...financial technologies; consulting with management to ensure agreement on system principles...Permanent employmentRemote workFlexible hours- ...year. We work in a hybrid environment giving you flexibility to manage working from home and being in office.Additional Benefits... ...facilities.JOB SUMMARYSalt Lake County is looking for a software engineer III to support PeopleSoft HCM and Financials. Modules include:...Full timeTemporary workWork at officeWork from homeFlexible hours
- ...responsible AI controls, and identify risks early so our products are reliable, scalable, production-ready, and trusted by millions of users.... ..., and business rules that protect customer trust.Partner with Engineering, Product, and Support throughout the software lifecycle to...Full time
- ...University, we are investing heavily in modern engineering platforms and internal tooling that... ...directly improve engineering velocity, reliability, and consistency across the... ...for service provisioning, environment management, and release automationPartner with DevOps...Full timeFlexible hours
- What We DoAt Goldman Sachs, our Engineers don't just make things - we make things possible. Change the world by connecting people and capital... ...efficient and accurate settlement of trades, helps the firm manage risk and comply with regulations, and establishes exceptional...Full timeTemporary workPart timeImmediate startRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Manager, Site Reliability Engineering. Be the first to apply!
- site services specialist Salt Lake City, UT
- construction site safety Salt Lake City, UT
- site leader Salt Lake City, UT
- official site Salt Lake City, UT
- website content developer Salt Lake City, UT
- on site coordinator Salt Lake City, UT
- IT site lead Salt Lake City, UT
- site safety Salt Lake City, UT
- junior website developer Salt Lake City, UT
- on-site clinical research associate (traveling/remote) Salt Lake City, UT

