Senior Site Reliability Engineer
QAD
Role Description
We are expanding our Site Reliability Engineering (SRE) team and seeking a highly skilled and passionate Senior SRE to join us. As a member of our growing SRE function, you will play a critical role in ensuring the reliability, scalability, and performance of our mission-critical services that power our customer experience. This is an exciting opportunity to shape our SRE practices, drive automation, and significantly impact our product's operational excellence.
- Drive Operational Excellence: Design, implement, and maintain highly available, scalable, and resilient systems that deliver exceptional customer experience.
- Datadog Expert: Be one of the go-to experts for Datadog, responsible for defining, implementing, and enforcing best practices for monitoring, alerting, logging, tracing, and synthetic testing across our entire AWS environment.
- Software Development for Reliability: Develop robust, well-tested, and maintainable software and tooling to automate operational tasks, create self-service capabilities for engineering teams, and enhance system reliability.
- Toil Reduction Champion: Identify and eliminate toil through automation, process improvements, and systematic problem-solving.
- Incident Management & Post-Mortems: Contribute to and evolve our incident response framework, participating in on-call rotations (using OpsGenie) and leading blameless post-mortems.
- Reliability Metrics & Goals: Collaborate with engineering teams to define, implement, and track Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.
- Infrastructure as Code: Leverage and contribute to our infrastructure as code (IaC) efforts, moving towards a fully automated environment using Terraform and GitHub Actions.
- System Design & Architecture: Provide SRE expertise in system design reviews, influencing architectural decisions to build reliability, observability, and scalability into our services.
- Knowledge Sharing & Mentorship: Document processes, build runbooks, and share your expertise with both the SRE team and broader engineering organization.
Qualifications
- Demonstrated experience operating and improving production systems at scale in an SRE, Production Engineering, or Platform Engineering role.
- Proven ability to rapidly build accurate mental models of complex distributed systems across infrastructure, applications, networking, identity, and observability domains.
- Strong troubleshooting skills with a methodical, evidence-driven approach to incident response and root cause analysis.
- Experience defining, implementing, and using Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets to guide reliability decisions.
- Excellent written and verbal communication skills, with the ability to explain complex technical issues clearly to both technical and non-technical audiences.
- Experience across several of the following areas:
- Kubernetes platforms, including Amazon EKS, and service mesh technologies such as Istio.
- Cloud infrastructure and services within AWS.
- Identity and access management systems, including Auth0 and AWS IAM.
- Networking fundamentals, including DNS, load balancing, routing, TLS, and connectivity troubleshooting.
- GitOps workflows and infrastructure automation using tools such as Flux and Terraform.
- Observability platforms and practices, including metrics, logs, traces, alerting, dashboards, and synthetic monitoring.
- CI/CD systems and engineering workflows.
- Application logging and distributed system debugging.
Requirements
- A strong SRE prioritizes service stability and customer impact during incidents.
- Slows down under pressure, gathers facts, and communicates clearly.
- Reduces operational complexity through automation and simplification.
- Identifies and eliminates toil through self-service tooling and process improvement.
- Demonstrates strong scripting and automation instincts.
- Brings a systems-thinking approach to problem-solving.
- Balances short-term remediation with long-term reliability improvements.
- Demonstrated ability to build and maintain automation, tooling, and self-service capabilities using one or more programming or scripting languages such as Python, Go, or Bash.
- Focuses on applying software engineering practices to improve reliability, reduce toil, and enhance developer productivity.
Benefits
- Calm and effective during high-severity incidents.
- Skilled at managing complex situations involving multiple teams and competing priorities.
- Able to lead blameless post-mortems and drive meaningful follow-up actions.
- Passionate about continuous improvement and fostering a culture of shared ownership.
$158.5k - $172k
...exceptional value they deserve. About The Opportunity As a Senior Engineer on the Runtime Automation team, you will design, automate,... ...environment. This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire...SeniorFull timeWork at office3 days per week$175k - $185k
...together. Come join our team as we develop new ways to improve the lives of working Americans. About the role: As the Senior Site Reliability Engineer, you will lead Branch’s effort to achieve greater reliability, performance, scalability, capacity and observability of...SeniorDaily paidRemote workHome officeFlexible hours- ...our Series B and have grown 800% over the last 12 months. Engineering at Ivo Engineers at Ivo are inventors. Ivo was first-to-... ...expect us to hit our SLAs. What? We’re looking for an Senior Site level Reliability Engineer as part of Infrastructure team to: Own uptime,...SeniorContract workWork at officeRemote workVisa sponsorshipRelocation packageFlexible hours
$150k - $190k
...operate. You'll be responsible for the reliability, performance, and availability of Develocity... ...the tooling you depend on, and with engineering teams to build reliability into how we... ..., relevant skills, qualifications, seniority, performance, and travel requirements....SeniorFull timeRemote workWork from homeShift work- Role Description Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure... ...for billions of people, billions of times a day. As a Senior Site Reliability Engineer, you will be: ~Designing, developing...SeniorFull timeWork at office
- Role Description Stack AV Site Reliability Engineers are responsible for enabling and ensuring our production systems meet their service-level objectives. Through the implementation of centralized observability and automation, the SRE team constantly ensures the health...SeniorFull time
- Role Description We’re looking for a Senior Site Reliability Engineer who takes ownership seriously — someone who designs for reliability, ships the automation, and stands behind it in production. You’ll work across cloud-native infrastructure on systems that process millions...SeniorFull time
$175k - $250k
...000.00/yr - $250,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On-Site only. Must live within commuting distance... ...scalability, performance, and reliability across environments. What You’ll Do Design...SeniorFull timeRemote workRelocationRelocation package- Role Description Join us as a Senior Site Reliability Engineer on our mission to turn payments into possibilities! The Site Reliability Engineering (SRE) team ensures the reliability, availability, scalability, and performance of a mission-critical payment orchestration...SeniorFull timeImmediate startRemote work
- ...Istio) ~Defining and monitoring Service-Level Objectives (SLOs) and Service-Level Agreements (SLAs) to ensure that systems meet reliability and performance targets ~Monitoring Tools like New Relic, Prometheus, Grafana, and/or Datadog ~OpenTelemetry knowledge for...SeniorFull timeRemote work
$54k - $150k
Role Description As Senior Site Reliability Engineer for Remote Build, you'll own the operational excellence and infrastructure strategy that makes Build's platform reliable, performant, and safe for customers. You'll report to the Engineering Manager and work closely...SeniorFull timeLocal areaRemote workHome officeFlexible hours$54k - $150k
Role Description As Senior Site Reliability Engineer for Remote Build, you'll own the operational excellence and infrastructure strategy that makes Build's platform reliable, performant, and safe for customers. You'll report to the Engineering Manager and work closely...SeniorFull timeLocal areaImmediate startRemote workHome officeFlexible hours- ...investors including Silver Lake Waterman, Moody’s, Sequoia Capital, GV and Riverwood Capital. About the Team: As a Senior Site Reliability Engineer, you will be a key technical leader driving the design and optimization of our Kubernetes-based infrastructure and CI/...SeniorFull timeRemote work
- ...Azure, Oracle, Cassandra, SQL Server, My SQL and Mongo DB Seniority level Seniority level Mid-Senior level Employment type Employment... ...new job is posted. Sign in to set job alerts for “Senior Site Reliability Engineer” roles. Bellevue, WA $204,000.00-$259,000.00 1 day ago...SeniorFull timeContract workRemote work
- Role Description The Senior Site Reliability Engineer is a technical leader responsible for architecting the reliability strategy for large-scale, distributed government systems. You will lead the implementation of the SRE framework, driving the adoption of SLO-based management...SeniorContract workRemote work
- Role Description Versant's Sports & Entertainment Digital Products division is seeking a Senior Site Reliability Engineer to help drive the reliability, scalability, and usability of internal developer platforms, tooling, and engineering workflows across a portfolio of...SeniorFull timeLocal areaRemote workWorldwide
- Role Description We’re looking for a Senior Platform Engineer to design, build, and operate the core services that power Optura’s AI Platform... ...end-to-end, from model and agent orchestration to routing, reliability, and observability. You will partner closely with product...SeniorFull timeRemote work
$190.8k - $267.1k
...is a unique opportunity to leave your mark on one of the most influential and trafficked corners of the internet. As a Senior Site Reliability Engineer on Reddit’s Infrastructure SRE team, you’ll use your knowledge of distributed systems and architecture to improve the...SeniorFull timeWork experience placementHome officeFlexible hours- ...to grow, adapt and use your skills consistently. Our customers rely on us in the moments that matter. Engineering delivers on that promise. The Senior Site Reliability Engineer is responsible for ensuring our SaaS products are fast, stable and optimized for our...SeniorFull timeWork experience placement
$125.04k - $187.56k
...services, including Finance, Legal, Sustainability, Commercial, Digital and E-commerce, Technology and more. Overview The Site Reliability Engineer (SRE) III is responsible for ensuring the scalability, reliability, and performance of production systems through automation...SeniorFull timeWork at officeRemote workFlexible hours$137.9k - $221.4k
...for someone to lead development aspects of the Infrastructure engineering team at ServiceTitan. You must have a strong background in... ...leadership and strong architectural thought process. Our Site Reliability and Infrastructure Engineering team is an investment by Cloud...SeniorFull timeImmediate startFlexible hours$130k - $170k
...Senior Site Reliability Engineer About Us Founded in 2014, we offer the industry’s first and only cloud‑based, fully‑customisable, end‑to‑end software solution to automate securities‑based lending from origination through the life of the loan. By combining thought...SeniorFull timeFlexible hoursShift work- ...optimize production infrastructure across CI/CD, cloud deployments, and security. You will collaborate with our internal product and engineering teams to keep services scalable, secure, and highly available. The role emphasizes GitHub Actions, Terraform, Vercel, AWS core...SeniorRemote work
$190k - $240k
Role Description As a Sr. Site Reliability Engineer (SRE) at ICD, you will play a critical role in ensuring the reliability and seamless operation of our global platform and AWS infrastructure to create scalable and highly reliable software systems. Job Responsibilities...SeniorFull timeWork at officeImmediate startFlexible hours$149.4k - $202k
...Noctua Technology is seeking a Senior Software Engineer specializing in Site Reliability Engineering to join their team. This role focuses on the reliability and performance of cloud-native applications, emphasizing Infrastructure as Code and automation. The ideal candidate...SeniorRemote work$125k - $135k
Role Description Vultr is seeking a highly skilled and experienced Senior Site Reliability Engineer to build and own the observability pipeline for the physical and provisioning infrastructure that powers Vultr's global datacenter footprint. The ideal candidate is a builder...SeniorFull timeWork at officeImmediate startRemote work- ...A leading livestream shopping platform is seeking a Senior Software Engineer for the Logistics Platform team. This role focuses on improving logistical data systems, enhancing buyer and seller experience, and fostering collaboration across departments. Ideal candidates...SeniorRemote work
- ...Passionate about designing and automating cloud platforms, the full-time Senior Azure Site Reliability Engineer will focus on improving reliability through Infrastructure as Code, automation, and observability while managing Azure cloud infrastructure and collaborating...SeniorFull timeRemote work
- ...Partner with software developers, platform engineers, and IT staff to improve system design,... ...requirements, service quality, reliability, security, and compliance needs. Drive continuous... ...Required: 8+ years of experience in Site Reliability Engineering, DevOps, Platform...SeniorWork at officeRemote work
$174k - $253k
...MINIMUM QUALIFICATIONS: Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical experience. 5... ...s degree in Computer Science or Engineering. ABOUT THE JOB: Site Reliability Engineering (SRE) is what you get when you treat operations...Senior
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- site reliability engineer Remote
- site reliability engineer remote Remote
- site reliability engineer sre Remote
- sr hr business partner Remote
- senior lighting artist Remote
- senior planner Remote
- senior hvac project manager Remote
- home instead senior care Remote
- research associate senior research associate Remote
- senior technical product manager Remote




















