Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

Full-time

QAD

Role Description

We are expanding our Site Reliability Engineering (SRE) team and seeking a highly skilled and passionate Senior SRE to join us. As a member of our growing SRE function, you will play a critical role in ensuring the reliability, scalability, and performance of our mission-critical services that power our customer experience. This is an exciting opportunity to shape our SRE practices, drive automation, and significantly impact our product's operational excellence.

  • Drive Operational Excellence: Design, implement, and maintain highly available, scalable, and resilient systems that deliver exceptional customer experience.
  • Datadog Expert: Be one of the go-to experts for Datadog, responsible for defining, implementing, and enforcing best practices for monitoring, alerting, logging, tracing, and synthetic testing across our entire AWS environment.
  • Software Development for Reliability: Develop robust, well-tested, and maintainable software and tooling to automate operational tasks, create self-service capabilities for engineering teams, and enhance system reliability.
  • Toil Reduction Champion: Identify and eliminate toil through automation, process improvements, and systematic problem-solving.
  • Incident Management & Post-Mortems: Contribute to and evolve our incident response framework, participating in on-call rotations (using OpsGenie) and leading blameless post-mortems.
  • Reliability Metrics & Goals: Collaborate with engineering teams to define, implement, and track Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.
  • Infrastructure as Code: Leverage and contribute to our infrastructure as code (IaC) efforts, moving towards a fully automated environment using Terraform and GitHub Actions.
  • System Design & Architecture: Provide SRE expertise in system design reviews, influencing architectural decisions to build reliability, observability, and scalability into our services.
  • Knowledge Sharing & Mentorship: Document processes, build runbooks, and share your expertise with both the SRE team and broader engineering organization.

Qualifications

  • Demonstrated experience operating and improving production systems at scale in an SRE, Production Engineering, or Platform Engineering role.
  • Proven ability to rapidly build accurate mental models of complex distributed systems across infrastructure, applications, networking, identity, and observability domains.
  • Strong troubleshooting skills with a methodical, evidence-driven approach to incident response and root cause analysis.
  • Experience defining, implementing, and using Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets to guide reliability decisions.
  • Excellent written and verbal communication skills, with the ability to explain complex technical issues clearly to both technical and non-technical audiences.
  • Experience across several of the following areas:
    • Kubernetes platforms, including Amazon EKS, and service mesh technologies such as Istio.
    • Cloud infrastructure and services within AWS.
    • Identity and access management systems, including Auth0 and AWS IAM.
    • Networking fundamentals, including DNS, load balancing, routing, TLS, and connectivity troubleshooting.
    • GitOps workflows and infrastructure automation using tools such as Flux and Terraform.
    • Observability platforms and practices, including metrics, logs, traces, alerting, dashboards, and synthetic monitoring.
    • CI/CD systems and engineering workflows.
    • Application logging and distributed system debugging.

Requirements

  • A strong SRE prioritizes service stability and customer impact during incidents.
  • Slows down under pressure, gathers facts, and communicates clearly.
  • Reduces operational complexity through automation and simplification.
  • Identifies and eliminates toil through self-service tooling and process improvement.
  • Demonstrates strong scripting and automation instincts.
  • Brings a systems-thinking approach to problem-solving.
  • Balances short-term remediation with long-term reliability improvements.
  • Demonstrated ability to build and maintain automation, tooling, and self-service capabilities using one or more programming or scripting languages such as Python, Go, or Bash.
  • Focuses on applying software engineering practices to improve reliability, reduce toil, and enhance developer productivity.

Benefits

  • Calm and effective during high-severity incidents.
  • Skilled at managing complex situations involving multiple teams and competing priorities.
  • Able to lead blameless post-mortems and drive meaningful follow-up actions.
  • Passionate about continuous improvement and fostering a culture of shared ownership.
Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in Remote vacancy
  • $158.5k - $172k

     ...exceptional value they deserve. About The Opportunity As a Senior Engineer on the Runtime Automation team, you will design, automate,...  ...environment. This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire... 
    Senior
    Full time
    Work at office
    3 days per week

    Grubhub

    Chicago, IL
    4 days ago
  • $175k - $185k

     ...together. Come join our team as we develop new ways to improve the lives of working Americans. About the role: As the Senior Site Reliability Engineer, you will lead Branch’s effort to achieve greater reliability, performance, scalability, capacity and observability of... 
    Senior
    Daily paid
    Remote work
    Home office
    Flexible hours

    Branch

    United States
    2 days ago
  •  ...our Series B and have grown 800% over the last 12 months. Engineering at Ivo Engineers at Ivo are inventors. Ivo was first-to-...  ...expect us to hit our SLAs. What? We’re looking for an Senior Site level Reliability Engineer as part of Infrastructure team to: Own uptime,... 
    Senior
    Contract work
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours

    Ivo Inc.

    San Francisco, CA
    3 days ago
  • $150k - $190k

     ...operate. You'll be responsible for the reliability, performance, and availability of Develocity...  ...the tooling you depend on, and with engineering teams to build reliability into how we...  ..., relevant skills, qualifications, seniority, performance, and travel requirements.... 
    Senior
    Full time
    Remote work
    Work from home
    Shift work

    Develocity

    United States
    2 days ago
  • Role Description Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure...  ...for billions of people, billions of times a day. As a Senior Site Reliability Engineer, you will be: ~Designing, developing... 
    Senior
    Full time
    Work at office

    Akamai

    Remote
    16 hours ago
  • Role Description Stack AV Site Reliability Engineers are responsible for enabling and ensuring our production systems meet their service-level objectives. Through the implementation of centralized observability and automation, the SRE team constantly ensures the health... 
    Senior
    Full time

    Stack AV

    Remote
    1 day ago
  • Role Description We’re looking for a Senior Site Reliability Engineer who takes ownership seriously — someone who designs for reliability, ships the automation, and stands behind it in production. You’ll work across cloud-native infrastructure on systems that process millions... 
    Senior
    Full time

    CertifyOS

    Remote
    5 days ago
  • $175k - $250k

     ...000.00/yr - $250,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On-Site only. Must live within commuting distance...  ...scalability, performance, and reliability across environments. What You’ll Do Design... 
    Senior
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    San Francisco, CA
    2 days ago
  • Role Description Join us as a Senior Site Reliability Engineer on our mission to turn payments into possibilities! The Site Reliability Engineering (SRE) team ensures the reliability, availability, scalability, and performance of a mission-critical payment orchestration... 
    Senior
    Full time
    Immediate start
    Remote work

    CellPoint Digital

    Remote
    6 days ago
  •  ...Istio) ~Defining and monitoring Service-Level Objectives (SLOs) and Service-Level Agreements (SLAs) to ensure that systems meet reliability and performance targets ~Monitoring Tools like New Relic, Prometheus, Grafana, and/or Datadog ~OpenTelemetry knowledge for... 
    Senior
    Full time
    Remote work

    Shippo

    Remote
    5 days ago
  • $54k - $150k

    Role Description As Senior Site Reliability Engineer for Remote Build, you'll own the operational excellence and infrastructure strategy that makes Build's platform reliable, performant, and safe for customers. You'll report to the Engineering Manager and work closely... 
    Senior
    Full time
    Local area
    Remote work
    Home office
    Flexible hours

    Remote

    Remote
    7 days ago
  • $54k - $150k

    Role Description As Senior Site Reliability Engineer for Remote Build, you'll own the operational excellence and infrastructure strategy that makes Build's platform reliable, performant, and safe for customers. You'll report to the Engineering Manager and work closely... 
    Senior
    Full time
    Local area
    Immediate start
    Remote work
    Home office
    Flexible hours

    Referral Board

    Remote
    16 hours ago
  •  ...investors including Silver Lake Waterman, Moody’s, Sequoia Capital, GV and Riverwood Capital. About the Team: As a Senior Site Reliability Engineer, you will be a key technical leader driving the design and optimization of our Kubernetes-based infrastructure and CI/... 
    Senior
    Full time
    Remote work

    SecurityScorecard

    New York, NY
    2 days ago
  •  ...Azure, Oracle, Cassandra, SQL Server, My SQL and Mongo DB Seniority level Seniority level Mid-Senior level Employment type Employment...  ...new job is posted. Sign in to set job alerts for “Senior Site Reliability Engineer” roles. Bellevue, WA $204,000.00-$259,000.00 1 day ago... 
    Senior
    Full time
    Contract work
    Remote work

    Signature IT World Inc

    Washington DC
    2 days ago
  • Role Description The Senior Site Reliability Engineer is a technical leader responsible for architecting the reliability strategy for large-scale, distributed government systems. You will lead the implementation of the SRE framework, driving the adoption of SLO-based management... 
    Senior
    Contract work
    Remote work

    Arctiq

    Remote
    4 days ago
  • Role Description Versant's Sports & Entertainment Digital Products division is seeking a Senior Site Reliability Engineer to help drive the reliability, scalability, and usability of internal developer platforms, tooling, and engineering workflows across a portfolio of... 
    Senior
    Full time
    Local area
    Remote work
    Worldwide

    Versant

    Remote
    16 hours ago
  • Role Description We’re looking for a Senior Platform Engineer to design, build, and operate the core services that power Optura’s AI Platform...  ...end-to-end, from model and agent orchestration to routing, reliability, and observability. You will partner closely with product... 
    Senior
    Full time
    Remote work

    Optura

    Remote
    5 days ago
  • $190.8k - $267.1k

     ...is a unique opportunity to leave your mark on one of the most influential and trafficked corners of the internet. As a Senior Site Reliability Engineer on Reddit’s Infrastructure SRE team, you’ll use your knowledge of distributed systems and architecture to improve the... 
    Senior
    Full time
    Work experience placement
    Home office
    Flexible hours

    Alien Blue

    Chicago, IL
    2 days ago
  •  ...to grow, adapt and use your skills consistently. Our customers rely on us in the moments that matter. Engineering delivers on that promise. The Senior Site Reliability Engineer is responsible for ensuring our SaaS products are fast, stable and optimized for our... 
    Senior
    Full time
    Work experience placement

    Donnelley Financial Solutions

    Remote
    4 days ago
  • $125.04k - $187.56k

     ...services, including Finance, Legal, Sustainability, Commercial, Digital and E-commerce, Technology and more. Overview The Site Reliability Engineer (SRE) III is responsible for ensuring the scalability, reliability, and performance of production systems through automation... 
    Senior
    Full time
    Work at office
    Remote work
    Flexible hours

    ViziRecruiter,LLC.

    Chicago, IL
    2 days ago
  • $137.9k - $221.4k

     ...for someone to lead development aspects of the Infrastructure engineering team at ServiceTitan. You must have a strong background in...  ...leadership and strong architectural thought process. Our Site Reliability and Infrastructure Engineering team is an investment by Cloud... 
    Senior
    Full time
    Immediate start
    Flexible hours

    ServiceTitan

    Remote
    2 days ago
  • $130k - $170k

     ...Senior Site Reliability Engineer About Us Founded in 2014, we offer the industry’s first and only cloud‑based, fully‑customisable, end‑to‑end software solution to automate securities‑based lending from origination through the life of the loan. By combining thought... 
    Senior
    Full time
    Flexible hours
    Shift work

    Supernova Technology™

    Chicago, IL
    2 days ago
  •  ...optimize production infrastructure across CI/CD, cloud deployments, and security. You will collaborate with our internal product and engineering teams to keep services scalable, secure, and highly available. The role emphasizes GitHub Actions, Terraform, Vercel, AWS core... 
    Senior
    Remote work

    Remote Leverage

    New York, NY
    3 days ago
  • $190k - $240k

    Role Description As a Sr. Site Reliability Engineer (SRE) at ICD, you will play a critical role in ensuring the reliability and seamless operation of our global platform and AWS infrastructure to create scalable and highly reliable software systems. Job Responsibilities... 
    Senior
    Full time
    Work at office
    Immediate start
    Flexible hours

    Tradeweb

    Remote
    5 days ago
  • $149.4k - $202k

     ...Noctua Technology is seeking a Senior Software Engineer specializing in Site Reliability Engineering to join their team. This role focuses on the reliability and performance of cloud-native applications, emphasizing Infrastructure as Code and automation. The ideal candidate... 
    Senior
    Remote work

    Noctua Technology

    Virginia, MN
    3 days ago
  • $125k - $135k

    Role Description Vultr is seeking a highly skilled and experienced Senior Site Reliability Engineer to build and own the observability pipeline for the physical and provisioning infrastructure that powers Vultr's global datacenter footprint. The ideal candidate is a builder... 
    Senior
    Full time
    Work at office
    Immediate start
    Remote work

    Vultr

    Remote
    5 days ago
  •  ...A leading livestream shopping platform is seeking a Senior Software Engineer for the Logistics Platform team. This role focuses on improving logistical data systems, enhancing buyer and seller experience, and fostering collaboration across departments. Ideal candidates... 
    Senior
    Remote work

    Whatnot

    Los Angeles, CA
    4 days ago
  •  ...Passionate about designing and automating cloud platforms, the full-time Senior Azure Site Reliability Engineer will focus on improving reliability through Infrastructure as Code, automation, and observability while managing Azure cloud infrastructure and collaborating... 
    Senior
    Full time
    Remote work

    Virtual Vocations Inc

    United States
    3 days ago
  •  ...Partner with software developers, platform engineers, and IT staff to improve system design,...  ...requirements, service quality, reliability, security, and compliance needs. Drive continuous...  ...Required: 8+ years of experience in Site Reliability Engineering, DevOps, Platform... 
    Senior
    Work at office
    Remote work

    GrabJobs

    Cincinnati, OH
    1 day ago
  • $174k - $253k

     ...MINIMUM QUALIFICATIONS: Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical experience. 5...  ...s degree in Computer Science or Engineering. ABOUT THE JOB: Site Reliability Engineering (SRE) is what you get when you treat operations... 
    Senior

    Socket

    Raleigh, NC
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!