Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Director, Site Reliability Engineering

Full-time

Jobgether

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Director, Site Reliability Engineering based in United States. As Director, Site Reliability Engineering, you will lead the teams and technical strategy responsible for keeping critical infrastructure reliable, scalable, and secure for millions of users. You will work across software, systems, automation, cloud infrastructure, and operational processes to solve complex reliability challenges at scale. The role combines strategic leadership with hands-on technical depth, including troubleshooting production systems and partnering closely with software engineering teams. You will help shape the future architecture and deployment practices of a large-scale, privacy-focused technology environment. You will also drive improvements in automation, observability, incident response, and engineering efficiency. This is a remote-first leadership opportunity with significant ownership, autonomy, and impact. Accountabilities: - Lead and develop Site Reliability Engineering teams responsible for the reliability, scalability, performance, and operational health of large-scale systems. - Define and execute the technical direction for infrastructure, deployment, reliability engineering, automation, and operational practices. - Lead high-impact and complex initiatives from initial proposal and planning through implementation, measurement, and postmortem. - Investigate and resolve sources of instability across high-traffic, distributed systems, identifying root causes and implementing sustainable remediation. - Establish and improve tools, services, monitoring, alerts, incident-response processes, and operational practices that identify and mitigate reliability risks. - Partner closely with software engineers to troubleshoot production issues, evaluate performance considerations, and implement appropriate code-level or infrastructure-level solutions. - Drive automation for infrastructure provisioning and configuration management to improve efficiency, scalability, consistency, and reliability.

- Leverage cloud-native architectures and services to strengthen system resilience and support continued growth. - Help ensure products and infrastructure meet established reliability standards while minimizing user impact during failures and incidents. - Identify emerging technical needs and opportunities to guide the long-term evolution of deployment and infrastructure architecture. - Support a culture of ownership, continuous improvement, measurable outcomes, and effective post-incident learning. Requirements: - 10+ years of relevant professional experience in Site Reliability Engineering, platform engineering, infrastructure engineering, software engineering, or related fields. - 4+ years of experience leading SRE or comparable engineering teams. - Experience participating in or managing 24/7 on-call operations for large-scale production environments. - Advanced programming experience and the ability to read, write, troubleshoot, and deploy software across production systems. - Strong experience with Linux administration and troubleshooting, web technologies, distributed systems, and high-traffic production environments. - Demonstrated ability to lead complex technical projects from ambiguous initial requirements through execution and postmortem. - Experience developing effective reliability tooling, services, monitoring, alerting, and incident-response capabilities. - Strong investigative and root-cause analysis skills, particularly within distributed and high-scale systems. - Experience designing and implementing infrastructure automation, provisioning, and configuration-management solutions. - Hands-on experience with cloud-native services and architectures, including application packaging and deployment using Docker and Docker Compose.

Vacancy posted 7 days ago
Similar jobs that could be interesting for youBased on the Director, Site Reliability Engineering in Remote vacancy
  • $81.5k - $141.3k

     ...branded generic medicines. Our 122,000 colleagues serve people in more than 160 countries.JOB DESCRIPTION:Position Title: Site Reliability Engineer IITeam: CRM DevOps Employment Type: Full-TimeAbout the RoleThis Site Reliability Engineer II position works on-site out of... 
    Suggested
    Remote work
    Shift work

    Abbott

    Sylmar, CA
    2 days ago
  • $106.7k - $177.9k

    OverviewJoin M&T Bank's Digital Banking organization and help drive the reliability, resiliency, and performance of the platforms our customers depend on every day. As a Senior Software Engineer, Site Reliability Engineering (SRE), you will play a key role in supporting... 
    Suggested
    Permanent employment
    Full time
    Work experience placement

    M&T Bank

    Wilmington, DE
    16 hours ago
  • $87.12k - $151.25k

     ...If you want to be part of an inclusive, adaptable, and forward-thinking organization, apply now.We are currently seeking a Site Reliability Engineer- REMOTE - Onsite Training to join our team in Memphis, Tennessee (US-TN), United States (US).Day-to Day Responsibilities:... 
    Suggested
    Full time
    Temporary work
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    NTT DATA

    Memphis, TN
    2 days ago
  • Site Reliability Engineers are responsible for ensuring the availability, reliability, scalability, and performance of the firm’s most critical customer-facing microservices that power all eCommerce channels. This role applies Google-inspired SRE principles to balance... 
    Suggested
    Local area
    Remote work
    Flexible hours
    Shift work

    O'Reilly Auto Parts

    Springfield, MO
    4 days ago
  • $96.8k - $145.2k

     ...If you want to be part of an inclusive, adaptable, and forward-thinking organization, apply now.We are currently seeking a Site Reliability Engineer (Onsite Hybrid) to join our team in Plano, Texas (US-TX), United States (US).Job Responsibilities Include: Own and manage... 
    Suggested
    Temporary work
    Work at office
    Remote work
    Flexible hours

    NTT DATA

    Plano, TX
    16 hours ago
  •  ...The Role:GIPHY is seeking a highly experienced Site Reliability Engineer to join our SRE team. You will help design, build, operate, and evolve the infrastructure that powers GIPHY, including our cloud environment, Kubernetes clusters, and CI/CD platforms.You will also... 
    Full time
    Work experience placement
    Remote work

    Shutterstock

    New York, NY
    16 hours ago
  • As the Senior Site Reliability Engineer, you will serve as a trusted technical resource responsible for deploying, validating, and operationalizing AI, HPC, Kubernetes, and enterprise infrastructure environments. This role transforms newly installed hardware into production... 
    Work at office
    Immediate start
    Worldwide
    Shift work

    Wesco International

    Atlanta, GA
    16 hours ago
  • $115.5k - $164.8k

     ...mission that matters at a company where you matter.Your ImpactAs an engineer on the APX SRE CloudOps team, you will spend a significant...  ...that replace what previously required human intervention with reliable, tested automation. You will also participate in on-call rotations... 
    Work experience placement
    Work at office
    Remote work

    Axon

    Washington DC
    16 hours ago
  • $130k - $180k

     ...building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference...  ...and AI R&D. The role Nebius is looking for a Site Reliability Engineer in Hardware Infrastructure team. This is a remote position... 
    Temporary work
    Immediate start
    Remote work

    Nebius

    Remote
    2 days ago
  • $180k - $195k

     ...is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer, NetBox Delivery based in United States. This is a senior engineering role focused on making a widely used network... 
    Temporary work
    Remote work

    Jobgether

    Oregon State
    1 day ago
  • $262k - $364k

     ...and training AI infrastructure from SRE side, ensuring it is reliable, scalable, cost effective and performant, while working closely...  ...qualifications:Master's degree in Computer Science or Engineering.Site Reliability Engineering (SRE) combines software and systems engineering... 

    Google

    Sunnyvale, CA
    16 hours ago
  • $166.3k - $238.3k

     ...and more intuitive with technology that simply works. The SRE Engineering Enablement org for our Network Platform team supports our CI...  ...customers are all engineers at Cisco.YOUR IMPACT As a Lead Site Reliability Engineer, you will architect, build, and evolve the... 
    Full time
    Temporary work
    Local area
    Remote work
    Flexible hours

    CISCO Systems

    Boston, MA
    21 hours ago
  • $125k - $250k

     ...we are reimagining how developers build reliable, scalable, event-driven applications without...  ...possible Partner closely with engineering teams to improve system resiliency and scalability...  ...For 5+ years of experience in Site Reliability Engineering, DevOps,... 
    Full time
    Immediate start
    Remote work
    Flexible hours

    Orkes

    United States
    21 hours ago
  • $180k - $230k

     ...Acceleration Job Description We're looking for a Senior SRE to own the reliability, scalability, and observability of our production systems. You'll work closely with platform and data engineering to keep high-throughput, data-intensive services running at the... 
    Work at office
    Local area
    Immediate start
    Remote work
    3 days per week

    GridCARE, Inc.

    Eastern, KY
    1 day ago
  •  ...A senior Site Reliability Engineer will join an established infrastructure function responsible for highly available, security-conscious cloud systems supporting complex business-critical workloads. You’ll take significant ownership of reliability, scalability, and... 
    Full time
    Remote work

    Motion Recruitment Partners LLC

    Eastern, KY
    1 day ago
  •  ...Senior Site Reliability Engineer Company: Sphera Work Type: Remote Employment: Full Time Location: US Seniority: Senior Level Technologies: Terraform, ARM templates, Kubernetes, Azure, SonarCloud, CheckPoint, Hadoop, Kafka, Presto, NewRelic, CI/CD, Linux, Windows, Redis... 
    Full time
    Remote work

    Sphera

    United States
    21 hours ago
  •  ...Blackpoint Cyber is in hyper-growth mode, fueled by a recent $190m series C round.  SUMMARY We're hiring a Senior Site Reliability Engineer to design, implement, and maintain our cloud and on-premise infrastructure and CI/CD pipelines, with a focus on automation... 
    Local area
    Remote work

    Blackpoint Cyber

    United States
    2 days ago
  •  ...ears, and hands on the ground at a government customer site, ensuring the reliability and performance of Twenty's mission-critical platform running...  ...of deep technical ownership and customer-facing engineering: you'll define how we measure reliability, lead incident... 
    Full time
    Contract work
    Remote work
    Flexible hours

    Twenty Inc.

    Arlington, VA
    21 hours ago
  •  ...Senior Site Reliability Engineer Bentley Systems Location Exton or Philadelphia, PA (Hybrid - 3 times a week in-office) Position Summary Are you ready to start a new journey with a team of energized professionals advancing and connecting the world's infrastructure... 
    Casual work
    Work at office
    Worldwide

    Bentley Systems

    Exton, PA
    1 day ago
  • $140k - $180k

     ...-making, and accelerated growth in the AI-driven world. Learn more at Opportunity We’re looking for a Senior Site Reliability Engineer to help build and scale a high-impact SRE function. You’ll be a technical leader on a team responsible for improving system... 
    Work experience placement
    Local area
    Remote work
    Visa sponsorship
    Work visa

    UJET

    United States
    2 days ago
  •  ...grow, adapt and use your skills consistently. Our customers rely on us in the moments that matter. Engineering delivers on that promise.   The Senior Site Reliability Engineer is responsible for ensuring our SaaS products are fast, stable and optimized for our customers... 
    Work experience placement
    Remote work
    Flexible hours

    Donnelley Financial Solutions

    United States
    4 days ago
  •  ...Site Reliability Engineer Company: GitLab Work Type: Remote Employment: Full Time Location: CA, US Seniority: Senior Level Technologies: Terraform, Ansible, Kubernetes, Go, Ruby, Jsonnet, Prometheus, ELK, Grafana Requirements: Senior-level SRE with strong Terraform/IaC... 
    Full time
    Remote work

    GitLab

    United States
    21 hours ago
  •  ...encourage you to apply. The Role  As a Senior Platform Engineer, you are a champion for DevOps and SRE culture and industry...  ...met. \n What You Will Be Doing Improving production reliability and system resilience within an SRE scoped team Championing... 
    Remote work
    Flexible hours

    Megaport

    United States
    4 days ago
  •  ...Engineering, Product, Design, and Marketing Engineering Compensation ~ Zone 1 Base Pay: $214K – $260K Superhuman offers...  ...role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them... 
    Worldwide
    Home office
    Flexible hours

    Superhuman

    Eastern, KY
    3 days ago
  • $141.8k - $195k

     ...work, grow fast, and bring their full selves to the herd. Why You’ll Love This Role Cribl Inc is seeking a Senior Site Reliability Engineer to join our mission where you will unlock the value of all observability data, as we expand our team in the U.S. Cribl provides... 
    Temporary work
    Remote work

    Cribl

    United States
    21 hours ago
  •  ...future of short-term rentals. ⭐ Role Overview Are you a systems-minded engineer who cares deeply about reliability, scalability, and production excellence? Join Lodgify as a Senior Site Reliability Engineer and help our engineering teams build and operate... 
    Temporary work
    Remote work
    Worldwide

    Lodgify

    United States
    3 hours ago
  • $175k - $250k

     ...Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On‑Site only. Must live within commuting distance of San Francisco or...  ...while ensuring scalability, performance, and reliability across environments. What You’ll Do Design, build... 
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    Washington DC
    2 days ago
  • $150k - $170k

     ...Senior Site Reliability Engineer – Zip CoJoin to apply for the Senior Site Reliability Engineer role at Zip CoAt Zip, we build cloud-native software applications that serve millions of customers and process billions of dollars in payments. We're looking for a seasoned... 
    Casual work
    Work at office
    Remote work

    ZIP

    New York, NY
    2 days ago
  • $120k - $155k

     ...and intelligence. If you want to push yourself and reshape a $200B+ market, we’re excited to talk to you! What will the Site Reliability Engineer do? We're looking for a Site Reliability Engineer who's passionate about building and maintaining reliable, scalable infrastructure... 
    Full time
    Immediate start
    Remote work
    Visa sponsorship
    Flexible hours

    Cintrifuse

    Springfield, VA
    1 day ago
  •  ...% uptime. You'll own SLOs, incident response, and production reliability for a system that processes millions of identity verifications...  ...Sentry error tracking, structured logging Implement chaos engineering practices to proactively identify failure modes Optimize... 
    Remote work

    Xident B.V.

    Union, NJ
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Director, Site Reliability Engineering. Be the first to apply!