Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer (GPU Clusters) - Hosting

$250k
Full-time

Looking for a role with plenty of growth opportunities?

Join a rapidly scaling AI cloud infrastructure provider building a next-generation GPU platform designed for AI training, experimentation, and inference at scale. The company is developing a fully featured AI cloud platform powered by renewable energy and is already operating with strong momentum across Europe, while now significantly expanding its footprint in the United States.

The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments powering GPU-intensive workloads. The role involves working closely with platform, ML, and infrastructure teams to improve reliability, automation, and observability across distributed compute environments while supporting long-term infrastructure growth and scalability.

Don’t miss out on this exciting opportunity and apply today!

Responsibilities:

  • Ensure the reliability, scalability, and performance of HPC and cloud infrastructure environments
  • Design, build, and maintain automation, observability, and monitoring frameworks for GPU compute clusters
  • Collaborate with ML, data, and platform engineering teams to deliver highly available infrastructure systems
  • Improve CI/CD pipelines, deployment workflows, and operational tooling
  • Contribute to infrastructure architecture discussions and long-term platform strategy
  • Diagnose performance bottlenecks across distributed systems and HPC workloads
  • Support and optimize Slurm-based GPU cluster environments
  • Participate in an on-call rotation supporting mission-critical infrastructure operations

Skills/Must Have:

  • Deep experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or related fields
  • Strong experience supporting HPC or large-scale distributed compute environments
  • Deep Linux expertise (Ubuntu/Debian preferred)
  • Strong scripting and automation skills using Python, Go, or Bash
  • Hands-on experience with public cloud platforms or modern GPU cloud providers
  • Strong understanding of networking fundamentals (DNS, TCP/IP, routing, performance optimization)
  • Experience with Infrastructure-as-Code tooling such as Terraform and Ansible
  • Proven experience operating Slurm-based GPU/HPC clusters
  • Ability to troubleshoot distributed systems and optimize workload scheduling/performance

Benefits:

  • Stock options
  • Bonus 
  • Remote working option and allowance 

Salary:

  • Circa $250,000 base salary 
Vacancy posted more than 2 months ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer (GPU Clusters) - Hosting in San Francisco, CA vacancy
  • $175k - $250k

     ...50,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco,...  ...unavailable. Modality: On-Site only. Must live within...  ...scalability, performance, and reliability across environments....  ...Manage and automate GPU compute clusters using tools such as Python... 
    Senior
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    San Francisco, CA
    1 day ago
  •  ...grown 800% over the last 12 months. Engineering at Ivo Engineers at Ivo are...  ...without sacrificing accuracy [2024] Clustering legal documents descended from the same...  ...our SLAs. What? We’re looking for an Senior Site level Reliability Engineer as part of Infrastructure team... 
    Senior
    Contract work
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours

    Ivo Inc.

    San Francisco, CA
    4 days ago
  • $159.2k - $301.6k

     ...Graphs on the cloud. In this reliability-focused role, you will own...  ...services—working closely with cluster orchestrators like Kubernetes...  ...'ll partner with the backend engineers building these APIs to make...  ...~5-10 years of experience in site reliability engineering, infrastructure... 
    Senior
    Temporary work
    Local area
    Worldwide

    Adobe

    San Francisco, CA
    4 days ago
  • $181k - $263k

     ...providing first line operational support. We are looking for a Senior Staff Site Reliability Engineer who will set the technical direction for reliability...  ..., and friendly people who love what they do. Fun: We host in-person and virtual events such as game nights, happy... 
    Senior
    Work from home
    Flexible hours
    Night shift

    LiveRamp

    San Francisco, CA
    1 day ago
  • $220k

     ...Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience... 
    Senior

    Perplexity

    San Francisco, CA
    1 day ago
  • $300k

     ...training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and automation of this GPU-powered infrastructure, ensuring...  ...operational backbone of one of the largest GPU clusters in private deployment. If you want... 
    Senior
    Permanent employment
    San Francisco, CA
    more than 2 months ago
  • $205k - $235k

     ...management. We have become a multibillion-dollar asset manager, and we have ambitious goals for the future.  As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage... 
    Senior
    Remote job
    Local area

    The Voleon Group

    Berkeley, CA
    more than 2 months ago
  •  ...advanced models run efficiently, reliably, and at scale. We build and...  ...the Role We’re hiring engineers to scale and optimize OpenAI’...  ...infrastructure across emerging GPU platforms. You’ll work across...  ...model execution on large AMD GPU clusters. You can thrive in this... 
    Full time

    OpenAI

    San Francisco, CA
    23 hours ago
  •  ...us and help build the platform engineers turn to to ship AI products....  ...Baseten is building its own GPU infrastructure for large-scale...  ...tuning, bad optics, RNIC issues, host kernel stalls, GPU driver problems...  ...Staff-level or senior staff-level experience building... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    23 hours ago
  • $232k - $319k

     ...millions of users a day. The service is hosted on Amazon Web Services (AWS) across multiple...  ...scale the service with great people and reliable, cost-effective, and efficient...  ...Accelerate the velocity of SRE and product engineering by developing robust platforms, powerful... 
    Senior
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta, Inc.

    San Francisco, CA
    23 hours ago
  •  ...Conviction. Join us and help build the platform engineers turn to to ship AI products. At...  ...for foundational engineers to lead our GPU Networking efforts, making RDMA a first-class...  ...networking performance on bleeding-edge clusters (H100/H200, B200/B300, GB200/300 NVL72),... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    23 hours ago
  •  ...on running the world’s largest, most reliable, and frictionless GPU fleet to support OpenAI’s general purpose...  ...-button automation for kubernetes cluster provisioning and upgrades...  ...Much more! About the Role As an engineer within Fleet infrastructure, you will... 
    Full time
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    23 hours ago
  •  ...these hyperscale supercomputers reliable and efficient during the training...  ...the Role We are looking for engineers to operate the next generation of compute clusters that power OpenAI’s frontier research...  ...bare-metal Linux environments, GPU hardware, and large-scale... 
    Full time

    OpenAI

    San Francisco, CA
    23 hours ago
  • $350k

     ...Site Reliability Engineer (SRE) San Francisco Thinking Machines Lab's mission is to empower humanity through advancing collaborative...  ...scale: deploying, operating, debugging, and tuning clusters handling heterogeneous GPU workloads. Logistics Location: This role... 
    Local area
    Visa sponsorship
    Work visa
    Relocation package

    Thinking Machines Lab

    San Francisco, CA
    14 hours ago
  •  ...A tech startup in San Francisco is looking for Site Reliability Engineers to enhance system reliability and performance. Ideal candidates have over 5 years of relevant experience and strong expertise in cloud infrastructure, including AWS and Kubernetes. The role involves... 
    Senior

    Breakout Tools

    San Francisco, CA
    28 minutes ago
  •  ...Network and Positioning Engine deliver centimeter-...  ...automation. Role Outcome Senior DevOps Engineers own...  ...One’s services running reliably at scale - from the...  ...run reliably across all hosted environments, even under...  ...Experience with distributed and clustered data stores (e.g.,... 
    Senior
    Immediate start

    Point One Navigation

    San Francisco, CA
    15 days ago
  •  ...Site Reliability Engineer (SRE) We're looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You'll partner with engineers and data scientists to build, automate... 
    Senior

    Alembic Technologies

    San Francisco, CA
    4 days ago
  • $140k - $205k

     ...Senior Technology Site Reliability Engineer Cooley is seeking a Senior Site Reliability Engineer to join the Infrastructure & Development Operationsteam. Position summary: The Senior Technology Site Reliability Engineer("SRE") is responsible for ensuring the reliability... 
    Senior
    Full time
    Temporary work
    Work at office
    Flexible hours
    Weekend work

    Cooley

    Daly City, CA
    3 days ago
  •  ...come shape the future and be part of a truly unique global culture at OutSystems! Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of software engineering and applies them to infrastructure and... 
    Senior
    Immediate start
    Remote work
    Worldwide

    OutSystems

    San Francisco, CA
    1 day ago
  • $117k - $209.33k

     ...Job Requisition ID # 26WD99273 Position Overview Want to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products. As part of a... 
    Senior
    For contractors

    Autodesk

    San Francisco, CA
    14 hours ago
  • $81.1k - $187k

     ...Site Reliability Engineer 3 We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving... 
    Senior
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    San Francisco, CA
    23 hours ago
  • $148.5k - $223.9k

     ...duplicating efforts. Job Category Software Engineering Job Details About Salesforce Salesforce...  ...future of Salesforce. Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with... 
    Senior
    Worldwide
    Weekend work

    Salesforce.Com Inc

    San Francisco, CA
    2 days ago
  • $166.9k - $225.9k

     ...Summary: Drata's SRE team operates as both a central engineering function and an embedded reliability practice. You'll be part of a close-knit SRE team...  ...What you'll bring: ~6+ years of experience in Site Reliability Engineering, Cloud Engineering, or building... 
    Senior
    Work at office
    Immediate start
    Worldwide
    Monday to Friday
    Flexible hours

    Drata Inc

    San Francisco, CA
    1 day ago
  • $320k

     ...Anthropic’s mission is to create reliable, interpretable, and steerable...  ...of committed researchers, engineers, policy experts, and business...  ..., streaming model outputs, or GPU-based serving infrastructure...  ...extremely collaborative group, and we host frequent research discussions... 
    Senior
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    23 hours ago
  •  ...into a horizontal automation engine adopted by HR, Finance, Legal...  ...enabling and supporting self-hosted deployments for enterprise customers...  ..., performance, and reliability of production systems through...  ...Kubernetes in production, including cluster management and workload... 
    Senior
    Full time

    Serval

    San Francisco, CA
    23 hours ago
  •  ...The team is hiring a Head of Platform/AI Cluster Management to oversee the strategic...  ...), including multi-tenancy, quotas, and GPU/host fleet management. Lead cluster operations...  ...services that ensure workload SLOs and reliable runtime execution. Define and implement... 
    Permanent employment
    San Francisco, CA
    more than 2 months ago
  • $220k - $235k

     ...Staff/Senior Staff Site Reliability Engineer Ironclad is the leading AI contracting platform that transforms agreements into assets. Contracts move faster, insights surface instantly, and agents push work forward, all with you in control. Whether you're buying or selling... 
    Senior
    Full time
    Contract work
    Work at office

    Ironclad Inc

    San Francisco, CA
    1 day ago
  • $255k - $405k

     ...Lambda is the #1 GPU Cloud for ML/AI teams training, fine-tuning and inferencing AI models, where engineers can easily, securely and affordably build, test and deploy AI products...  ...portfolio includes on-prem GPU systems, hosted GPUs across public & private clouds and managed... 
    Senior
    Full time
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda

    San Francisco, CA
    23 hours ago
  • $210.8k - $272.8k

    About Thumbtack Thumbtack helps millions of people confidently care for their homes. About the Site Reliability Engineering Team The Site Reliability Engineering team focuses on creating and maintaining a reliable, secure, and scalable platform vital for a seamless user... 
    Senior
    Local area

    Thumbtack

    San Francisco, CA
    23 hours ago
  • $161.3k - $241.9k

     ...looking for a Production Engineer to help build and...  ...move quickly and operate reliable services at scale. In this...  ...platform, including cluster provisioning, upgrades,...  ...software, infrastructure, site reliability, or...  .... Experience operating GPU fleets, high-performance... 
    Senior

    Harvey

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer (GPU Clusters) - Hosting. Be the first to apply!