Senior Site Reliability Engineer (GPU Clusters) - Hosting
$250kLooking for a role with plenty of growth opportunities?
Join a rapidly scaling AI cloud infrastructure provider building a next-generation GPU platform designed for AI training, experimentation, and inference at scale. The company is developing a fully featured AI cloud platform powered by renewable energy and is already operating with strong momentum across Europe, while now significantly expanding its footprint in the United States.
The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments powering GPU-intensive workloads. The role involves working closely with platform, ML, and infrastructure teams to improve reliability, automation, and observability across distributed compute environments while supporting long-term infrastructure growth and scalability.
Don’t miss out on this exciting opportunity and apply today!
Responsibilities:
- Ensure the reliability, scalability, and performance of HPC and cloud infrastructure environments
- Design, build, and maintain automation, observability, and monitoring frameworks for GPU compute clusters
- Collaborate with ML, data, and platform engineering teams to deliver highly available infrastructure systems
- Improve CI/CD pipelines, deployment workflows, and operational tooling
- Contribute to infrastructure architecture discussions and long-term platform strategy
- Diagnose performance bottlenecks across distributed systems and HPC workloads
- Support and optimize Slurm-based GPU cluster environments
- Participate in an on-call rotation supporting mission-critical infrastructure operations
Skills/Must Have:
- Deep experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or related fields
- Strong experience supporting HPC or large-scale distributed compute environments
- Deep Linux expertise (Ubuntu/Debian preferred)
- Strong scripting and automation skills using Python, Go, or Bash
- Hands-on experience with public cloud platforms or modern GPU cloud providers
- Strong understanding of networking fundamentals (DNS, TCP/IP, routing, performance optimization)
- Experience with Infrastructure-as-Code tooling such as Terraform and Ansible
- Proven experience operating Slurm-based GPU/HPC clusters
- Ability to troubleshoot distributed systems and optimize workload scheduling/performance
Benefits:
- Stock options
- Bonus
- Remote working option and allowance
Salary:
- Circa $250,000 base salary
$175k - $250k
...50,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco,... ...unavailable. Modality: On-Site only. Must live within... ...scalability, performance, and reliability across environments.... ...Manage and automate GPU compute clusters using tools such as Python...SeniorFull timeRemote workRelocationRelocation package- ...grown 800% over the last 12 months. Engineering at Ivo Engineers at Ivo are... ...without sacrificing accuracy [2024] Clustering legal documents descended from the same... ...our SLAs. What? We’re looking for an Senior Site level Reliability Engineer as part of Infrastructure team...SeniorContract workWork at officeRemote workVisa sponsorshipRelocation packageFlexible hours
$159.2k - $301.6k
...Graphs on the cloud. In this reliability-focused role, you will own... ...services—working closely with cluster orchestrators like Kubernetes... ...'ll partner with the backend engineers building these APIs to make... ...~5-10 years of experience in site reliability engineering, infrastructure...SeniorTemporary workLocal areaWorldwide$181k - $263k
...providing first line operational support. We are looking for a Senior Staff Site Reliability Engineer who will set the technical direction for reliability... ..., and friendly people who love what they do. Fun: We host in-person and virtual events such as game nights, happy...SeniorWork from homeFlexible hoursNight shift$220k
...Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience...Senior$300k
...training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and automation of this GPU-powered infrastructure, ensuring... ...operational backbone of one of the largest GPU clusters in private deployment. If you want...SeniorPermanent employment$205k - $235k
...management. We have become a multibillion-dollar asset manager, and we have ambitious goals for the future. As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage...SeniorRemote jobLocal area- ...advanced models run efficiently, reliably, and at scale. We build and... ...the Role We’re hiring engineers to scale and optimize OpenAI’... ...infrastructure across emerging GPU platforms. You’ll work across... ...model execution on large AMD GPU clusters. You can thrive in this...Full time
- ...us and help build the platform engineers turn to to ship AI products.... ...Baseten is building its own GPU infrastructure for large-scale... ...tuning, bad optics, RNIC issues, host kernel stalls, GPU driver problems... ...Staff-level or senior staff-level experience building...Full timeFlexible hours
$232k - $319k
...millions of users a day. The service is hosted on Amazon Web Services (AWS) across multiple... ...scale the service with great people and reliable, cost-effective, and efficient... ...Accelerate the velocity of SRE and product engineering by developing robust platforms, powerful...SeniorPermanent employmentLocal areaWorldwideFlexible hours- ...Conviction. Join us and help build the platform engineers turn to to ship AI products. At... ...for foundational engineers to lead our GPU Networking efforts, making RDMA a first-class... ...networking performance on bleeding-edge clusters (H100/H200, B200/B300, GB200/300 NVL72),...Full timeFlexible hours
- ...on running the world’s largest, most reliable, and frictionless GPU fleet to support OpenAI’s general purpose... ...-button automation for kubernetes cluster provisioning and upgrades... ...Much more! About the Role As an engineer within Fleet infrastructure, you will...Full timeWork at officeRelocation package
- ...these hyperscale supercomputers reliable and efficient during the training... ...the Role We are looking for engineers to operate the next generation of compute clusters that power OpenAI’s frontier research... ...bare-metal Linux environments, GPU hardware, and large-scale...Full time
$350k
...Site Reliability Engineer (SRE) San Francisco Thinking Machines Lab's mission is to empower humanity through advancing collaborative... ...scale: deploying, operating, debugging, and tuning clusters handling heterogeneous GPU workloads. Logistics Location: This role...Local areaVisa sponsorshipWork visaRelocation package- ...A tech startup in San Francisco is looking for Site Reliability Engineers to enhance system reliability and performance. Ideal candidates have over 5 years of relevant experience and strong expertise in cloud infrastructure, including AWS and Kubernetes. The role involves...Senior
- ...Network and Positioning Engine deliver centimeter-... ...automation. Role Outcome Senior DevOps Engineers own... ...One’s services running reliably at scale - from the... ...run reliably across all hosted environments, even under... ...Experience with distributed and clustered data stores (e.g.,...SeniorImmediate start
- ...Site Reliability Engineer (SRE) We're looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You'll partner with engineers and data scientists to build, automate...Senior
$140k - $205k
...Senior Technology Site Reliability Engineer Cooley is seeking a Senior Site Reliability Engineer to join the Infrastructure & Development Operationsteam. Position summary: The Senior Technology Site Reliability Engineer("SRE") is responsible for ensuring the reliability...SeniorFull timeTemporary workWork at officeFlexible hoursWeekend work- ...come shape the future and be part of a truly unique global culture at OutSystems! Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of software engineering and applies them to infrastructure and...SeniorImmediate startRemote workWorldwide
$117k - $209.33k
...Job Requisition ID # 26WD99273 Position Overview Want to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products. As part of a...SeniorFor contractors$81.1k - $187k
...Site Reliability Engineer 3 We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving...SeniorTemporary workImmediate startFlexible hoursShift work$148.5k - $223.9k
...duplicating efforts. Job Category Software Engineering Job Details About Salesforce Salesforce... ...future of Salesforce. Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with...SeniorWorldwideWeekend work$166.9k - $225.9k
...Summary: Drata's SRE team operates as both a central engineering function and an embedded reliability practice. You'll be part of a close-knit SRE team... ...What you'll bring: ~6+ years of experience in Site Reliability Engineering, Cloud Engineering, or building...SeniorWork at officeImmediate startWorldwideMonday to FridayFlexible hours$320k
...Anthropic’s mission is to create reliable, interpretable, and steerable... ...of committed researchers, engineers, policy experts, and business... ..., streaming model outputs, or GPU-based serving infrastructure... ...extremely collaborative group, and we host frequent research discussions...SeniorFull timeWork at officeVisa sponsorshipFlexible hours- ...into a horizontal automation engine adopted by HR, Finance, Legal... ...enabling and supporting self-hosted deployments for enterprise customers... ..., performance, and reliability of production systems through... ...Kubernetes in production, including cluster management and workload...SeniorFull time
- ...The team is hiring a Head of Platform/AI Cluster Management to oversee the strategic... ...), including multi-tenancy, quotas, and GPU/host fleet management. Lead cluster operations... ...services that ensure workload SLOs and reliable runtime execution. Define and implement...Permanent employment
$220k - $235k
...Staff/Senior Staff Site Reliability Engineer Ironclad is the leading AI contracting platform that transforms agreements into assets. Contracts move faster, insights surface instantly, and agents push work forward, all with you in control. Whether you're buying or selling...SeniorFull timeContract workWork at office$255k - $405k
...Lambda is the #1 GPU Cloud for ML/AI teams training, fine-tuning and inferencing AI models, where engineers can easily, securely and affordably build, test and deploy AI products... ...portfolio includes on-prem GPU systems, hosted GPUs across public & private clouds and managed...SeniorFull timeWork at officeLocal areaWork from homeFlexible hours$210.8k - $272.8k
About Thumbtack Thumbtack helps millions of people confidently care for their homes. About the Site Reliability Engineering Team The Site Reliability Engineering team focuses on creating and maintaining a reliable, secure, and scalable platform vital for a seamless user...SeniorLocal area$161.3k - $241.9k
...looking for a Production Engineer to help build and... ...move quickly and operate reliable services at scale. In this... ...platform, including cluster provisioning, upgrades,... ...software, infrastructure, site reliability, or... .... Experience operating GPU fleets, high-performance...Senior
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer (GPU Clusters) - Hosting. Be the first to apply!
- site reliability engineer San Francisco, CA
- site reliability engineer sre San Francisco, CA
- site reliability engineer remote San Francisco, CA
- senior business analyst San Francisco, CA
- senior cost estimator San Francisco, CA
- senior manager tax San Francisco, CA
- senior automation engineer San Francisco, CA
- senior devops San Francisco, CA
- senior recruiter San Francisco, CA
- senior property manager San Francisco, CA



