Senior SRE - AI Infra: Scale GPU Clusters & HPC
Hamilton Barnes Associates Limited
Hamilton Barnes Associates Limited is seeking an experienced SRE/Infrastructure Engineer to help build a seed-stage AI infrastructure company with large-scale GPU clusters for training and inference. You design, deploy, and maintain compute environments and ensure reliability across Slurm and Kubernetes. You will implement automation, observability, and IaC practices, collaborating with ML and platform teams to optimize GPU utilization, data flow, and latency. #J-18808-Ljbffr Hamilton Barnes Associates Limited
$250k
...opportunities? Join a rapidly scaling AI cloud infrastructure... ...a next-generation GPU platform designed for... ...company is looking for a Senior / Staff Site... ...and scale large-scale HPC and cloud environments... ...frameworks for GPU compute clusters Collaborate with ML,...SeniorFull timeRemote work$150k - $300k
...Intellect in San Francisco seeks a Solutions Architect for GPU Infrastructure who will transform client requirements into robust systems capable of training advanced AI models. Responsibilities include designing GPU cluster architectures, deploying orchestration systems, and...Suggested- ...distributed ML training and inference clusters Develop efficient, scalable... ...pipelines to manage petabyte-scale datasets and model training... ...Analyze, profile and debug low-level GPU operations to optimize... ...GCP, AWS, or Azure) and their ML/AI service offerings Familiarity...Senior
$210k - $240k
...high-performance AI platform supporting GPU infrastructure, distributed... ...deeply hands-on senior engineer with... .... General SRE, DevOps, cloud infrastructure... ...for GPU or HPC environments RDMA... ...validation GPU cluster, HPC, or AI/ML... ...startup or rapidly scaling technical...SeniorImmediate start$250k
...development? Join a seed-stage AI infrastructure company building large-scale training and inference... ...with a single managed GPU cluster that quickly reached... ...+ years of experience in SRE, DevOps, or Infrastructure... ...high-performance computing (HPC) or AI/ML training infrastructure...Senior- STN Inc in San Francisco is seeking an experienced AI Infrastructure Engineer to design, deploy, and manage large-scale GPU clusters for AI training and inference workloads. You will optimize GPU utilization, tune NCCL, CUDA, UCX, and Slurm, and work across storage, networking...Senior
- Magic AI, Inc. is seeking a engineer for the Supercomputing Platform & Infrastructure to design, build, and operate large-scale GPU infrastructure powering model training and inference. You will... ...environments, manage Kubernetes clusters, and ensure reproducibility and...Visa sponsorshipRelocation package
- A cutting-edge AI technology company based in San Francisco is seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate...Senior
- A leading AI technology company in San Francisco is looking for a Senior Software Engineer to build scalable infrastructure for large‑scale training and fine-tuning of foundation models. You will design... ...training systems and optimize GPU utilization while collaborating with...Senior
- Linuxcareers in San Francisco is building AI research infrastructure. You will design, deploy, and operate large-scale GPU clusters powering training, evaluation, and serving for the research team. The role emphasizes extending orchestration with Kubernetes/Slurm, building...
$250k
...development? Join a seed-stage AI infrastructure company building large-scale training and inference... ...with a single managed GPU cluster that quickly reached... ...+ years of experience in SRE, DevOps, or Infrastructure... ...high-performance computing (HPC) or AI/ML training...SeniorPermanent employment$160k - $225k
Cacheflow is seeking a Senior Software Engineer for AI Runtime at Databricks, located in San Francisco. You will be instrumental in building and scaling systems for large-scale GPU training, ensuring high throughput and resilience in training across expansive fleets of...Senior- Cogent, an Applied AI Lab in the cybersecurity space, is hiring a Senior Storage Infrastructure Engineer to own data storage, protection, and scale. You will implement backup and encryption policies across Aurora RDS, Clickhouse, and MSK, and build dashboards to monitor...Senior
$179k - $218k
...only vertically integrated AI infrastructure company... ...urgency, who believe in the scale of our ambition and... ...bridged.We are seeking a Senior Staff Data Center Operations Engineer, GPU Hardware Architecture to... ...needed to maintain peak cluster health.The Strategic BridgeFor...SeniorTemporary work$175k - $220k
Together AI is looking for a Product Manager to join our fast-growing team in San Francisco. In this role, you will drive day-to-... ...work across our AI infrastructure, focusing on observability and GPU clusters. The ideal candidate will possess strong data skills and will...- Oracle is seeking a senior leader to serve as the Strategic OCI & AI Infrastructure Lead within the Global AI Infrastructure... ...outcomes. The Lead will own and scale demanding AI and cloud... ...international markets, including GPU cluster architecture and multi-million-dollar...Senior
$15k
...company that applies state-of-the-art AI and machine learning techniques... ...catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to... ...work will provide a world-class HPC platform for researchers to...SeniorWork at officeLocal areaRemote work- Salesforce, Inc. is seeking a Senior Software Engineer to join the team responsible for the voice infrastructure at scale. The role focuses on deploying, maintaining, and monitoring voice services across multiple regions, ensuring high availability and quality for customers...Senior
- ...re looking for an experienced HPC infrastructure engineer to... ...probably the largest anime AI training cluster in the world . You’ll serve... ...our researchers and the bare GPU machines, helping to make sure... ...your love of anime and large-scale GPU systems. You’re familiar...Work at officeVisa sponsorship
$152.5k - $205k
...power trusted, internet-scale financial innovation. Learn... ...be responsible for:As a Senior Site Reliability... ...critical digital-assets, AI, and application workloads... ...role is for an experienced SRE or infrastructure... ...troubleshooting production clusters and containerized workloads...SeniorFlexible hours- ...to design, deploy, and operate large-scale storage for FAC and related platforms,... ...integration with the CoreHPC compute cluster. This role supports AI, data science, and computational research... ...teams. Experience with ZFS, VAST, and HPC schedulers is essential. #J-18808-...
$150k - $300k
...agentic models to the infra that enables anyone... ...at frontier scale, adapting models to... ...Solutions Architect for GPU Infrastructure, you... ...’s most advanced AI models. We recently... ...design optimal GPU cluster architectures Create... ..., inference, and HPC workloads Present...$200k - $400k
...vLLM as the world's AI inference engine and... ...looking for a hands-on cluster administration... ...the high-performance GPU compute infrastructure... ...performance GPU and HPC clusters across neo-... ...operate, debug, and scale compute across providers... ...ML infrastructure, SRE, platform...Remote workVisa sponsorship$250k
...career? Join a rapidly scaling AI cloud infrastructure... ...building next-generation GPU platforms for large-... ...company is looking for a Senior Storage Engineer with experience... ...supporting AI and HPC workloads Manage and... ...training and inference clusters Work closely with...SeniorFull timeRemote work- ...Cloud, is a leader in AI cloud... ...superintelligence. One person, one GPU. If you'd like... ...We are seeking a Senior Software Engineer... ...training and inference at scale. As a Senior... ...for end-to-end cluster lifecycle management... ...Familiarity with HPC and traditional job...SeniorFull timeWork at officeLocal areaWork from homeFlexible hours
- ...Cloud, is a leader in AI cloud... ...superintelligence. One person, one GPU. If you'd like... ...-performance AI clusters by welding together... ...for an experienced Senior Software Engineer... ...that provisions, scales, heals, and meters... ...Ceph at 100 PB+ in HPC or AI environments....SeniorFull timeWork at officeLocal areaWork from homeFlexible hours
- ...Superintelligence Cloud, is a leader in AI cloud infrastructure... .... One person, one GPU. If you'd like to... ...Help to build and scale Lambda's high performance... ...hardware for new and existing clusters Ensure high... ...Ansible/Salt Hands-on with HPC/AI networking: RoCEv2 and...SeniorWork at officeLocal areaWork from homeFlexible hours
- Inferact is seeking a hands-on cluster administration engineer to own and operate its high-performance GPU compute infrastructure. You will ensure health, availability,... ...standardize provisioning, operation, debugging, and scaling of compute resources. The role directly...Senior
- A pioneering technology company in San Francisco is seeking a Head of Platform/AI Cluster Management to lead AI and platform initiatives. Responsibilities include overseeing scheduler management and optimizing performance across cross-functional teams. The ideal candidate...
- ...people interact with the web by building AI agents that can reliably do everyday... ...Michele Catasta, etc. Responsibilities: Scale infra for post-training of multimodal LLMs (CPT... ...for: Experience with ML infrastructure (GPU clusters) and supporting networking (NCCL) Experience...Work at officeRelocationVisa sponsorship
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior SRE - AI Infra: Scale GPU Clusters & HPC. Be the first to apply!
- senior network engineer remote San Francisco, CA
- senior app developer San Francisco, CA
- senior manager legal San Francisco, CA
- sr project manager San Francisco, CA
- senior account executive San Francisco, CA
- senior manager strategic initiatives San Francisco, CA
- senior staff systems engineer San Francisco, CA
- senior commercial counsel San Francisco, CA
- senior data manager San Francisco, CA
- senior safety manager San Francisco, CA




