Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior SRE - AI Infra: Scale GPU Clusters & HPC

Hamilton Barnes Associates Limited

Hamilton Barnes Associates Limited is seeking an experienced SRE/Infrastructure Engineer to help build a seed-stage AI infrastructure company with large-scale GPU clusters for training and inference. You design, deploy, and maintain compute environments and ensure reliability across Slurm and Kubernetes. You will implement automation, observability, and IaC practices, collaborating with ML and platform teams to optimize GPU utilization, data flow, and latency. #J-18808-Ljbffr Hamilton Barnes Associates Limited

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior SRE - AI Infra: Scale GPU Clusters & HPC in San Francisco, CA vacancy
  • $250k

     ...opportunities? Join a rapidly scaling AI cloud infrastructure...  ...a next-generation GPU platform designed for...  ...company is looking for a Senior / Staff Site...  ...and scale large-scale HPC and cloud environments...  ...frameworks for GPU compute clusters Collaborate with ML,... 
    Senior
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  • $150k - $300k

     ...Intellect in San Francisco seeks a Solutions Architect for GPU Infrastructure who will transform client requirements into robust systems capable of training advanced AI models. Responsibilities include designing GPU cluster architectures, deploying orchestration systems, and... 
    Suggested

    Prime Intellect

    San Francisco, CA
    4 days ago
  •  ...distributed ML training and inference clusters Develop efficient, scalable...  ...pipelines to manage petabyte-scale datasets and model training...  ...Analyze, profile and debug low-level GPU operations to optimize...  ...GCP, AWS, or Azure) and their ML/AI service offerings Familiarity... 
    Senior

    Kindredventures

    San Francisco, CA
    2 days ago
  • $210k - $240k

     ...high-performance AI platform supporting GPU infrastructure, distributed...  ...deeply hands-on senior engineer with...  .... General SRE, DevOps, cloud infrastructure...  ...for GPU or HPC environments RDMA...  ...validation GPU cluster, HPC, or AI/ML...  ...startup or rapidly scaling technical... 
    Senior
    Immediate start

    Stratitech

    San Francisco, CA
    3 days ago
  • $250k

     ...development? Join a seed-stage AI infrastructure company building large-scale training and inference...  ...with a single managed GPU cluster that quickly reached...  ...+ years of experience in SRE, DevOps, or Infrastructure...  ...high-performance computing (HPC) or AI/ML training infrastructure... 
    Senior

    Hamilton Barnes Associates Limited

    San Francisco, CA
    1 day ago
  • STN Inc in San Francisco is seeking an experienced AI Infrastructure Engineer to design, deploy, and manage large-scale GPU clusters for AI training and inference workloads. You will optimize GPU utilization, tune NCCL, CUDA, UCX, and Slurm, and work across storage, networking... 
    Senior

    STN Inc

    San Francisco, CA
    2 days ago
  • Magic AI, Inc. is seeking a engineer for the Supercomputing Platform & Infrastructure to design, build, and operate large-scale GPU infrastructure powering model training and inference. You will...  ...environments, manage Kubernetes clusters, and ensure reproducibility and... 
    Visa sponsorship
    Relocation package

    Magic AI Corp.

    San Francisco, CA
    3 days ago
  • A cutting-edge AI technology company based in San Francisco is seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate... 
    Senior

    Reflection AI

    San Francisco, CA
    17 hours ago
  • A leading AI technology company in San Francisco is looking for a Senior Software Engineer to build scalable infrastructure for large‑scale training and fine-tuning of foundation models. You will design...  ...training systems and optimize GPU utilization while collaborating with... 
    Senior

    Baseten

    San Francisco, CA
    17 hours ago
  • Linuxcareers in San Francisco is building AI research infrastructure. You will design, deploy, and operate large-scale GPU clusters powering training, evaluation, and serving for the research team. The role emphasizes extending orchestration with Kubernetes/Slurm, building... 

    Linuxcareers

    San Francisco, CA
    1 day ago
  • $250k

     ...development? Join a seed-stage AI infrastructure company building large-scale training and inference...  ...with a single managed GPU cluster that quickly reached...  ...+ years of experience in SRE, DevOps, or Infrastructure...  ...high-performance computing (HPC) or AI/ML training... 
    Senior
    Permanent employment
    San Francisco, CA
    4 days ago
  • $160k - $225k

    Cacheflow is seeking a Senior Software Engineer for AI Runtime at Databricks, located in San Francisco. You will be instrumental in building and scaling systems for large-scale GPU training, ensuring high throughput and resilience in training across expansive fleets of... 
    Senior

    Cacheflow

    San Francisco, CA
    2 days ago
  • Cogent, an Applied AI Lab in the cybersecurity space, is hiring a Senior Storage Infrastructure Engineer to own data storage, protection, and scale. You will implement backup and encryption policies across Aurora RDS, Clickhouse, and MSK, and build dashboards to monitor... 
    Senior

    Cogent

    San Francisco, CA
    3 days ago
  • $179k - $218k

     ...only vertically integrated AI infrastructure company...  ...urgency, who believe in the scale of our ambition and...  ...bridged.We are seeking a Senior Staff Data Center Operations Engineer, GPU Hardware Architecture to...  ...needed to maintain peak cluster health.The Strategic BridgeFor... 
    Senior
    Temporary work

    Crusoe

    San Francisco, CA
    5 days ago
  • $175k - $220k

    Together AI is looking for a Product Manager to join our fast-growing team in San Francisco. In this role, you will drive day-to-...  ...work across our AI infrastructure, focusing on observability and GPU clusters. The ideal candidate will possess strong data skills and will... 

    Together AI

    San Francisco, CA
    17 hours ago
  • Oracle is seeking a senior leader to serve as the Strategic OCI & AI Infrastructure Lead within the Global AI Infrastructure...  ...outcomes. The Lead will own and scale demanding AI and cloud...  ...international markets, including GPU cluster architecture and multi-million-dollar... 
    Senior

    Socket

    San Francisco, CA
    3 days ago
  • $15k

     ...company that applies state-of-the-art AI and machine learning techniques...  ...catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to...  ...work will provide a world-class HPC platform for researchers to... 
    Senior
    Work at office
    Local area
    Remote work

    The Voleon Group

    Berkeley, CA
    6 days ago
  • Salesforce, Inc. is seeking a Senior Software Engineer to join the team responsible for the voice infrastructure at scale. The role focuses on deploying, maintaining, and monitoring voice services across multiple regions, ensuring high availability and quality for customers... 
    Senior

    Salesforce

    San Francisco, CA
    2 days ago
  •  ...re looking for an experienced HPC infrastructure engineer to...  ...probably the largest anime AI training cluster in the world . You’ll serve...  ...our researchers and the bare GPU machines, helping to make sure...  ...your love of anime and large-scale GPU systems. You’re familiar... 
    Work at office
    Visa sponsorship

    Spellbrush

    San Francisco, CA
    24 days ago
  • $152.5k - $205k

     ...power trusted, internet-scale financial innovation. Learn...  ...be responsible for:As a Senior Site Reliability...  ...critical digital-assets, AI, and application workloads...  ...role is for an experienced SRE or infrastructure...  ...troubleshooting production clusters and containerized workloads... 
    Senior
    Flexible hours

    Circle

    San Francisco, CA
    7 days ago
  •  ...to design, deploy, and operate large-scale storage for FAC and related platforms,...  ...integration with the CoreHPC compute cluster. This role supports AI, data science, and computational research...  ...teams. Experience with ZFS, VAST, and HPC schedulers is essential. #J-18808-... 

    UCSF Health

    San Francisco, CA
    4 days ago
  • $150k - $300k

     ...agentic models to the infra that enables anyone...  ...at frontier scale, adapting models to...  ...Solutions Architect for GPU Infrastructure, you...  ...’s most advanced AI models. We recently...  ...design optimal GPU cluster architectures Create...  ..., inference, and HPC workloads Present... 

    Prime Intellect

    San Francisco, CA
    4 days ago
  • $200k - $400k

     ...vLLM as the world's AI inference engine and...  ...looking for a hands-on cluster administration...  ...the high-performance GPU compute infrastructure...  ...performance GPU and HPC clusters across neo-...  ...operate, debug, and scale compute across providers...  ...ML infrastructure, SRE, platform... 
    Remote work
    Visa sponsorship

    Inferact Inc.

    San Francisco, CA
    2 days ago
  • $250k

     ...career? Join a rapidly scaling AI cloud infrastructure...  ...building next-generation GPU platforms for large-...  ...company is looking for a Senior Storage Engineer with experience...  ...supporting AI and HPC workloads Manage and...  ...training and inference clusters Work closely with... 
    Senior
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  •  ...Cloud, is a leader in AI cloud...  ...superintelligence. One person, one GPU. If you'd like...  ...We are seeking a Senior Software Engineer...  ...training and inference at scale. As a Senior...  ...for end-to-end cluster lifecycle management...  ...Familiarity with HPC and traditional job... 
    Senior
    Full time
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda

    San Francisco, CA
    17 hours ago
  •  ...Cloud, is a leader in AI cloud...  ...superintelligence. One person, one GPU. If you'd like...  ...-performance AI clusters by welding together...  ...for an experienced Senior Software Engineer...  ...that provisions, scales, heals, and meters...  ...Ceph at 100 PB+ in HPC or AI environments.... 
    Senior
    Full time
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda

    San Francisco, CA
    17 hours ago
  •  ...Superintelligence Cloud, is a leader in AI cloud infrastructure...  .... One person, one GPU. If you'd like to...  ...Help to build and scale Lambda's high performance...  ...hardware for new and existing clusters Ensure high...  ...Ansible/Salt Hands-on with HPC/AI networking: RoCEv2 and... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Corporation

    San Francisco, CA
    1 day ago
  • Inferact is seeking a hands-on cluster administration engineer to own and operate its high-performance GPU compute infrastructure. You will ensure health, availability,...  ...standardize provisioning, operation, debugging, and scaling of compute resources. The role directly... 
    Senior

    Inferact

    San Francisco, CA
    4 days ago
  • A pioneering technology company in San Francisco is seeking a Head of Platform/AI Cluster Management to lead AI and platform initiatives. Responsibilities include overseeing scheduler management and optimizing performance across cross-functional teams. The ideal candidate... 

    Hamilton Barnes Associates Limited

    San Francisco, CA
    1 day ago
  •  ...people interact with the web by building AI agents that can reliably do everyday...  ...Michele Catasta, etc. Responsibilities: Scale infra for post-training of multimodal LLMs (CPT...  ...for: Experience with ML infrastructure (GPU clusters) and supporting networking (NCCL) Experience... 
    Work at office
    Relocation
    Visa sponsorship

    Yutori

    San Francisco, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior SRE - AI Infra: Scale GPU Clusters & HPC. Be the first to apply!