Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead AI Infrastructure Engineer: GPU Clusters & Reliability

Luma AI

Luma AI in San Francisco is seeking a leader to define reliability for a frontier AI infrastructure. You will architect and operate large GPU environments, pushing the limits of training and inference while partnering with research and product to scale systems and improve availability. The role demands deep Linux/distributed systems expertise, strong Kubernetes mastery, and a track record of delivering scalable, production‑grade infrastructure. #J-18808-Ljbffr Luma AI

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Lead AI Infrastructure Engineer: GPU Clusters & Reliability in San Francisco, CA vacancy
  • Accenture is seeking a seasoned AI Infrastructure Architect in San Francisco to design and implement...  ..., including accelerated computing clusters, model serving endpoints, and secure governance...  ...The role requires deep experience with GPU/DPUs, networking, storage, and... 
    Suggested

    Accenture

    San Francisco, CA
    9 hours ago
  •  ...partnering with a rapidly growing AI infrastructure company to own the core cluster infrastructure powering a heterogeneous...  ...cloud. You will manage large CPU/GPU/accelerator clusters, bare‑metal...  ...schedulers. This role focuses on reliability, observability, and automation to... 
    Suggested

    Acceler8 Talent

    San Francisco, CA
    9 hours ago
  • Together Computer Inc is hiring a Customer Support Engineer in San Francisco. The ideal candidate will support customers with AI solutions and tackle complex technical challenges related to GPU clusters. With a strong foundation in AI and customer service, you'll work closely... 
    Suggested
    Remote job

    Together Computer Inc

    San Francisco, CA
    9 hours ago
  • $225k - $250k

     ...year. Cline is the leading open‑source AI coding agent. Trusted...  ...the world's largest engineering organizations and individual...  ...code on your infrastructure. Teams choose Cline...  ...ll design for scale, reliability, and developer...  ...resource utilization (GPU scheduling, autoscaling... 
    Suggested

    Céline

    San Francisco, CA
    9 hours ago
  • $250k

     ...opportunities? Join a rapidly scaling AI cloud infrastructure provider building a next-generation GPU platform designed for AI...  ...for a Senior / Staff Site Reliability Engineer to support and scale large-...  ...frameworks for GPU compute clusters Collaborate with ML, data,... 
    Suggested
    Permanent employment
    Remote work
    San Francisco, CA
    a month ago
  •  ...the next generation of enterprise AI infrastructure. As a Staff AI Infrastructure Engineer, you will design, build, and...  ...distributed systems, Kubernetes, GPU infrastructure, high‑performance...  ...to deliver secure, scalable, and reliable systems spanning edge deployments... 

    Seekr

    San Francisco, CA
    9 hours ago
  • Sciforium is looking for a Senior HPC & GPU Infrastructure Engineer to manage the health and performance of our GPU compute cluster. This role involves hands-on Linux systems engineering and maintaining the ML software stack including CUDA and PyTorch. The ideal candidate... 
    Flexible hours

    Sciforium

    San Francisco, CA
    9 hours ago
  • Beam is seeking a role-focused engineer to own the health and reliability of our GPU compute fleet in a fast-growing AI inference platform. You will build and own metrics pipelines, alerts, and a unified health view across thousands of GPUs in production. You will automate... 

    Beam

    San Francisco, CA
    9 hours ago
  • $190k - $270k

    AI Chopping Block, Inc. is hiring an AI Infrastructure Engineer to ensure smooth operations of user-facing services and production systems in San Francisco. The...  ...implementing best practices for availability and reliability. Competitive salary range is $190,000 - $270,000... 

    AI Chopping Block, Inc.

    San Francisco, CA
    4 days ago
  •  ...intersection of platform engineering, site reliability, and applied ML...  ...operability of Meshy’s AI model serving stack,...  ...with core engineering infrastructure. The team operates a...  ...governance. Develop CPU/GPU resource‑management...  ...and training share a cluster. Drive unified... 
    Work at office
    Remote work
    Flexible hours

    MeshyAI

    San Francisco, CA
    9 hours ago
  • About Luma AI A new class of intelligence...  .... It is an infrastructure challenge at the edge...  ...rapidly scaling 10k+ GPU fleets, pushing...  ..., throughput, and reliability hard enough that yesterday...  ...Infrastructure Engineering team is a systems...  ...evolve as cluster size and concurrency... 

    Luma

    San Francisco, CA
    9 hours ago
  • cursor in New York, New York, seeks talented engineers to enhance their ML Infrastructure team, focusing on building a robust and scalable coding model...  ...researchers, improve training systems, and automate GPU cluster management while collaborating in a flat organizational... 

    cursor

    San Francisco, CA
    5 days ago
  • arcee.ai is seeking a Compute Infra Specialist in San Francisco to help manage and scale the infrastructure for AI workloads. This hands-on role involves coordinating across teams and ensuring efficient GPU resource allocation while enhancing customer experience with our... 

    arcee.ai

    San Francisco, CA
    1 day ago
  • Together AI is seeking an AI Infrastructure Engineer to keep production systems running smoothly and automate operations with mature engineering discipline...  ...software skills with pragmatic operations and design reliable, scalable systems for high concurrency. Required: 5+... 

    AI Chopping Block

    San Francisco, CA
    2 days ago
  • $216k - $270k

    As a Software Engineer on the Machine Learning Infrastructure team, you will build the “Operating...  ...” for our large-scale GPU clusters. You will architect a high...  ..., networking, and reliability challenges that emerge at...  ...compute into breakthrough AI. You will: Architect and... 
    Full time
    For contractors

    Scale AI

    San Francisco, CA
    9 hours ago
  • Brain Co. in San Francisco is seeking a Backend Engineer to build the technical capabilities for scaling AI products. The role involves designing and maintaining...  ...in building production systems, is driven by reliability and usability, and thrives in collaborative environments... 

    brainco

    San Francisco, CA
    9 hours ago
  •  ...Francisco Compute seeks an experienced systems software engineer to contribute to its UEFI bare-metal infrastructure, VM platform, and other low-level components. You...  ...fault-tolerant distributed systems, work with GPU clusters, and implement Kubernetes operators, while... 

    San Francisco Compute

    San Francisco, CA
    2 days ago
  • OpenAI is seeking an Infrastructure Operations Engineer to operate and improve large-...  ...Ethernet fabrics supporting GPU clusters, storage, and management...  ...response across a global AI network. You will partner...  ...teams to raise reliability, perform RCA, and automate... 

    OpenAI

    San Francisco, CA
    3 days ago
  • $180k - $200k

     ...States Who We Are Lightning AI is the company behind PyTorch...  ...For Lightning AI is seeking a GPU & Compute Infrastructure Engineer to join our Infrastructure Engineering...  ...automation, improving reliability, and enabling efficient cluster bring‑up for AI/ML and HPC workloads... 
    Remote work
    Work from home
    Flexible hours

    Lightning AI

    San Francisco, CA
    9 hours ago
  • techire ai is seeking an experienced ML Platform engineer to craft scalable infrastructure for training, evaluation, and production deployment of frontier AI models in a highly...  .... You will tackle distributed training, GPU orchestration, scheduling, and fault-tolerant... 

    techire ai

    San Francisco, CA
    3 days ago
  • $177.5k - $248k

     ...understanding in healthcare. Our AI-powered platform was...  ..., technologists, and engineers working together to...  .... The Role As an AI Infrastructure Engineer at Abridge,...  ...scalable Kubernetes clusters for AI model inference...  ...workflows and enhance GPU utilization for ML workloads... 
    Hourly pay
    Full time
    Flexible hours

    Abridge

    San Francisco, CA
    2 days ago
  • A leading AI research company in San Francisco is seeking a software engineer for its Fleet High Performance Computing team. In this role, you'll ensure the reliability and uptime of the compute fleet, working with automation systems and monitoring tools. Ideal candidates... 

    Jobleads-US

    San Francisco, CA
    1 day ago
  • Gravity Engineering Services Pvt Ltd. is seeking a Senior Software Engineer for AI Runtime in San Francisco, California. The role involves...  ...architecture of a managed GPU training platform, solving...  ...and enhancing performance and reliability. Ideal candidates will have 5... 

    Gravity Engineering Services Pvt Ltd.

    San Francisco, CA
    9 hours ago
  • $224k - $284k

     ...Atoms builds Physical AI - real‑world robots...  ...into something more reliable, more scalable, and more...  ...We are roboticists, engineers, operators, and builders...  ...and automate our GPU training clusters, including provisioning...  .... Design our infrastructure to scale smoothly as... 
    Full time
    Work at office
    Flexible hours

    Atoms

    San Francisco, CA
    9 hours ago
  • $190k - $270k

    AI Chopping Block, Inc. is hiring an AI Infrastructure Engineer in San Francisco, California. This full-time role involves ensuring smooth operation of user-facing services and production systems, alongside building and running infrastructure with Ansible, Terraform, and... 
    Full time

    AI Chopping Block, Inc.

    San Francisco, CA
    4 days ago
  • An innovative studio is seeking an AI Infrastructure Engineer to enhance their ML infrastructure for groundbreaking anime games. This role involves designing and implementing cutting-edge inference architectures to support various platforms. As part of a small, agile team... 
    Worldwide

    Spellbrush

    San Francisco, CA
    9 hours ago
  •  ...happen to be the world's leading generative AI studio — we're the team behind...  ..., is looking for an AI Infrastructure Engineer to join us in building out...  ...both on-prem and multi-cloud clusters. But most importantly, you...  ...understanding of GPU’s handling large workloads... 
    Work experience placement
    Work at office
    Visa sponsorship

    Spellbrush

    San Francisco, CA
    3 days ago
  • $160k - $230k

    An innovative AI technology firm is seeking a passionate Customer Support Engineer to tackle complex technical challenges. You will support customers in building solutions and collaborate with various teams to enhance customer satisfaction. The role requires 3+ years in... 

    Together AI

    San Francisco, CA
    9 hours ago
  • A leading technology firm in San Francisco is seeking a Senior Site Reliability Engineer to maintain and improve cloud infrastructure. The ideal candidate has over 5 years of experience as an SRE or DevOps engineer and strong expertise in Kubernetes. This role focuses on... 

    TechChain Talent

    San Francisco, CA
    9 hours ago
  • $255k

     ...hyperscale supercomputers reliable and efficient...  ...are looking for engineers to operate the...  ...generation of compute clusters that power OpenAI'...  ...with hands-on infrastructure work on our largest...  ...Linux environments, GPU hardware, and large...  ...OpenAI is an AI research and deployment... 

    OpenAI

    San Francisco, CA
    9 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead AI Infrastructure Engineer: GPU Clusters & Reliability. Be the first to apply!