Lead AI Infrastructure Engineer: GPU Clusters & Reliability
Luma AI
Luma AI in San Francisco is seeking a leader to define reliability for a frontier AI infrastructure. You will architect and operate large GPU environments, pushing the limits of training and inference while partnering with research and product to scale systems and improve availability. The role demands deep Linux/distributed systems expertise, strong Kubernetes mastery, and a track record of delivering scalable, production‑grade infrastructure. #J-18808-Ljbffr Luma AI
- Accenture is seeking a seasoned AI Infrastructure Architect in San Francisco to design and implement... ..., including accelerated computing clusters, model serving endpoints, and secure governance... ...The role requires deep experience with GPU/DPUs, networking, storage, and...Suggested
- ...partnering with a rapidly growing AI infrastructure company to own the core cluster infrastructure powering a heterogeneous... ...cloud. You will manage large CPU/GPU/accelerator clusters, bare‑metal... ...schedulers. This role focuses on reliability, observability, and automation to...Suggested
- Together Computer Inc is hiring a Customer Support Engineer in San Francisco. The ideal candidate will support customers with AI solutions and tackle complex technical challenges related to GPU clusters. With a strong foundation in AI and customer service, you'll work closely...SuggestedRemote job
$225k - $250k
...year. Cline is the leading open‑source AI coding agent. Trusted... ...the world's largest engineering organizations and individual... ...code on your infrastructure. Teams choose Cline... ...ll design for scale, reliability, and developer... ...resource utilization (GPU scheduling, autoscaling...Suggested$250k
...opportunities? Join a rapidly scaling AI cloud infrastructure provider building a next-generation GPU platform designed for AI... ...for a Senior / Staff Site Reliability Engineer to support and scale large-... ...frameworks for GPU compute clusters Collaborate with ML, data,...SuggestedPermanent employmentRemote work- ...the next generation of enterprise AI infrastructure. As a Staff AI Infrastructure Engineer, you will design, build, and... ...distributed systems, Kubernetes, GPU infrastructure, high‑performance... ...to deliver secure, scalable, and reliable systems spanning edge deployments...
- Sciforium is looking for a Senior HPC & GPU Infrastructure Engineer to manage the health and performance of our GPU compute cluster. This role involves hands-on Linux systems engineering and maintaining the ML software stack including CUDA and PyTorch. The ideal candidate...Flexible hours
- Beam is seeking a role-focused engineer to own the health and reliability of our GPU compute fleet in a fast-growing AI inference platform. You will build and own metrics pipelines, alerts, and a unified health view across thousands of GPUs in production. You will automate...
$190k - $270k
AI Chopping Block, Inc. is hiring an AI Infrastructure Engineer to ensure smooth operations of user-facing services and production systems in San Francisco. The... ...implementing best practices for availability and reliability. Competitive salary range is $190,000 - $270,000...- ...intersection of platform engineering, site reliability, and applied ML... ...operability of Meshy’s AI model serving stack,... ...with core engineering infrastructure. The team operates a... ...governance. Develop CPU/GPU resource‑management... ...and training share a cluster. Drive unified...Work at officeRemote workFlexible hours
- About Luma AI A new class of intelligence... .... It is an infrastructure challenge at the edge... ...rapidly scaling 10k+ GPU fleets, pushing... ..., throughput, and reliability hard enough that yesterday... ...Infrastructure Engineering team is a systems... ...evolve as cluster size and concurrency...
- cursor in New York, New York, seeks talented engineers to enhance their ML Infrastructure team, focusing on building a robust and scalable coding model... ...researchers, improve training systems, and automate GPU cluster management while collaborating in a flat organizational...
- arcee.ai is seeking a Compute Infra Specialist in San Francisco to help manage and scale the infrastructure for AI workloads. This hands-on role involves coordinating across teams and ensuring efficient GPU resource allocation while enhancing customer experience with our...
- Together AI is seeking an AI Infrastructure Engineer to keep production systems running smoothly and automate operations with mature engineering discipline... ...software skills with pragmatic operations and design reliable, scalable systems for high concurrency. Required: 5+...
$216k - $270k
As a Software Engineer on the Machine Learning Infrastructure team, you will build the “Operating... ...” for our large-scale GPU clusters. You will architect a high... ..., networking, and reliability challenges that emerge at... ...compute into breakthrough AI. You will: Architect and...Full timeFor contractors- Brain Co. in San Francisco is seeking a Backend Engineer to build the technical capabilities for scaling AI products. The role involves designing and maintaining... ...in building production systems, is driven by reliability and usability, and thrives in collaborative environments...
- ...Francisco Compute seeks an experienced systems software engineer to contribute to its UEFI bare-metal infrastructure, VM platform, and other low-level components. You... ...fault-tolerant distributed systems, work with GPU clusters, and implement Kubernetes operators, while...
- OpenAI is seeking an Infrastructure Operations Engineer to operate and improve large-... ...Ethernet fabrics supporting GPU clusters, storage, and management... ...response across a global AI network. You will partner... ...teams to raise reliability, perform RCA, and automate...
$180k - $200k
...States Who We Are Lightning AI is the company behind PyTorch... ...For Lightning AI is seeking a GPU & Compute Infrastructure Engineer to join our Infrastructure Engineering... ...automation, improving reliability, and enabling efficient cluster bring‑up for AI/ML and HPC workloads...Remote workWork from homeFlexible hours- techire ai is seeking an experienced ML Platform engineer to craft scalable infrastructure for training, evaluation, and production deployment of frontier AI models in a highly... .... You will tackle distributed training, GPU orchestration, scheduling, and fault-tolerant...
$177.5k - $248k
...understanding in healthcare. Our AI-powered platform was... ..., technologists, and engineers working together to... .... The Role As an AI Infrastructure Engineer at Abridge,... ...scalable Kubernetes clusters for AI model inference... ...workflows and enhance GPU utilization for ML workloads...Hourly payFull timeFlexible hours- A leading AI research company in San Francisco is seeking a software engineer for its Fleet High Performance Computing team. In this role, you'll ensure the reliability and uptime of the compute fleet, working with automation systems and monitoring tools. Ideal candidates...
- Gravity Engineering Services Pvt Ltd. is seeking a Senior Software Engineer for AI Runtime in San Francisco, California. The role involves... ...architecture of a managed GPU training platform, solving... ...and enhancing performance and reliability. Ideal candidates will have 5...
$224k - $284k
...Atoms builds Physical AI - real‑world robots... ...into something more reliable, more scalable, and more... ...We are roboticists, engineers, operators, and builders... ...and automate our GPU training clusters, including provisioning... .... Design our infrastructure to scale smoothly as...Full timeWork at officeFlexible hours$190k - $270k
AI Chopping Block, Inc. is hiring an AI Infrastructure Engineer in San Francisco, California. This full-time role involves ensuring smooth operation of user-facing services and production systems, alongside building and running infrastructure with Ansible, Terraform, and...Full time- An innovative studio is seeking an AI Infrastructure Engineer to enhance their ML infrastructure for groundbreaking anime games. This role involves designing and implementing cutting-edge inference architectures to support various platforms. As part of a small, agile team...Worldwide
- ...happen to be the world's leading generative AI studio — we're the team behind... ..., is looking for an AI Infrastructure Engineer to join us in building out... ...both on-prem and multi-cloud clusters. But most importantly, you... ...understanding of GPU’s handling large workloads...Work experience placementWork at officeVisa sponsorship
$160k - $230k
An innovative AI technology firm is seeking a passionate Customer Support Engineer to tackle complex technical challenges. You will support customers in building solutions and collaborate with various teams to enhance customer satisfaction. The role requires 3+ years in...- A leading technology firm in San Francisco is seeking a Senior Site Reliability Engineer to maintain and improve cloud infrastructure. The ideal candidate has over 5 years of experience as an SRE or DevOps engineer and strong expertise in Kubernetes. This role focuses on...
$255k
...hyperscale supercomputers reliable and efficient... ...are looking for engineers to operate the... ...generation of compute clusters that power OpenAI'... ...with hands-on infrastructure work on our largest... ...Linux environments, GPU hardware, and large... ...OpenAI is an AI research and deployment...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead AI Infrastructure Engineer: GPU Clusters & Reliability. Be the first to apply!
- lead algorithm engineer San Francisco, CA
- lead product engineer San Francisco, CA
- lead backend developer San Francisco, CA
- lead operating engineer San Francisco, CA
- lead industrial engineer San Francisco, CA
- lead solutions engineer San Francisco, CA
- lead automation engineer San Francisco, CA
- lead engineer San Francisco, CA
- lead infrastructure engineer San Francisco, CA
- lead network engineer San Francisco, CA

