AI Infra Engineer: Scale GPU Clusters & On-Prem AI
Jobleads-US
SpaceX is seeking a Software Engineer for AI Infra (Starshield) to design, deploy, and scale GPU/CPU infrastructure across Top Secret data centers. You will automate on-premise Kubernetes AI clusters and supporting OS environments, while collaborating closely with AI engineers to deliver highly scalable products.
The role covers Site Reliability Engineering, DevOps, and GPU platforms with extensive automation, monitoring, and performance optimization.
#J-18808-Ljbffr Jobleads-US- ...SpaceX is hiring a Sr. Software Engineer for AI Infrastructure (Starshield) in... ...Palo Alto to design, deploy, and scale software and GPU infrastructure supporting... ...missions. You will automate on-prem resources, build scalable AI clusters, and collaborate with AI engineers...Suggested
- ...NVIDIA Corporation seeks a Senior AI/ML Performance and Efficiency Engineer for GPU Clusters to advance AI efficiency across research workloads. You will... ...The role demands 5+ years in managing large-scale compute infra, strong ML tooling knowledge, and hands-on experience...Suggested
- ...SpaceX in Palo Alto is seeking a Software Engineer for AI Infrastructure (Starshield). You will design, operate, and scale GPU-backed infrastructure to support critical national... ...on‑premise compute resources and Kubernetes clusters. Collaborate with AI engineers to build...Suggested
- ...Cloud seeks a Senior Systems Software Engineer with deep expertise in distributed systems... ...own hard technical problems at large scale, optimizing AI infrastructure in production and shaping... ...run on Kubernetes across thousands of GPU nodes. You will collaborate with researchers...Suggested
- ...Sunnyvale, CA seeks a senior software engineer to design high-performance AI frameworks for large-scale distributed systems. You will optimize GPU-driven workloads, integrate LLMs, and lead test-automation infra across Kubernetes-based clusters. You are expected to have...Suggested
- ...Job Overview: LiveX AI is building the next generation of realtime... ...a strong Machine Learning Engineer to help us train, fine‑tune,... ...data pipelines for large‑scale video, audio, and multimodal... ...deployment infrastructure on GPU clusters (PyTorch, DeepSpeed, FSDP, Ray...
- ...Peano AI is seeking a Founding Large-Scale Mid-Training/RL Infrastructure Engineer to design, run, and scale our in-house training stack. You will work with thousands of GPUs and TPU accelerators to optimize pretraining and RL-based methods for large language models,...
- ...Google LLC's YouTube team seeks a Senior Software Engineer, AI/ML, to advance scalable systems and AI-powered features around video and... ...background, and a drive to innovate in information retrieval and large-scale infrastructure. The position emphasizes collaboration, code...
$174k - $252k
...Google is seeking software engineers to design and optimize large-scale systems and AI-powered products. You will work across information retrieval, distributed computing, ML, NLP, and security to push technology forward at scale. Individual pay is determined by factors...$190k - $260k
...artificial intelligence (AI) powered technology... ...GigaFusionNet to large-scale world models – depends... ...throughput. We are looking for engineers who make model training... ...stalling a single GPU, sharding data and models... ...across multi-node GPU clusters, applying data, tensor,...Temporary workWork at officeVisa sponsorshipFlexible hours$152k - $241.5k
...unlimited potential of AI to define the next era of... ...computing. An era in which our GPU acts as the brains of... ...that analyzes large-scale datacenter workloads on GPU-accelerated clusters. You will turn telemetry... ...container, GPU, and systems engineers. When useful, you will...Full timeRemote work- ...Google LLC in Mountain View, CA seeks software engineers to develop next‑gen AI systems and large‑scale ML products. You will translate research into scalable, product‑level improvements across information retrieval, AI quality, and system design. Collaboration and...
$120k - $195k
...of the team.Responsibilities: AI is at the core of how LinkedIn... ...trust platforms. As an AI Software Engineer you will own end-to-end machine... ...run in production at LinkedIn scale. You won't just train models,... ...prove efficacy, and managing the GPU fleets that run them at scale...For contractorsWork at officeImmediate startFlexible hours- ...and scientific discovery to powering AI and the technologies people rely on every... ...a highly motivated PhD AI Systems & GPU Performance Engineering Intern to join our team. In this role... ...TensorFlow and familiarity with large-scale AI model training and inference...Full timeSummer workInternshipSummer internshipWorldwide
$139k - $204k
...The Essential Cloud for AI™. Built for pioneers by... ...innovators to build and scale AI with confidence. Trusted... ...role As part of the Cluster Orchestration team, you... ...efficiently across massive GPU clusters. By building... ...Do As a Senior Software Engineer I (IC3), you will own multiple...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$182k - $242k
...The Essential Cloud for AI™. Built for pioneers by... ...to build and scale AI with confidence. Trusted... ...cloud for high-performance GPU infrastructure across AI... ...inference. Our stack is engineered for speed, scale, and... ...including workload setup, cluster configuration, runbooks...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours- Senior Lead Software Engineer Be an integral part of an agile... ...platforms optimized for AI/ML workloads. Partner... ...by automation at scale. Required qualifications... ...containerization (Docker), including cluster operations and... ...understanding of NVIDIA GPU infrastructure software...For contractors
- Senior Lead Software Engineer Be an integral part of an agile... ...platforms optimized for AI/ML workloads. Partner... ...by automation at scale. Required qualifications... ...containerization (Docker), including cluster operations and... ...understanding of NVIDIA GPU infrastructure software...For contractors
$124.8k - $283.8k
...are looking for an experienced Senior AI Infrastructure Engineer to join our team. This role owns the technical... ..., to spec, and operated reliably at scale.Key Responsibilities• Conduct on-site... .../interns on hardware diagnostics, cluster tooling, and best practices.• Track...Full timeOverseasRelocation package- ...Description Gauss Labs builds Industrial AI for the world's leading... ...tuning, and serving them in large-scale production systems. As a Senior/Staff AI Engineer, you will turn ML research into robust... .../parallel training (multi-GPU/multi-node). Experience deploying...Full timeShift work
$200k - $270k
...Description Samsung SDS America AI Team is researching the next... ...looking for a Senior Physical AI Engineer to join the team developing... ...that support manufacturing at scale across thousands of factory lines... ...foundation models Experience with GPU acceleration and distributed...WorldwideFlexible hours$180k - $275k
...About the role Own verification of our AI compute core - tensor pipelines, MAC arrays... ...accelerator program from first silicon through scale-out. What you'll do Own verification... ...verifying compute-heavy ASIC blocks (GPU, NPU, DSP, or AI accelerator) ~ Strong understanding...H1bVisa sponsorshipWork visa- ...powering the future of physical AI. Founded in 2017 and now... ...looking for a performance engineer who specializes in making large-scale machine learning workloads... ...- it is throughput, cluster goodput, and cost per unit... ...run that wastes 30% of its GPU-hours on stalled data loaders...Full timeFor contractorsFor subcontractorCasual workWork at officeRemote workDay shift
$185k - $275k
...The Essential Cloud for AI™. Built for pioneers by... ...innovators to build and scale AI with confidence. Trusted... ...role As part of the Cluster Orchestration team, you... ...efficiently across massive GPU clusters. By building... ...What You'll Do As a Staff Engineer, you will be a technical...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours- ...Model Optimization & Deployment Engineer, you will focus on bringing... ..., production-ready large-scale models to our on-vehicle stack... ...maximize memory bandwidth on AI accelerators. Write production... ...efficiency optimization for GPU clusters. Experience with end-to-...Temporary workRelocation package
$184k - $287.5k
...into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers... ...We are looking for a dedicated engineer for the Senior Systems Software... ...focusing on GPU Performance at Scale. At NVIDIA, this role is...- ...NVIDIA Corporation in Santa Clara, California, seeks a Senior Developer Technology Engineer, Artificial Intelligence, to research parallel algorithms and optimize AI workloads on modern CPU/GPU architectures. You will work with experts across industry and academia to...
$152k - $241.5k
...We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join our AI Efficiency efforts. As an Engineer, you will have a pivotal... ...+ years of experience designing and operating large scale compute infrastructure Strong understanding of...- ...Corporation in Santa Clara, CA is seeking a Developer Technology Engineer to drive end-to-end agentic AI deployment with RTX & DGX platforms. You will... ...internal teams and with external developers to advance GPU-powered AI workflows. The role requires 5+ years in local...Local area
$137.2k - $206.6k
...US and Dubai, we're now scaling manufacturing and preparing... ...for a Staff Software Engineer with broad software experience and a strong AI focus to help bring production... ...Experience operating GPU compute for AI workloads... ...optimization or on-prem/hybrid infrastructure...Temporary work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Infra Engineer: Scale GPU Clusters & On-Prem AI. Be the first to apply!





