Lead AI Systems Engineer — Scale Core ML & GPU Clusters
Transluce
Transluce, a fast-moving research lab in San Francisco, is seeking an exceptional AI systems engineer to lead the design and development of our core ML stack, building scalable systems that can leverage thousands of GPUs and handle trillion-token databases. As an early member of a highly collaborative team, you will move fast, innovate from the ground up, and deliver high-impact tooling with cross-organizational reach, including open-source components that support AI evaluation and public policy #J-18808-Ljbffr Transluce
- Linuxcareers in San Francisco is building AI research infrastructure. You will design, deploy, and operate large-scale GPU clusters powering training, evaluation, and serving for the research team. The role emphasizes extending orchestration with Kubernetes/Slurm, building...Suggested
$250k
...Join a rapidly scaling AI cloud infrastructure... ...a next-generation GPU platform designed... ...Site Reliability Engineer to support and scale... ...closely with platform, ML, and... ...frameworks for GPU compute clusters Collaborate with... ...available infrastructure systems Improve CI/CD...SuggestedFull timeRemote work- ...s most dynamic AI companies, like... ...build the platform engineers turn to to ship... ...operating system for distributed... ...modal workloads scale, the network is... ...foundational engineers to lead our GPU Networking... ...bleeding-edge clusters (H100/H200, B20... ...a variety of ML startups,...SuggestedFull timeFlexible hours
- Salesforce is seeking an accessibility engineer to accelerate the development of accessible AI powered products and services across the enterprise... .... As a subject matter expert with AI/ML experience, you design prompt systems, build evaluation frameworks, and enable scalable...Suggested
- ...progress of AI applications... ...scientist can scale an ML application from... ...to the cluster without needing... ...distributed systems expert.... ...looking for engineers with systems... ...About the Ray Core Team The Ray... ...you will: Leading cross-team projects... ...Knowledge of GPU programming is...SuggestedFull timeWork experience placement
- ...software development with AI-powered formal... ...Join our team as an AI Engineer and help us push the boundaries... ...refine the data and ML pipelines for scaled distributed training... ...with Kubernetes clusters and distributed compute... ...Multi-node and multi-GPU training Mathematical...Full timeContract work
$240k - $280k
...a Software Engineer to build the systems that treat infrastructure... ...inference clusters without a... ...stand up, scale, or tear... ...functioning AI cluster for... ...bring-up to GPU driver/CUDA... ...the inference/ML platform... ...Requirements Core requirements... ...to leading open-source...Full time$220k
...Perplexity is looking for an engineer to join their team in San Francisco. You will work on... ...engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving... ...experience in software engineering with a focus on ML inference, familiarity with deep learning...$220k - $300k
...so companies can scale without losing... ...raised $204M from leading venture capital... ...customer operations. AI is reshaping... ...effort is our Core AI Platform team... ...applied AI and engineering talent whose work... ...pipelines, evaluation systems) that all... ...technical fluency in AI/ML systems,...Work at officeImmediate startRemote workWork from homeMonday to Friday- iframe.ai is hiring a Customer Cluster Engineer to own three to five reserved-capacity accounts, managing their training and inference performance end... ...kernel tuning. You have 5+ years in distributed- or ML-systems engineering, strong PyTorch/FSDP/Megatron-LM debugging...Remote job
- ...the web by building AI agents that can reliably... ...Responsibilities: Scale infra for post-training... ...closely with product engineers to translate cutting‑... ...: Experience with ML infrastructure (GPU clusters) and supporting networking... ...latency) Low level systems experience (Triton,...Work at officeRelocationVisa sponsorship
$250k
A Series A Funded start-up in California is seeking a Systems Engineer to design and optimize systems handling complex ML pipelines. The role involves building scalable infrastructure, developing CI/CD pipelines, and ensuring system performance. Key qualifications include...$207k - $290k
...About JazzX AI: Vision:... ...enterprises don't scale expertise—... ...bringing AI systems to market that... ...experienced AI Engineer with deep... ...You will lead the design, development... ...product, core platform engineering... ...in AI/ML engineering,... ...(Kubernetes, GPU/TPU clusters, and cloud ML...WorldwideFlexible hours- ...to incubate and scale revolutionary AI‑powered enterprise... ...and intelligent systems. The result is a... ...pacesetters who lead their industries... ...an experienced AI Engineer with deep expertise... ...experience in AI/ML engineering, including... ...(Kubernetes, GPU/TPU clusters, and cloud ML platforms...Flexible hours
$106.9k - $176.5k
...world. The opportunity We are seeking AI Systems Engineers to own the security and trust fabric of... ...and cryptographic lifecycle at production scale across multiple environments. Strong understanding... ...OPA), and API gateways. Exposure to AI/ML workloads and the specific trust...Work experience placementSummer holidayRemote workFlexible hours$350k
...tech stack for understanding and debugging AI systems. We build world-class, AI-backed... ...are looking for an exceptional AI systems engineer to lead the design and development of our core ML stack , building systems that can scale to thousands of GPUs and performantly query...Visa sponsorshipFlexible hours- ...world's most dynamic AI companies, like... ...the platform engineers turn to to ship AI... ...building its own GPU infrastructure for large-scale inference. As we... ...high-density NVIDIA systems, the hardest... ...We are hiring a Lead Software Engineer... ...to a variety of ML startups, offering...Full timeFlexible hours
- ...software development with AI-powered formal... ...our team as an AI Engineer and help us push... ...LLMs) pipelines for scaled distributed training... ...optimizing machine learning systems, including general... ...experience in ML Infra, DataOps,... ...Proficiency with Kubernetes clusters and distributed...Full timeContract work
$200k - $350k
...operations across multiple systems and platforms.... ...seeking a Senior AI Engineer to develop, deploy... ...AI, and modern ML frameworks. Build... ...infrastructure. GPU optimization. RAG... ...systems. Startup or scale-up experience.... ...technology that sits at the core of critical...Remote jobFull timeImmediate start- AI Systems Engineer - Codex Core Agents Location San Francisco Employment Type Full time... ...virtualization, cloud platforms, or ML systems. Enjoy working... ...strong ownership, and can lead scoped or multi-team AI... ...runtimes, inference optimization, GPU systems, benchmarking,...Full timeWork at officeLocal areaRelocation packageFlexible hours
$230k - $385k
AI Systems Engineer - Codex Core Agents About The Team The Codex Core Agents team builds... ...across low-level systems and ML workflows, able to debug Codex... ..., inference/runtime stack, GPU fleet, and product surface.... ...show strong ownership, and can lead scoped or multi-team AI...- Innovaccer is seeking a Principal AI Engineer to build production-grade AI systems, including LLM-based solutions, AI agents, and RAG-driven workflows. You... ...design, develop, and deploy AI-powered applications at scale. The ideal candidate will own the full AI development...
$200k - $350k
...operations across multiple systems and platforms.... ...seeking a Staff AI Engineer to lead the design and... ...production AI systems at scale. What You'll Do... ...of engineering or ML experience. ~... ...Kubernetes and GPU infrastructure.... ...that sits at the core of critical business...Remote jobFull timeImmediate start- ...About the Team The Codex Core Agent team builds the kernel of... ...That means working across the systems that make Codex actually function... ...We’re looking for applied AI engineers to help bring Codex agents from... ...Python and comfortable with modern ML tooling. Have worked on...Full time
- ...Limited is seeking an ambitious infrastructure engineer to help build and operate automated systems that bring GPU clusters from bare machines to customer-ready... ...tenant lifecycle, run Kubernetes and Postgres at scale, and contribute to Kubernetes operators that power...Remote job
- ...client is a well-funded AI startup building production-grade ML infrastructure used by... ...for a Senior AI/ML Engineer to own model training pipelines, evaluation systems, and inference serving at scale. Full-time, on-site in... ...distributed training, GPU optimization, or inference...Full time
- ...da Vinci surgical system and Ion—have transformed... ....We’re a team of engineers, clinicians, and... ...a Senior Systems GPU Engineer - AI & Robotics, you... ...research, SW/ HW/ ML engineering, regulatory... ...& Strategy: Lead complex, end-to-end... ...while serving as a core contributor to team...Local areaWorldwideFlexible hours
- ...The AI Infrastructure team at Zensors builds the engine that powers our visual sensing... ...Engineer in ML Runtime & Optimization... ...: Optimizing Core ML Pipelines:... ..., and operating systems to facilitate... ...understanding of GPU hardware performance... ...or cloud-scale inference serving...Full time
$350k
...tools to make AI work for their... ...are scientists, engineers, and builders who... ...and systems engineers to help... ...architecting and scaling the core infrastructure... ...infrastructure for the clusters to reliably and... ...clusters with GPU workloads, or building... ...with GPU/ML workflows or...Full timeLocal areaImmediate startVisa sponsorshipWork visaRelocation packageFlexible hours- ...to be the world's leading generative AI studio — we're... ...AI Infrastructure Engineer to join us in building... ...out end-to-end ML infrastructure to... ...distributed systems. You have familiarity... ...and multi-cloud clusters. But most... ...understanding of GPU’s handling large...Work experience placementWork at officeVisa sponsorship
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead AI Systems Engineer — Scale Core ML & GPU Clusters. Be the first to apply!
- lead algorithm engineer San Francisco, CA
- lead web developer San Francisco, CA
- lead network engineer San Francisco, CA
- lead infrastructure engineer San Francisco, CA
- lead system engineer San Francisco, CA
- lead operating engineer San Francisco, CA
- lead app. developer San Francisco, CA
- lead engineer San Francisco, CA
- ai research engineer San Francisco, CA
- senior ai engineer San Francisco, CA



