RL Systems Engineer: Large-Scale GPU Training
Applied Compute
Applied Compute is seeking a research scientist to design, implement, and optimize the large-scale training infrastructure powering our reinforcement learning stack in a San Francisco office. You’ll work with researchers to ensure the RL system is fast, reliable, and capable of days-long runs with minimal intervention. You will design and optimize training pipelines across GPUs, build observability tooling, and collaborate on post-training capabilities for production deployments. #J-18808-Ljbffr Applied Compute
$225k
...Manufacturing Co is looking for a Software Engineer on the Inference & RL Systems team in San Francisco. The role... ...high reliability for RL and post-training workflows. The ideal candidate will... ...fundamentals and experience with large-scale systems. Compensation includes a competitive...Training- ...Technical Staff to design and operate distributed systems for serving models in production and driving large-scale post-training workflows. You will work where model execution... ...latency, throughput, and reliability of RL and training loops. You will own the infrastructure...Training
- Linuxcareers in San Francisco is building AI research infrastructure. You will design, deploy, and operate large-scale GPU clusters powering training, evaluation, and serving for the research team. The role emphasizes extending orchestration with Kubernetes/Slurm, building...Training
- Magic AI, Inc. is seeking a engineer for the Supercomputing Platform & Infrastructure to design, build, and operate large-scale GPU infrastructure powering model training and inference. You will implement Terraform-driven IaC across cloud and hybrid environments, manage...TrainingVisa sponsorshipRelocation package
- ...in San Francisco is looking for a Senior Software Engineer to build scalable infrastructure for large‑scale training and fine-tuning of foundation models. You will design distributed training systems and optimize GPU utilization while collaborating with cross-functional...Training
- ...shipping excellence. We seek engineers with strong intrinsic... ...We’re looking for a systems engineer with HPC or... ...programming experience to help scale AI inference. You’ll... ...systems to optimize GPU performance at the... ...Familiarity with distributed training/inference frameworks (...TrainingFull timeWork at office
$100k - $150k
...vertically integrated AI cloud engineered for AI. We own and... ...energy, data centres, GPU superclusters,... ...-on with GPU, HPC, or large-scale data centre estates.... ...Engineering, including training content, workshops, and... ...failures. ~ Linux systems engineering at scale....TrainingFull timeRemote workFlexible hours- Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance. This...
$150k - $300k
...anyone to create, train, and deploy them... ...with the full RL post-training stack... ...at frontier scale, adapting models... ...Solutions Architect for GPU Infrastructure,... ...‑ready systems capable of training... ...our world‑class engineering team while having... ...of successful large‑scale deployments...Training$250k
...opportunities? Join a rapidly scaling AI cloud... ...building a next-generation GPU platform designed for AI training, experimentation, and... ...Staff Site Reliability Engineer to support and scale large-scale HPC and cloud... ...available infrastructure systems Improve CI/CD...TrainingFull timeRemote work$264.8k - $331k
...our society. At Scale, our mission is... ...of the art post-training algorithms to reach... ...ML Sys Research Engineer, you'll work on... ...next-gen Agent RL training platform, support large scale training,... ...optimize our ML system. Your customer will... ...of the modern GPU cluster Experience...TrainingFull time- ...in San Francisco is seeking a Staff ML Systems Engineer to design and prototype algorithms,... ...style systems, while profiling across GPU, networking, and memory to improve latency... ...and cost. You will also co-design RL and post-training pipelines, drive performance improvements...Training
$310k
A leading AI research organization is looking for a Software Engineer for their Platform Systems team in San Francisco. You will design and build systems for large-scale AI training workloads, focusing on reliability and performance. Ideal candidates should have a deep...Training$264.8k - $331k
Machine Learning Systems Research Engineer, Agent Post-training - Enterprise GenAI AI is becoming... ...of our society. At Scale, our mission is to... ...for our next-gen Agent RL training platform, support large scale training, and research... ...of the modern GPU cluster Experience with...TrainingFull timeContract workFor contractorsFor subcontractorWork at office$215k - $260k
...who believe in the scale of our ambition and... ...Production / Sustaining Engineer to strengthen Crusoe’s Hardware Systems Engineering team... ...bring-up to large-scale production—while... ...across Crusoe Cloud’s GPU- and CPU-based... ...across:PCIe (link training, topology, performance...Training$300 per month
...urgency, who believe in the scale of our ambition and... ...a Staff Hardware Systems Engineer to strengthen Crusoe’s... ...prototype bring-up to large-scale production while... ...across Crusoe Cloud’s GPU- and CPU-based infrastructure... ...studies across training and inference - dense,...TrainingTemporary work- ...research firm in San Francisco is seeking talent to build and optimize GPU infrastructure for large-scale model inference and training workloads. The ideal candidate will have hands-on experience with GPU systems and optimization techniques, actively contributing to synthetic...Training
$150k - $300k
Prime Intellect in San Francisco seeks a Solutions Architect for GPU Infrastructure who will transform client requirements into robust systems capable of training advanced AI models. Responsibilities include designing GPU cluster architectures, deploying orchestration systems...Training- Physical Intelligence seeks a systems-focused ML infrastructure engineer to own scheduling, placement, and cluster management for large-scale model training. You will design multi-tenant schedulers, manage heterogeneous GPU/TPU clusters, and ensure fault-tolerant operations...Training
- ...Labs is seeking an expert specialized in training systems to enhance performance and stability... ...multimodal generative models, concentrating on GPU-level optimizations and distributed... ...skills in PyTorch, experience with large-scale training, and practical judgment for profiling...TrainingRemote job
$350k
Mirendil is looking for engineers to build infrastructure for frontier reasoning models at their San Francisco location. This role focuses on large-scale reinforcement learning (RL) model training and requires a solid understanding of engineering principles. The ideal...Training- ...recruiting a Member of Technical Staff - Research Engineer to own and optimize large-scale generative-model training systems in a hybrid SF/remote setting. You’ll work... ...performance, memory footprint, and stability across GPU clusters. You’ll implement GPU-level...TrainingRemote work
$160k - $225k
Cacheflow is seeking a Senior Software Engineer for AI Runtime at Databricks, located in San Francisco. You will be instrumental in building and scaling systems for large-scale GPU training, ensuring high throughput and resilience in training across expansive fleets of...Training- ...seeking a senior ML infrastructure engineer to design and optimize distributed training systems and performance-critical... ...hardware utilization across multi‑node GPU clusters. Join a team focused on... ...performance improvements for large‑scale runs. #J-18808-Ljbffr GenesisAITraining
$165k - $206k
...hiring the world’s best engineers, scientists, designers,... ...:The Enterprise Systems Engineer who leads with... ...requirements.Identify, pilot, and scale low-friction AI use... ...UAT script generation, training material creation, and... ...validation, and working with large data sets.Exceptional...Training$205k - $270k
...interpretable, and steerable AI systems. We want AI to be safe... ...researchers, engineers, policy experts, and... ...reliability Assist with training and onboarding of new... ...supporting rapid growth and scaling financial systems... ...team on just a few large-scale research efforts...TrainingWork at officeVisa sponsorshipFlexible hours$220k - $320k
...Help us build the systems that train specialized AI models... ...evaluation, and planet-scale hosting. We are a... ...ten-person team of engineers who work in-person... ...have the autonomy, a large compute budget / GPU reservation, and technical... ...techniques in SFT, RL, and model...TrainingFull timeWork at office$250k
A Series A Funded start-up in California is seeking a Systems Engineer to design and optimize systems handling complex ML pipelines. The role involves building scalable infrastructure, developing CI/CD pipelines, and ensuring system performance. Key qualifications include...- ...the most advanced Large Language Models. Overview... ...talented MLOps Engineers with deep, hands-... ...involves AI model training and evaluation... ...data for frontier AI systems. This is a W-2... ...experience with PyTorch at scale. Experience... ...optimizing custom GPU kernels using Triton...TrainingFull timeWeekday work
- Eventual Inc. in San Francisco seeks a Systems Engineer for the Dataloading team to turn multi-petabyte video corpora into tensors... ...paths, and scalable loaders to keep GPUs fed during large-scale model training. You’ll grow with NVL72, CUDA, and SLURM, shaping data movement...TrainingWork at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to RL Systems Engineer: Large-Scale GPU Training. Be the first to apply!
- lead system engineer San Francisco, CA
- distributed systems engineer San Francisco, CA
- operations support system engineer San Francisco, CA
- computer systems engineer San Francisco, CA
- system performance engineer San Francisco, CA
- unix linux systems engineer San Francisco, CA
- microsoft systems engineer San Francisco, CA
- mission system engineer San Francisco, CA
- system engineer remote San Francisco, CA
- application system engineer San Francisco, CA



