Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

RL Systems Engineer: Large-Scale GPU Training

Applied Compute

Applied Compute is seeking a research scientist to design, implement, and optimize the large-scale training infrastructure powering our reinforcement learning stack in a San Francisco office. You’ll work with researchers to ensure the RL system is fast, reliable, and capable of days-long runs with minimal intervention. You will design and optimize training pipelines across GPUs, build observability tooling, and collaborate on post-training capabilities for production deployments. #J-18808-Ljbffr Applied Compute

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the RL Systems Engineer: Large-Scale GPU Training in San Francisco, CA vacancy
  • $225k

     ...Manufacturing Co is looking for a Software Engineer on the Inference & RL Systems team in San Francisco. The role...  ...high reliability for RL and post-training workflows. The ideal candidate will...  ...fundamentals and experience with large-scale systems. Compensation includes a competitive... 
    Training

    Dormont Manufacturing Co

    San Francisco, CA
    2 days ago
  •  ...Technical Staff to design and operate distributed systems for serving models in production and driving large-scale post-training workflows. You will work where model execution...  ...latency, throughput, and reliability of RL and training loops. You will own the infrastructure... 
    Training

    Magic AI Corp.

    San Francisco, CA
    3 days ago
  • Linuxcareers in San Francisco is building AI research infrastructure. You will design, deploy, and operate large-scale GPU clusters powering training, evaluation, and serving for the research team. The role emphasizes extending orchestration with Kubernetes/Slurm, building... 
    Training

    Linuxcareers

    San Francisco, CA
    4 days ago
  • Magic AI, Inc. is seeking a engineer for the Supercomputing Platform & Infrastructure to design, build, and operate large-scale GPU infrastructure powering model training and inference. You will implement Terraform-driven IaC across cloud and hybrid environments, manage... 
    Training
    Visa sponsorship
    Relocation package

    Magic AI Corp.

    San Francisco, CA
    1 day ago
  •  ...in San Francisco is looking for a Senior Software Engineer to build scalable infrastructure for large‑scale training and fine-tuning of foundation models. You will design distributed training systems and optimize GPU utilization while collaborating with cross-functional... 
    Training

    Baseten

    San Francisco, CA
    3 days ago
  •  ...shipping excellence. We seek engineers with strong intrinsic...  ...We’re looking for a systems engineer with HPC or...  ...programming experience to help scale AI inference. You’ll...  ...systems to optimize GPU performance at the...  ...Familiarity with distributed training/inference frameworks (... 
    Training
    Full time
    Work at office

    Vast.ai Inc.

    San Francisco, CA
    1 day ago
  • $100k - $150k

     ...vertically integrated AI cloud engineered for AI. We own and...  ...energy, data centres, GPU superclusters,...  ...-on with GPU, HPC, or large-scale data centre estates....  ...Engineering, including training content, workshops, and...  ...failures. ~ Linux systems engineering at scale.... 
    Training
    Full time
    Remote work
    Flexible hours

    Nscale

    San Francisco, CA
    2 days ago
  • Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance. This... 

    Vast.ai Inc.

    San Francisco, CA
    1 day ago
  • $150k - $300k

     ...anyone to create, train, and deploy them...  ...with the full RL post-training stack...  ...at frontier scale, adapting models...  ...Solutions Architect for GPU Infrastructure,...  ...‑ready systems capable of training...  ...our world‑class engineering team while having...  ...of successful large‑scale deployments... 
    Training

    Prime Intellect

    San Francisco, CA
    2 days ago
  • $250k

     ...opportunities? Join a rapidly scaling AI cloud...  ...building a next-generation GPU platform designed for AI training, experimentation, and...  ...Staff Site Reliability Engineer to support and scale large-scale HPC and cloud...  ...available infrastructure systems Improve CI/CD... 
    Training
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  • $264.8k - $331k

     ...our society. At Scale, our mission is...  ...of the art post-training algorithms to reach...  ...ML Sys Research Engineer, you'll work on...  ...next-gen Agent RL training platform, support large scale training,...  ...optimize our ML system. Your customer will...  ...of the modern GPU cluster Experience... 
    Training
    Full time

    Scale AI

    San Francisco, CA
    16 hours ago
  •  ...in San Francisco is seeking a Staff ML Systems Engineer to design and prototype algorithms,...  ...style systems, while profiling across GPU, networking, and memory to improve latency...  ...and cost. You will also co-design RL and post-training pipelines, drive performance improvements... 
    Training

    Together

    San Francisco, CA
    2 days ago
  • $310k

    A leading AI research organization is looking for a Software Engineer for their Platform Systems team in San Francisco. You will design and build systems for large-scale AI training workloads, focusing on reliability and performance. Ideal candidates should have a deep... 
    Training

    OpenAI

    San Francisco, CA
    2 days ago
  • $264.8k - $331k

    Machine Learning Systems Research Engineer, Agent Post-training - Enterprise GenAI AI is becoming...  ...of our society. At Scale, our mission is to...  ...for our next-gen Agent RL training platform, support large scale training, and research...  ...of the modern GPU cluster Experience with... 
    Training
    Full time
    Contract work
    For contractors
    For subcontractor
    Work at office

    Scale LLP

    San Francisco, CA
    2 days ago
  • $215k - $260k

     ...who believe in the scale of our ambition and...  ...Production / Sustaining Engineer to strengthen Crusoe’s Hardware Systems Engineering team...  ...bring-up to large-scale production—while...  ...across Crusoe Cloud’s GPU- and CPU-based...  ...across:PCIe (link training, topology, performance... 
    Training

    Crusoe

    San Francisco, CA
    16 hours ago
  • $300 per month

     ...urgency, who believe in the scale of our ambition and...  ...a Staff Hardware Systems Engineer to strengthen Crusoe’s...  ...prototype bring-up to large-scale production while...  ...across Crusoe Cloud’s GPU- and CPU-based infrastructure...  ...studies across training and inference - dense,... 
    Training
    Temporary work

    Crusoe

    San Francisco, CA
    1 day ago
  •  ...research firm in San Francisco is seeking talent to build and optimize GPU infrastructure for large-scale model inference and training workloads. The ideal candidate will have hands-on experience with GPU systems and optimization techniques, actively contributing to synthetic... 
    Training

    Reflection

    San Francisco, CA
    1 day ago
  • $150k - $300k

    Prime Intellect in San Francisco seeks a Solutions Architect for GPU Infrastructure who will transform client requirements into robust systems capable of training advanced AI models. Responsibilities include designing GPU cluster architectures, deploying orchestration systems... 
    Training

    Prime Intellect

    San Francisco, CA
    2 days ago
  • Physical Intelligence seeks a systems-focused ML infrastructure engineer to own scheduling, placement, and cluster management for large-scale model training. You will design multi-tenant schedulers, manage heterogeneous GPU/TPU clusters, and ensure fault-tolerant operations... 
    Training

    Physical Intelligence

    San Francisco, CA
    2 days ago
  •  ...Labs is seeking an expert specialized in training systems to enhance performance and stability...  ...multimodal generative models, concentrating on GPU-level optimizations and distributed...  ...skills in PyTorch, experience with large-scale training, and practical judgment for profiling... 
    Training
    Remote job

    Black Forest Labs

    San Francisco, CA
    3 days ago
  • $350k

    Mirendil is looking for engineers to build infrastructure for frontier reasoning models at their San Francisco location. This role focuses on large-scale reinforcement learning (RL) model training and requires a solid understanding of engineering principles. The ideal... 
    Training

    Mirendil

    San Francisco, CA
    3 days ago
  •  ...recruiting a Member of Technical Staff - Research Engineer to own and optimize large-scale generative-model training systems in a hybrid SF/remote setting. You’ll work...  ...performance, memory footprint, and stability across GPU clusters. You’ll implement GPU-level... 
    Training
    Remote work

    BlackForestLabs

    San Francisco, CA
    2 days ago
  • $160k - $225k

    Cacheflow is seeking a Senior Software Engineer for AI Runtime at Databricks, located in San Francisco. You will be instrumental in building and scaling systems for large-scale GPU training, ensuring high throughput and resilience in training across expansive fleets of... 
    Training

    Cacheflow

    San Francisco, CA
    6 hours ago
  •  ...seeking a senior ML infrastructure engineer to design and optimize distributed training systems and performance-critical...  ...hardware utilization across multi‑node GPU clusters. Join a team focused on...  ...performance improvements for large‑scale runs. #J-18808-Ljbffr GenesisAI
    Training

    GenesisAI

    San Francisco, CA
    4 days ago
  • $165k - $206k

     ...hiring the world’s best engineers, scientists, designers,...  ...:The Enterprise Systems Engineer who leads with...  ...requirements.Identify, pilot, and scale low-friction AI use...  ...UAT script generation, training material creation, and...  ...validation, and working with large data sets.Exceptional... 
    Training

    Juul

    San Francisco, CA
    3 days ago
  • $205k - $270k

     ...interpretable, and steerable AI systems. We want AI to be safe...  ...researchers, engineers, policy experts, and...  ...reliability Assist with training and onboarding of new...  ...supporting rapid growth and scaling financial systems...  ...team on just a few large-scale research efforts... 
    Training
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    1 day ago
  • $220k - $320k

     ...Help us build the systems that train specialized AI models...  ...evaluation, and planet-scale hosting. We are a...  ...ten-person team of engineers who work in-person...  ...have the autonomy, a large compute budget / GPU reservation, and technical...  ...techniques in SFT, RL, and model... 
    Training
    Full time
    Work at office

    Inference

    San Francisco, CA
    16 hours ago
  • $250k

    A Series A Funded start-up in California is seeking a Systems Engineer to design and optimize systems handling complex ML pipelines. The role involves building scalable infrastructure, developing CI/CD pipelines, and ensuring system performance. Key qualifications include... 

    Acceler8 Talent

    San Francisco, CA
    1 day ago
  •  ...the most advanced Large Language Models. Overview...  ...talented MLOps Engineers with deep, hands-...  ...involves AI model training and evaluation...  ...data for frontier AI systems. This is a W-2...  ...experience with PyTorch at scale. Experience...  ...optimizing custom GPU kernels using Triton... 
    Training
    Full time
    Weekday work

    Obsidian

    San Francisco, CA
    16 hours ago
  • Eventual Inc. in San Francisco seeks a Systems Engineer for the Dataloading team to turn multi-petabyte video corpora into tensors...  ...paths, and scalable loaders to keep GPUs fed during large-scale model training. You’ll grow with NVL72, CUDA, and SLURM, shaping data movement... 
    Training
    Work at office

    Eventual Inc.

    San Francisco, CA
    16 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to RL Systems Engineer: Large-Scale GPU Training. Be the first to apply!