Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

RL Systems Engineer: Large-Scale GPU Training

Applied Compute

Applied Compute is seeking a research scientist to design, implement, and optimize the large-scale training infrastructure powering our reinforcement learning stack in a San Francisco office. You’ll work with researchers to ensure the RL system is fast, reliable, and capable of days-long runs with minimal intervention. You will design and optimize training pipelines across GPUs, build observability tooling, and collaborate on post-training capabilities for production deployments. #J-18808-Ljbffr Applied Compute

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the RL Systems Engineer: Large-Scale GPU Training in San Francisco, CA vacancy
  • $225k

     ...Manufacturing Co is looking for a Software Engineer on the Inference & RL Systems team in San Francisco. The role...  ...high reliability for RL and post-training workflows. The ideal candidate will...  ...fundamentals and experience with large-scale systems. Compensation includes a competitive... 
    Training

    Dormont Manufacturing Co

    San Francisco, CA
    1 day ago
  •  ...Manufacturing Co is looking for a Software Engineer for their Pre-training Systems team in San Francisco. Your primary...  ...that trains long-context models at scale, tackling challenges related to...  ...engineering fundamentals, experience with large models, and a proactive attitude... 
    Training

    Dormont Manufacturing Co

    San Francisco, CA
    2 days ago
  •  ...in San Francisco is looking for a Senior Software Engineer to build scalable infrastructure for large‑scale training and fine-tuning of foundation models. You will design distributed training systems and optimize GPU utilization while collaborating with cross-functional... 
    Training

    Baseten

    San Francisco, CA
    2 days ago
  •  ...shipping excellence. We seek engineers with strong intrinsic...  ...We’re looking for a systems engineer with HPC or...  ...programming experience to help scale AI inference. You’ll...  ...systems to optimize GPU performance at the bleeding...  ...with distributed training/inference frameworks (bonus... 
    Training
    Full time
    Work at office

    Vast

    San Francisco, CA
    6 hours ago
  • $350k

    Thinking Machines Lab is looking for a skilled engineer to operate and automate large GPU clusters in San Francisco, California....  ...Python or Rust, as well as experience with large-scale clusters and container orchestration systems. The role offers an expected salary range of... 
    Training
    Visa sponsorship

    Thinking Machines Lab

    San Francisco, CA
    5 days ago
  •  ...Manufacturing Co is seeking a Staff ML Platform Engineer to join their team. This role focuses on building infrastructure for large-scale training of ML models, requiring deep knowledge of PyTorch and multi-node systems. Join a dynamic environment poised for growth and... 
    Training

    Dormont Manufacturing Co

    San Francisco, CA
    5 days ago
  • Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance. This... 

    Vast.ai Inc.

    San Francisco, CA
    5 days ago
  • $150k - $300k

     ...anyone to create, train, and deploy them...  ...with the full RL post-training stack...  ...at frontier scale, adapting models...  ...Solutions Architect for GPU Infrastructure,...  ...‑ready systems capable of training...  ...our world‑class engineering team while having...  ...of successful large‑scale deployments... 
    Training

    Prime Intellect

    San Francisco, CA
    1 day ago
  • $216k - $270k

    Scale AI in Seattle is looking for a Software Engineer on the Machine Learning Infrastructure team to design and implement a high-performance training platform. The role involves building a multi-tenant orchestration layer for GPU clusters, optimizing job lifecycles, and... 
    Training

    Scale AI

    San Francisco, CA
    5 days ago
  • Gravity Engineering Services Pvt Ltd. is searching for an engineer to design, build, and operate a GPU supercomputing environment. You'll automate large GPU clusters and write software for cluster management, contributing to high-performance computing. The ideal candidate... 
    Training

    Gravity Engineering Services Pvt Ltd.

    San Francisco, CA
    5 days ago
  • $310k

    A leading AI research organization is looking for a Software Engineer for their Platform Systems team in San Francisco. You will design and build systems for large-scale AI training workloads, focusing on reliability and performance. Ideal candidates should have a deep... 
    Training

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...Group is seeking a Research Engineer (Robotics / Foundation Models...  ...You will help design scalable training pipelines, build data infrastructure...  ...across pretraining and RL post-training, own the full...  ...and optimize performance on large GPU clusters. Visa sponsorship is... 
    Training
    Full time
    Visa sponsorship

    Brahma Consulting Group

    San Francisco, CA
    2 days ago
  •  ...and inference. Distributed training, 1000+ K8s GPU clusters, petabyte scale data pipelines, etc. We...  ...datastores, job orchestration systems, and streaming pipelines...  ...and scaling for large-scale GPU job processing...  ...researchers on evolving RL infrastructure Applied ML... 
    Training

    krea.ai

    San Francisco, CA
    5 days ago
  • $175k - $225k

     ...Principal Systems Engineer – GPU Supercluster Bringup We are building AI infrastructure for frontier-scale workloads. Our platform is designed for...  ...and lead the bringup of large-scale GPU clusters (hundreds...  ...methodology for AI training workloads. Deployment... 
    Training
    Flexible hours

    Nscale

    San Francisco, CA
    3 days ago
  • $150k - $300k

    Prime Intellect in San Francisco seeks a Solutions Architect for GPU Infrastructure who will transform client requirements into robust systems capable of training advanced AI models. Responsibilities include designing GPU cluster architectures, deploying orchestration systems... 
    Training

    Prime Intellect

    San Francisco, CA
    1 day ago
  •  ...research firm in San Francisco is seeking talent to build and optimize GPU infrastructure for large-scale model inference and training workloads. The ideal candidate will have hands-on experience with GPU systems and optimization techniques, actively contributing to synthetic... 
    Training

    Reflection

    San Francisco, CA
    5 days ago
  •  ...seeking a Member of Technical Staff for RL Infrastructure in San Francisco. This role...  ...building infrastructure for distributed RL training and inference across thousands of GPUs....  .... Candidates should have strong software engineering experience, particularly in building infrastructure... 
    Training

    Vmax

    San Francisco, CA
    3 days ago
  • $350k

    Mirendil is looking for engineers to build infrastructure for frontier reasoning models at their San Francisco location. This role focuses on large-scale reinforcement learning (RL) model training and requires a solid understanding of engineering principles. The ideal... 
    Training

    Mirendil

    San Francisco, CA
    2 days ago
  • $172k - $209k

     ...who believe in the scale of our ambition and...  ...Production / Sustaining Engineer to strengthen Crusoe’s Hardware Systems Engineering team...  ...bring‑up to large‑scale production—while...  ...across Crusoe Cloud’s GPU‑ and CPU‑based...  ...across: PCIe (link training, topology, performance... 
    Training
    Temporary work

    Crusoe Energy Systems

    San Francisco, CA
    3 days ago
  •  ...Labs is seeking an expert specialized in training systems to enhance performance and stability...  ...multimodal generative models, concentrating on GPU-level optimizations and distributed...  ...skills in PyTorch, experience with large-scale training, and practical judgment for profiling... 
    Training
    Remote job

    Black Forest Labs

    San Francisco, CA
    2 days ago
  •  ...infrastructure. As a Staff AI Infrastructure Engineer, you will design, build, and operate the platforms that enable large‑scale training, serving, evaluation, and deployment...  .... You will work across distributed systems, Kubernetes, GPU infrastructure, high‑performance... 
    Training

    Seekr

    San Francisco, CA
    5 days ago
  •  ...technology firm in San Francisco is seeking systems-oriented candidates to enhance their...  ...distributed systems, particularly within GPU environments. The position offers the chance...  ...work on cutting-edge projects involving large-scale data processing and custom infrastructure... 

    krea.ai

    San Francisco, CA
    5 days ago
  •  ...recruiting a Member of Technical Staff - Research Engineer to own and optimize large-scale generative-model training systems in a hybrid SF/remote setting. You’ll work...  ...performance, memory footprint, and stability across GPU clusters. You’ll implement GPU-level... 
    Training
    Remote work

    BlackForestLabs

    San Francisco, CA
    1 day ago
  • $160k - $225k

    Cacheflow is seeking a Senior Software Engineer for AI Runtime at Databricks, located in San Francisco. You will be instrumental in building and scaling systems for large-scale GPU training, ensuring high throughput and resilience in training across expansive fleets of... 
    Training

    Cacheflow

    San Francisco, CA
    4 days ago
  • Gravity Engineering Services Pvt Ltd. is seeking a Senior Software Engineer for AI Runtime...  ...driving the architecture of a managed GPU training platform, solving complex training challenges...  ...+ years of experience with distributed systems, GPU training, and a strong... 
    Training

    Gravity Engineering Services Pvt Ltd.

    San Francisco, CA
    5 days ago
  •  ...seeking a senior ML infrastructure engineer to design and optimize distributed training systems and performance-critical...  ...hardware utilization across multi‑node GPU clusters. Join a team focused on...  ...performance improvements for large‑scale runs. #J-18808-Ljbffr Genesis AI
    Training

    Genesis AI

    San Francisco, CA
    3 days ago
  •  ...Design, deploy, and maintain large distributed ML training and inference clusters...  ...pipelines to manage petabyte-scale datasets and model training...  ...profile and debug low-level GPU operations to optimize performance...  ...task management systems and scalable model serving... 
    Training

    Kindredventures

    San Francisco, CA
    4 days ago
  • $280k

     ...and steerable AI systems. We want AI to be...  ...committed researchers, engineers, policy experts,...  ...algorithms at our scale often requires...  ...record of solving large-scale systems problems...  ...modelsImplement GPU kernels to adapt our...  ...of education, training, and/or experienceRequired... 
    Training
    Work at office
    Visa sponsorship
    Flexible hours

    Gravity Engineering Services Pvt Ltd.

    San Francisco, CA
    5 days ago
  • About the Role Scale's Physical AI business unit is...  ...for processing, training, and fine‑tuning on data...  ...Physical AI. As an ML Systems Engineer on the Physical AI team...  ...of experience building large‑scale, high‑performance...  ..., including GPU‑level algorithm optimizations... 
    Training

    Gravity Engineering Services Pvt Ltd.

    San Francisco, CA
    5 days ago
  •  ...Research, etc.), where we built large-scale training infrastructure powering...  ...Research Scientist (Efficient ML Systems) at Goaly, you will research...  ...directly shape how modern LLMs and RL systems are trained and...  ...Freedom: Access to abundant GPU cluster resources—don't let your... 
    Training
    Full time
    Work at office

    Goaly AI

    San Francisco, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to RL Systems Engineer: Large-Scale GPU Training. Be the first to apply!