Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Systems Engineer: Distributed GPU Training & RL Pipelines

Nebius B.V.

Nebius B.V. is building an AI training and model post-training capability focused on frontier models. This role owns the infrastructure for large-scale training, RL experiments, and production-grade workflows. You will work at the intersection of distributed systems, GPU performance, and ML framework integration. The role requires strong Python and PyTorch engineering skills, hands-on experience with distributed model training, and the ability to optimize throughput and memory across multi-GPU #J-18808-Ljbffr Nebius B.V.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the ML Systems Engineer: Distributed GPU Training & RL Pipelines in Palo Alto, CA vacancy
  • Waymo is hiring engineers to build ML systems that train and improve pre-trained models for the Waymo Driver, tackling reinforcement...  ...and modeling engineers to deploy RL techniques across CPU/GPU/TPU, design scalable training pipelines, and evaluate results. This role offers... 
    Pipeline
    Training

    Neura Market

    Mountain View, CA
    2 days ago
  •  ...is building an AI training and model post-training...  ...scale training and RL experiments...  ...the intersection of distributed systems, GPU performance, model...  ...training frameworks, RL pipelines, and production engineering. Your responsibilities...  ..., large-scale ML systems, or GPU cluster... 
    Pipeline
    Training

    Nebius B.V.

    Palo Alto, CA
    4 days ago
  •  ...for the multimodal video, training, and RL pipelines that power frontier...  ...world model teams. 3+ yrs distributed systems / ML infra. About Orbifold AI...  ...hiring a Machine Learning Engineer to scale and optimize the...  ...deployments to maximize GPU/CPU utilization and reduce... 
    Pipeline
    Training

    Orbifold AI

    Palo Alto, CA
    4 days ago
  • $300k - $400k

     ...will own the systems layer that makes...  ...model training and inference...  ...coupled to the RL feedback loop...  ...communication and GPU kernels to...  ...the training pipeline Engage directly...  ..., or distributed training collective...  ...distributed ML systems to identify...  ...scientists, engineers, and problem-... 
    Pipeline
    Training
    Visa sponsorship
    Flexible hours
    Shift work

    Periodic Labs

    Menlo Park, CA
    4 days ago
  •  ...alongside evolving ML workloads, and...  ...forefront of training AI models for...  ...and production engineering for chip designs...  ...robust, efficient pipelines for model fine‑...  ...the end‑to‑end RL workflow—from...  ...coding tasks. Systems Engineering:...  ...with large‑scale distributed systems, high‑performance... 
    Pipeline
    Training

    Architect, Inc.

    Palo Alto, CA
    2 days ago
  • Intel is seeking a Machine Learning Engineer / Data Scientist to join our team...  ...for agent harness, develop RL environments and reward models, and run training to improve capabilities for agentic...  ...engineering and agent memory, and optimize pipelines from data ingestion to deployment... 
    Pipeline
    Training

    PVH (Tommy Hilfiger/Calvin Klein)

    Santa Clara, CA
    2 days ago
  •  ...the Technical Staff, you will train AI models for chip design, verification...  ...like TSMC. You will co-design RL environments, implement post-...  ..., and build end-to-end ML pipelines that translate research into reliable, production-ready systems. #J-18808-Ljbffr Architect... 
    Pipeline
    Training

    Architect Labs

    Palo Alto, CA
    2 days ago
  •  ...scale reinforcement learning systems that power large-scale models. You will run post-training RL pipelines across thousands of GPUs, crafting...  ...while integrating inference engines like vLLM and SGLang. You’ll...  ...in a fast, scalable, multi-GPU setting. #J-18808-Ljbffr... 
    Pipeline
    Training

    Luma

    Redwood City, CA
    1 day ago
  •  ...infrastructure company in California seeks a Member of Technical Staff — Training to design and optimize large-scale distributed training systems for frontier AI models. Candidates should have 5+ years of experience in ML systems and be proficient in Python along with another... 
    Training

    RadixArk

    Palo Alto, CA
    3 days ago
  • $174.9k - $261.3k

     ...The Data Labeling Engineering team designs, builds...  ...tools and pipelines that power autonomous...  ...engineering, and AI/ML, defining the strategies...  ...create reliable training data at scale. Our...  ...direct impact on systems that unblock the...  ...experience building robust distributed platforms and... 
    Pipeline
    Training
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    2 days ago
  • $250k - $350k

     ...Staff level Inference Engineers to accelerate the...  ...acceleration, GPU parallelism, advanced...  ...our creative AI systems deliver industry-leading...  ...inference pipelines, implement state-of...  ...computing kernels and distributed workloads using...  ...production.Improve Training Efficiency: (Bonus... 
    Pipeline
    Training
    Work at office
    3 days per week

    Pika

    Palo Alto, CA
    3 days ago
  • $150k - $230k

     ...recommendation systems, and adtech....  ...Machine Learning Engineer to drive the post-training of our large...  ...reinforcement learning (RL). You will own...  ...the full pipeline: continuous...  ...mid-to-large GPU clusters, applying distributed-training...  ...engineering for ML. You can independently... 
    Pipeline
    Training
    Full time
    Local area
    Work from home

    News Break

    Mountain View, CA
    2 days ago
  • Rhoda AI is hiring a Senior/Staff-level Research Engineer to ensure our robot-learning pipeline is reliable from data collection through model training, inference, and real-robot evaluation. You will build validation systems, observability, and robust operating practices... 
    Pipeline
    Training

    Socket

    Mountain View, CA
    3 days ago
  •  ...generation of generalist robotic systems in Mountain View, CA. We are...  ...or staff-level Research Engineer or ML Systems Engineer to make the robot-learning pipeline reliable and measurable from end...  ...through dataset generation, model training, inference, and real-robot... 
    Pipeline
    Training

    RHODA

    Mountain View, CA
    4 days ago
  • $90.1k - $191.8k

     ...The Data Labeling Engineering team designs, builds,...  ...machine labeling tools and pipelines that power autonomous...  ...engineering , and ML , defining labeling strategies...  ...that create reliable training data at scale. Our...  ...and work directly on systems that unblock the next... 
    Pipeline
    Training
    Full time
    Work experience placement
    Local area
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Mountain View, CA
    4 days ago
  •  ...for a Senior MLOps engineer to work closely with...  ...to build and deploy ML models on a modern MLOps...  ...batch model serving systems, hyper-parameter...  ...and maintain robust pipelines for distributed training on GPU-enabled clusters to...  ...Experience with Ray, vLLM, RL libraries such as... 
    Pipeline
    Training

    JP Morgan Chase

    Palo Alto, CA
    1 day ago
  • $180k

     ...is to create AI systems that can...  ...and focused on engineering excellence. This...  ...ROLE: As an ML Infrastructure...  ...building, and scaling GPU compute infrastructure, training frameworks, and...  ...data pipelines and integrating...  ...environments, distributed systems, GPU infrastructure... 
    Pipeline
    Training
    Temporary work
    Work experience placement

    SpaceXAI

    Palo Alto, CA
    2 days ago
  • $204k - $259k

     ...The Waymo ML Frameworks & Efficiency...  ...including pre-training and post-...  ...are looking for engineers with ML system expertise to...  ...reinforcement learning (RL), building...  ...(i.e. CPU/GPU/TPU)....  ...end RL training pipeline for efficient...  ...and low-latency distributed reply buffers... 
    Pipeline
    Training
    Full time
    Remote work

    Waymo

    Mountain View, CA
    1 day ago
  • $250k - $320k

    Staff Infrastructure Engineer We are partnered...  ...platform, GPU‑based training clusters, and data processing pipelines that drive real‑time...  ...role in scaling systems for both research...  ...experience in Software / ML Infrastructure...  ...experience with distributed systems and GPU... 
    Pipeline
    Training
    Full time
    Immediate start

    Strativ Group

    Menlo Park, CA
    1 day ago
  • $224k - $356.5k

     ..., we are building agentic systems that can reason about, build...  ...the meta-layer of modern ML: the agents, tooling, pipelines, and feedback loops that...  ...are looking for exceptional engineers who are passionate about...  ...curation, evaluation, debugging, training orchestration, and... 
    Pipeline
    Training
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...with all of their business systems through natural language...  ...with Moveworks’ Reasoning Engine and natural language capabilities...  ...help build cutting edge ML infrastructure for...  ...responsibilities including distributed training and inference pipeline for large language models(... 
    Pipeline
    Training
    Work at office
    Remote work
    Flexible hours

    Moveworks

    Mountain View, CA
    3 days ago
  •  ...looking for an Applied ML Engineer to design, evaluate,...  ...recommendation and ranking systems that power how content...  ...models, bandits, and RL-based approaches....  ...ranking, and re-ranking pipelines using offline and online...  ...) and experience training, evaluating, and deploying... 
    Pipeline
    Training
    Full time

    Darwin

    Palo Alto, CA
    1 day ago
  •  ...gaming and embedded systems. Grounded in a...  ...are hiring AI / ML Platform Engineers to build the...  ...agent execution, distributed training and inference, experiment...  ...management, and GPU cluster...  ...research workflows for RL systems,...  ...and evaluation pipelines.Partner with Applied... 
    Pipeline
    Training

    AMD

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

     ...our team of innovative engineers who are building an...  ...insights and automation for GPU fleets. As an ML Engineer on this team...  ...real-time streaming pipelines, detecting anomalies...  ...they impact AI training and inference.The core...  ...algorithms directly in systems languages for latency... 
    Pipeline
    Training
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...world-class conversational AI system. Our mission is to deliver...  ...are seeking to hire an ML Engineer to join our growing team! This...  ...and implementing ML training pipelines as we improve and scale our...  ...Demonstrated success working in large distributed teams to solve novel... 
    Pipeline
    Training
    Full time
    Temporary work
    Part time

    IBM

    Mountain View, CA
    3 days ago
  • $124k - $250k

     ...member of our software engineering infra team, you'll solve...  ...availability, globally distributed ecosystem platform of...  ...development of novel new systems that integrate into...  ...be exposed to the whole pipeline of model delivery, including training, serving, and optimizations... 
    Pipeline
    Training

    AppLovin

    Palo Alto, CA
    2 days ago
  • $153.2k - $234.1k

     ...hardware and battery systems to intuitive...  ...scenarios. As a Senior ML Infra Engineer, you will work on...  ...generation, training, evaluation and iteration...  ...model training pipelines that are...  ...working on large-scale distributed systems,...  ...training across large GPU/CPU clusters or specialized... 
    Pipeline
    Training
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    6 hours ago
  •  ...are looking for a senior ML infrastructure engineer to build and evolve the systems that support model training, deployment, and...  ...computing ML training pipelines, job orchestration, compute...  ...safety. Experience with distributed training, GPU infrastructure, workload... 
    Pipeline
    Training

    Maxinsights Corporation

    Santa Clara, CA
    23 days ago
  • xAI is seeking an engineer for the RL infrastructure team to help with low precision RL training and inference. You will design and optimize the inference stack for...  ...building and optimizing large-scale distributed systems and proficiency in Python, C++, or Rust, along... 
    Training

    xAI

    Palo Alto, CA
    21 hours ago
  •  ...seeking a Machine Learning Engineer to scale and optimize the ML infrastructure behind our multimodal data pipelines. You will work with...  ..., and sensor data into training, evaluation, and RL signals for frontier robotics...  ..., fault-tolerant systems and delivering end-to-end... 
    Pipeline
    Training

    Orbifold AI

    Palo Alto, CA
    21 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Systems Engineer: Distributed GPU Training & RL Pipelines. Be the first to apply!