Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Machine Learning Engineer, Model Training and Reinforcement Learning

$195.2k - $262.2k
Full-time

Nebius

About Nebius:

Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.

Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.

Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.

The role

Nebius Token Factory is building an AI training and model post-training capability for frontier model improvement. This role owns the infrastructure that makes large-scale training and RL experiments possible, reliable, reproducible, and efficient. The work sits at the intersection of distributed systems, GPU performance, model training frameworks, RL pipelines, and production engineering.

A Senior Machine Learning Engineer owns substantial ML work end to end. They can translate an ambiguous capability goal into concrete experiments, implement and debug training and RL recipes, build the supporting data and systems, and deliver measurable improvements in model quality, experiment throughput, and reliability. They are deeply hands-on and can independently debug both model-behavior failures and distributed training failures.

Your responsibilities:

  • Design and run model-training and post-training experiments, including SFT, continued pretraining, preference optimization (DPO/IPO/KTO), and RL methods such as RLHF/RLAIF, PPO, and GRPO.

  • Build reward functions, judge models, verifiers, task environments, and evaluation sets for reasoning, coding, tool use, and agentic workflows.

  • Create synthetic data and data pipelines, including teacher-student generation, self-play, rejection sampling, filtering, and quality scoring.

  • Analyze model-behavior failures and turn them into targeted data, reward, or algorithm improvements.

  • Build and maintain distributed training and RL infrastructure using frameworks such as Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, verl, slime, AReaL, or OpenRLHF.

  • Implement and debug parallelism strategies (tensor, pipeline, sequence/context, expert, and data parallelism) and build reliable rollout, reward-serving, checkpointing, and experiment-orchestration components.

  • Profile and improve GPU utilization, memory usage, communication efficiency, training throughput, and inference/serving performance.

  • Design rigorous evaluations and ablations for capability, instruction following, reasoning, tool use, safety, and regression risk.

  • Write clear experiment plans, design docs, benchmark reports, and runbooks, and partner across research and platform teams.

Must-haves:

  • Strong Python and PyTorch engineering skills, with the ability to move quickly from idea to experiment to working system.

  • Hands-on experience across at least two of: model training, post-training/RL, applied modeling, data pipelines, or large-scale ML systems.

  • Ability to design rigorous experiments with baselines, ablations, metrics, and failure analysis.

  • Practical understanding of modern LLM behavior, instruction tuning, preference optimization, and evaluation challenges.

  • Practical understanding of transformer training bottlenecks, memory pressure, communication overhead, and checkpointing.

  • Ability to reason quantitatively about model quality, throughput, utilization, reliability, cost, and research velocity.

  • Strong communication skills and ability to collaborate with researchers, engineers, and leadership.

Nice-to-haves:

  • Experience with LLM post-training, RL, agents, reward modeling, synthetic data, or model evaluation.

  • Experience with RL frameworks or pipelines such as verl, slime, AReaL, OpenRLHF, TRL, or custom PPO/GRPO/RLHF systems.

  • Experience with Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, Slurm, or Kubernetes on large GPU clusters.

  • Familiarity with NCCL, CUDA, Triton, Nsight, InfiniBand/RDMA, and H100/H200/B200 clusters, or with model serving and inference optimization.

  • Publications, open-source contributions, or production impact in LLM post-training, RL, reasoning, coding models, synthetic data, distributed training, or evaluation.

  • Experience designing agent environments, tool-use tasks, or verifier-based rewards.

Key employee benefits in the US:

  • Health insurance: 100% company-paid medical, dental, and vision coverage for employees and families.

  • 401(k) plan: Up to 4% company match with immediate vesting.

  • Parental leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers.

  • Remote work reimbursement: Up to $85/month for mobile and internet.

  • Disability & life insurance : Company-paid short-term, long-term and life insurance coverage.

Pay Transparency

We offer competitive compensation and benefits packages. Actual compensation will be determined based on job-related factors, including experience, skills, qualifications, the level at which the candidate is hired, and geographic location, consistent with applicable law.

Base Compensation Range

$195,200—$262,200 USD

Benefits & Perks:

  • Competitive compensation
  • Career growth and learning opportunities
  • Flexibility and ownership
  • Collaborative and innovative culture
  • Opportunity to work on impactful AI projects
  • International environment and talented teams

What's it like to work at Nebius:

Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI

Equal Opportunity Statement:

Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law.

Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire.

If you need accommodations during the application process, please let us know.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior Machine Learning Engineer, Model Training and Reinforcement Learning in Palo Alto, CA vacancy
  • $174.72k - $295.68k

     ...cutting-edge R&D in AI, machine learning, and smart...  ...time Machine Learning Engineer / Research Scientist to drive the modeling and algorithmic development...  ...infrastructure experts to design, train, and deploy large-...  ..., imitation and reinforcement learning) to improve model... 
    Senior
    Training
    Full time

    XPENG Motors

    Santa Clara, CA
    4 days ago
  • $174.72k - $295.68k

     ...cutting-edge R&D in AI, machine learning, and smart...  ...seeking Machine Learning Engineers with strong expertise in generative modeling and large-scale deep...  ...learned simulator for training and evaluating driving...  ...including model-based reinforcement learning and Vision-Language... 
    Senior
    Training
    Full time

    XPENG Motors

    Santa Clara, CA
    4 days ago
  • $262k - $361k

     ...AI, software, engineering, and product to...  ...global scale. Learn more about our...  ...’s multi-year machine learning...  ...engines, scaling models across multimodal...  ...sensing data, reinforcement learning for...  ...by mentoring senior and staff-level...  ...experience building, training, and deploying... 
    Senior
    Training
    Full time
    Remote work
    Flexible hours

    X Company

    Mountain View, CA
    4 days ago
  • $213k - $263k

     ...the foundation for training and validating...  ...advanced ML and engineering team that leverages...  ...computer vision, deep learning, and generative...  ...will report to a Senior Staff Technical...  .../ multimodal models (e.g., Gemini) to...  ...Fine-tuning and Reinforcement Learning (RL) techniques... 
    Senior
    Training
    Full time
    Remote work

    Waymo

    Mountain View, CA
    1 day ago
  • $210.3k - $273.4k

     .... We’re looking for a Senior Machine Learning Engineer to lead the development of...  ...AI, multi-agent systems, reinforcement learning, and intelligent...  ...projects. * Develop AI models for the Unity AI agentic...  ...Employee Assistance Program | Training and development programs... 
    Senior
    Training
    Work at office
    Worldwide

    Unity

    Mountain View, CA
    4 days ago
  • $204k - $259k

     ...automated algorithms. The learned metrics team is a strategic bet to use machine learning to ensure...  ...machine learning models to deliver training and evaluation data...  ...researchers and software engineers who are passionate...  ...~ Experience in reinforcement learning, transfer learning... 
    Senior
    Training
    Full time
    Work experience placement
    Remote work

    Waymo

    Mountain View, CA
    2 days ago
  • $195k - $230k

     ...visit About the RoleWe are looking for a Senior Machine Learning Engineer to help evolve our large-scale...  ...business metrics.Own systems from offline training online inference A/B experimentation...  ...issues related to data quality, model drift, and system performance in production... 
    Senior
    Training
    Full time
    Local area
    Work from home

    News Break

    Mountain View, CA
    12 hours ago
  • $281k - $356k

     ...the focus is scaling. This means larger models trained on more data for better generalization....  ...innovations. You'll build active learning and ML-aided labeling workflows to tackle...  ...You'll partner closely with Product and Engineering teams, especially those focused on data... 
    Senior
    Training
    Full time
    Temporary work
    Remote work

    Waymo

    Mountain View, CA
    16 hours ago
  • $230k - $265k

     ...alongside industry-veteran scientists and engineers. As a Senior Machine Learning Engineer, you’ll bring your strong...  ...the design and implementation of training, fine-tuning, post-training, and...  ...for large language and speech models using PyTorch and/or JAX, making principled... 
    Senior
    Training
    Permanent employment

    Otter.ai

    Mountain View, CA
    4 days ago
  • $184k - $287.5k

    We are seeking a Senior Machine Learning Engineer to join our end‑to‑end autonomous driving team! You will help build, train, and deploy large‑scale E2E driving models that leverage VLM/VLA architectures...  ...scale behavior cloning, or reinforcement/imitation learning for... 
    Senior
    Training
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $188.5k - $282.7k

     ...Rubrik's Semantic AI Governance Engine, which is the first system...  ...'s custom small language models act as judges on every agent...  ...model lifecycle: curating data, training small models, serving them...  ...) in Computer Science, Machine Learning, Computer Engineering, Statistics... 
    Senior
    Training
    Permanent employment
    Local area

    Rubrik

    Palo Alto, CA
    3 days ago
  • $170k - $225k

     ...skill, awareness, and learning capabilities. Our...  ...the Role As a Senior/Staff Machine Learning Engineer, you will be working...  ...imitation learning, reinforcement learning, and other...  ...’s internal world model. Dexterity has...  ...curation, labeling, training, evaluation,... 
    Senior
    Training
    Full time
    Worldwide

    Dexterity

    Redwood City, CA
    16 hours ago
  • $175k - $296k

     ...of the future. We are looking for a full-time Machine Learning Engineer, with deep knowledge and strong enthusiasm towards establishing a state-of-art ML infrastructure for training very large foundation model and accelerating model training/inference. Our... 
    Senior
    Training
    Full time

    XPeng Motors

    Santa Clara, CA
    more than 2 months ago
  • $224k - $356.5k

     ...are seeking exceptional Senior Machine Learning and Simulation Engineers to join NVIDIA's Autonomous...  ...including deep learning, reinforcement learning, end-to-end driving and Physics AI models. The successful candidate...  ...) framework in order to train advanced end-to-end AV models... 
    Senior
    Training
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $242k - $349k

     ...for developing Foundation Models for ML Agents and planning...  ...an ML Agents and Planning Machine Learning Engineer you will work on the bleeding...  ...imitation learning and reinforcement learning to generate driving...  ...Experience with training and deploying transformer-... 
    Senior
    Training
    Full time
    Temporary work
    Relocation package

    Zoox

    Foster, CA
    4 days ago
  •  ...LLMs, our proprietary models, and a sophisticated Agentic...  ..., and continuously learn and adapt.Moveworks is...  ...with Moveworks’ Reasoning Engine and natural language...  ...RoleWe are looking for a Machine Learning Engineer to...  ...including distributed training and inference pipeline... 
    Senior
    Training
    Work at office
    Remote work
    Flexible hours

    Moveworks

    Mountain View, CA
    16 hours ago
  • $70 - $85 per hour

     ...Senior Machine Learning Engineer (LLM Evaluation & AI Agents) *Location:* Menlo Park...  ...role focuses on evaluating model performance, identifying failure...  ...improve model quality and training outcomes. You'll work...  ...systems * Exposure to Reinforcement Learning (RL) or... 
    Senior
    Training
    Contract work
    Temporary work

    TEKsystems

    Menlo Park, CA
    5 days ago
  • $213k - $263k

     .... states. Waymo's Systems Engineering team works together to blend...  ...Manager. You will: Lead ML model development end-to-end for...  ...and Python ~ Ability to learn and utilize new frameworks quickly...  ..., experience, relevant training and education, and skill level... 
    Senior
    Training
    Full time
    Remote work

    Waymo

    Mountain View, CA
    2 days ago
  • $324k

     ...trajectory. As a Principal ML Engineer, you will lead the...  ...development of machine learning algorithms for our next...  ...in Foundation Models, Perception, Simulation...  ...imitation learning, reinforcement learning, and model scaling...  ...metrics, data, and training frameworksDevelop new... 
    Senior
    Training
    Full time
    Temporary work
    Immediate start
    Relocation package

    Zoox

    Foster, CA
    2 days ago
  • $213k - $263k

     ...the lifecycle of the machine learning workflow, including feature...  ...management, model development, optimization...  ...We are looking for engineers with ML software & systems...  ...you will report to the Senior Manager of Runtime and...  ...experience, relevant training and education, and skill... 
    Senior
    Training
    Full time
    Remote work

    Waymo

    Mountain View, CA
    2 days ago
  • $204k - $259k

     ...team builds the system which learns the spatial-temporal...  ...diverse set of sensors, enabling engineers like you to (1) develop...  ...world data, to (2) develop models and model training at scale, to (3) analyze real...  ...~5+ years of experience in Machine Learning, with a focus on large... 
    Senior
    Training
    Full time
    Remote work

    Waymo

    Mountain View, CA
    2 days ago
  • $193.93k - $291.15k

     ...partner-led business model, Nuro is working toward...  ...looking for a Software Engineer to join our Sensor...  ...engineer with robotics and machine learning expertise to develop...  ...requirementsRole is scoped as a Senior/Staff IC with the...  ...-on experience in training and evaluating modern... 
    Senior
    Training
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    2 days ago
  •  ...LLMs, our proprietary models, and a sophisticated Agentic...  ..., and continuously learn and adapt. Moveworks...  ...with Moveworks’ Reasoning Engine and natural language...  ...software engineer with machine learning expertise to join...  ...also go beyond model training to achieve state-of-the... 
    Senior
    Training
    Full time
    Work at office
    Immediate start
    Remote work
    Flexible hours

    ServiceNow

    Mountain View, CA
    17 days ago
  • $251k - $310k

     ...realistic environments for testing and training the Waymo Driver. Our team is a...  ...and collaborative group of software engineers, machine learning (ML) engineers, and data scientists....  ...Driver. By applying machine learning, we model the real world, including realistic agents... 
    Senior
    Training
    Full time
    Remote work

    Waymo

    Mountain View, CA
    2 days ago
  • $213k - $263k

     ....S. states. Software Engineering builds the brains of Waymo...  ...-making and deep learning, while collaborating with...  ...APIs to Perception models, to abstract sensor hardware...  ...through algorithms or machine learning. Projects on...  ...experience, relevant training and education, and... 
    Senior
    Training
    Full time
    Remote work

    Waymo

    Mountain View, CA
    2 days ago
  • $210k - $225k

     ...s where you come in. The Machine Learning team forms the core of Suki'...  ...data quality and relevance for training and evaluation. You'll...  ...relentlessly. You think business and engineering problems are like puzzles...  ...moving at least one NLP model to an enterprise production... 
    Senior
    Training
    Full time

    Suki

    Redwood City, CA
    16 hours ago
  • $174.72k - $295.68k

     ...through cutting-edge R&D in AI, machine learning, and smart connectivity.Our...  ...and is not limited to: LLM model fine tuning, PTQ, QAT, on-...  ...numerical consistency with training models, and productionize LLM...  ...Python programming and software engineering skills.Ability to work... 
    Senior
    Training
    Full time

    XPENG Motors

    Santa Clara, CA
    3 days ago
  • $242k - $290k

     .../HybridAs a Perception Engineer, you will be instrumental...  ...cutting-edge detection models, utilizing fused sensor...  ...In this role, you will:Train ML models, perform...  ...frameworksExperience deploying learned models into...  ...intersection of robotics, machine learning, and design, Zoox... 
    Senior
    Training
    Full time
    Temporary work
    Relocation package

    Zoox

    Foster, CA
    4 days ago
  • $195.2k - $361.2k

     ...design. Small, efficient models run directly on the user's machine (AI PC, edge, on-prem,...  ...seeking a **Machine Learning Engineer / Data Scientist** to join...  ...agent harness and post-training pipelines, develops RL...  ...supervised fine tuning and reinforcement learning.Ability to own... 
    Senior
    Training
    Full time
    Internship
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    4 days ago
  • $224k - $356.5k

    We are looking for outstanding Machine Learning Engineers to join our Physical AI teams. As the pioneers...  ...leverages state-of-the-art multimodal models and diffusion techniques to simulate...  ...environments, ensuring our AI agents are trained on the most diverse and rigorous data... 
    Senior
    Training
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Machine Learning Engineer, Model Training and Reinforcement Learning. Be the first to apply!