Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Systems Engineer

Nebius B.V.

Nebius Token Factory is building an AI training and model post-training capability for frontier model improvement. This role owns the infrastructure that makes large-scale training and RL experiments possible, reliable, reproducible, and efficient. The work sits at the intersection of distributed systems, GPU performance, model training frameworks, RL pipelines, and production engineering. Your responsibilities Build and maintain distributed training infrastructure for SFT, continued pretraining, preference optimization, and RL workloads. Integrate and extend frameworks such as Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, verl, slime, AReaL, OpenRLHF, or equivalent internal systems. Implement and debug parallelism strategies including tensor, pipeline, sequence/context, expert, and data parallelism. Build reliable rollout, reward model serving, replay/data buffer, checkpointing, evaluation, and experiment orchestration components for RL training. Profile and improve GPU utilization, communication efficiency, memory usage, and training throughput. Diagnose failures across NCCL, CUDA, PyTorch, Ray, schedulers, storage, networking, and checkpointing layers. Create reproducible training runs, launch scripts, dashboards, runbooks, and operational tooling for research users. Partner with research scientists to turn algorithmic training recipes into scalable, debuggable systems. Write clear design docs, incident reports, benchmark reports, and operating guides. Must-haves Strong Python and PyTorch engineering skills. Hands-on experience with distributed model training, large-scale ML systems, or GPU cluster workloads. Practical understanding of transformer training bottlenecks, memory pressure, gradient/optimizer state, communication overhead, and checkpointing. Experience debugging production or research training jobs across multiple GPUs or nodes. Ability to reason quantitatively about throughput, utilization, memory, reliability, cost, and research velocity. Strong communication skills and ability to collaborate with researchers, ML engineers, platform engineers, and leadership. Nice-to-haves Experience with Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, Slurm, Kubernetes, or large internal training platforms. Experience with RL infrastructure frameworks such as verl, slime, AReaL, OpenRLHF, TRL, or custom PPO/GRPO/RLHF systems. Familiarity with NCCL, CUDA, Triton, Nsight, InfiniBand, RDMA, RoCE, H100/H200/B200 clusters, or storage/network bottlenecks. Experience supporting SFT, DPO, PPO, GRPO, RLAIF, reward model serving, rollout generation, or agent training workloads. Open-source contributions to distributed training, RL infrastructure, PyTorch, Ray, Megatron, DeepSpeed, or related systems. Strong Python and PyTorch engineering skills, Hands-on experience with distributed model training, large-scale ML systems, or GPU cluster workloads, Practical understanding of transformer training bottlenecks, memory pressure, gradient/optimizer state, communication overhead, and checkpointing, Experience debugging production or research training jobs across multiple GPUs or nodes, Ability to reason quantitatively about throughput, utilization, memory, reliability, cost, and research velocity, Strong communication skills and ability to collaborate with researchers, ML engineers, platform engineers, and leadership #J-18808-Ljbffr Nebius B.V.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the ML Systems Engineer in Palo Alto, CA vacancy
  • $300k - $400k

     ...possible. About the Role You will own the systems layer that makes our frontier model...  ...Profiling and benchmarking distributed ML systems to identify and eliminate bottlenecks...  ...team of the world’s best — the scientists, engineers, and problem-solvers who don’t just follow... 
    Suggested
    Visa sponsorship
    Flexible hours
    Shift work

    Periodic Labs

    Menlo Park, CA
    4 days ago
  • $224k - $356.5k

     ...the next phase, we are building agentic systems that can reason about, build, evaluate, and...  ...about creating the meta-layer of modern ML: the agents, tooling, pipelines, and feedback...  .... We are looking for exceptional engineers who are passionate about the idea of AI-native... 
    Suggested

    NVIDIA

    Santa Clara, CA
    5 days ago
  • $150k - $300k

     ...Careers. Overview: We are seeking an accomplished Senior Staff ML Engineer who will serve as a technical leader for Generative AI...  ...team of AI and software engineers to design, develop, and deploy systems that ensure scalability, reliability and usability of generative... 
    Suggested
    Hourly pay
    Full time
    Work experience placement
    Local area

    GEICO

    Palo Alto, CA
    3 days ago
  •  ...About the Role We’re looking for an Applied ML Engineer to design, evaluate, and scale recommendation and ranking systems that power how content, ads, and interactive experiences are selected and surfaced in real time. This role focuses on decision-making systems, with... 
    Suggested
    Full time

    Darwin

    Palo Alto, CA
    1 day ago
  •  ...+ metadata lake Experience: 6+ years industry overall experience with 3+ in ML Infra or MLE Expertise: back end software engineering strength with recent industry exp making ML systems more reliable/scalable (with opportunities to help improve model quality in the... 
    Suggested
    Full time
    Part time
    Immediate start

    Greylock Partners

    Palo Alto, CA
    9 hours ago
  •  ...Oracle is seeking a Principal AI Agent / ML Software Engineer in Santa Clara, California, to provide technical leadership in developing next-generation AI systems on Oracle Cloud Infrastructure. The ideal candidate will have extensive experience in building scalable AI... 
    Full time

    Oracle

    Santa Clara, CA
    9 hours ago
  •  ...electric vehicle company in Palo Alto is seeking a Sr./Staff ML Engineer to contribute to the development of advanced machine learning...  ...extensive experience in working with complex datasets in safety-critical systems. Competitive compensation package offered. #J-18808-Ljbffr
    Full time

    Rivian

    Palo Alto, CA
    9 hours ago
  • Grindr is seeking a Staff ML Engineer in Palo Alto to architect innovative machine learning systems aimed at enhancing user connections. You will play a pivotal role in developing scalable recommendation frameworks and cutting-edge solutions, directly impacting the LGBTQ+... 
    Full time

    Grindr

    Palo Alto, CA
    9 hours ago
  •  ...content intelligence platform is seeking a Staff Machine Learning Engineer in Mountain View, California. This role offers the opportunity to provide technical leadership for advanced recommendation systems and AI initiatives, guiding large-scale projects across teams.... 
    Full time

    News Break

    Mountain View, CA
    9 hours ago
  • A leading e-commerce company is seeking a Senior Staff Machine Learning Engineer in Mountain View, CA. You will be responsible for developing machine learning models and optimizing advertising features for an innovative platform. Ideal candidates have at least 8 years... 
    Full time

    Coupang

    Mountain View, CA
    9 hours ago
  • $208k - $244k

     ...Join to apply for the Staff ML Engineer role at Grindr Join to apply for the Staff ML Engineer role at Grindr Get AI-powered advice on...  ...our long term ML strategy. Recommendations That Reshape: Build systems that match millions to their next big moment, adapting to a range... 
    Full time
    Casual work
    Work at office
    Immediate start
    Flexible hours

    Grindr

    Palo Alto, CA
    9 hours ago
  • Job Title: ML Engineer What You Will Own End‑to‑End ML Lifecycle across real products: data ingestion, feature design, model selection,...  ..., monitoring and iteration. No handoffs. Production‑grade ML systems built with PyTorch or TensorFlow, focusing on latency, reliability... 

    MetAntz

    Palo Alto, CA
    3 days ago
  • A growing AI technology startup is seeking an ML Engineer to design and deploy production-grade systems. The role involves using Python and collaborating with teams to optimize customer interactions through advanced AI applications. Candidates should have a degree in Computer... 

    Catalyst Labs, LLC

    Mountain View, CA
    5 days ago
  • A cutting-edge AI company is seeking a passionate Machine Learning Engineer to join their Applied Safety team in Palo Alto, California. You will drive innovative ML solutions to enhance user safety and compliance with X’s Terms of Service. Candidates should have 5+ years... 
    Full time

    xAI

    Palo Alto, CA
    9 hours ago
  •  ...A financial technology company is seeking a Principal Machine Learning Engineer in Mountain View, California. This role involves leading AI strategy and deploying AI/ML solutions across financial products. Candidates should have over 10 years of experience in ML development... 
    Full time

    Intuit Inc.

    Mountain View, CA
    9 hours ago
  • $283.64k - $425.96k

     ...investors. About The Role Nuro is looking for a Head of Systems Engineering to own the systems backbone that enables the Nuro Driver to...  ...integration) and the validation challenges specific to ML‑based and end‑to‑end AI systems in real‑world environments.... 
    Odd job
    Full time
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    9 hours ago
  • $197k - $266.5k

    A leading financial software firm is seeking a Staff Machine Learning Engineer to join their vibrant team in Mountain View, CA. The ideal candidate will have over 6 years of experience and strong knowledge of data science tools such as Python and SQL. Responsibilities... 
    Full time

    Intuit Inc.

    Mountain View, CA
    9 hours ago
  • A tech innovation company in Mountain View, California, is seeking a Machine Learning Research Engineer. You'll design sophisticated robot learning algorithms to enhance dexterous manipulation in home environments. The role requires 3+ years in machine learning for robotics... 
    Full time

    Sunday Robotics

    Mountain View, CA
    9 hours ago
  •  ...the clean energy transition. About The Role As an AI/ML Engineer at Powerline, you will be instrumental in developing, optimizing...  ...and prediction. ~ Familiarity with MLOps and building AI/ML systems end to end. ~ Strong communication skills and ability to... 
    Full time

    Powerline

    Palo Alto, CA
    1 day ago
  •  ...A leading automotive technology company is seeking an experienced Engineering Manager in Palo Alto, California. This role will involve guiding the development of robust streaming and analytics pipelines, leading a team of data professionals, and ensuring data security... 
    Full time

    Rivian and Volkswagen Group Technologies

    Palo Alto, CA
    9 hours ago
  •  ...A leading technology company in Mountain View is seeking a Senior Software Engineer focused on AI/ML for YouTube. The role involves writing and testing code, designing recommendation systems, and collaborating with peers. Candidates should have a Bachelor's degree, extensive... 
    Full time

    Google Inc.

    Mountain View, CA
    9 hours ago
  •  ...We bridge this exact gap by applying deep systems programming, software-defined networking,...  ...at UT Austin and world-renowned ML systems researcher with a pedigree spanning...  ...Seniority ~5+ years of production experience engineering ML systems, OR a PhD from a top-tier... 
    Shift work

    Success Matcher Recruitment

    Sunnyvale, CA
    25 days ago
  • $205k - $235k

     ...Year.    As a  Senior Machine Learning Engineer , you will be responsible for building machine learning models/systems and innovative web applications that deliver the...  ...~ Experience with building and evolving ML Training and Inferencing systems at significant... 
    Full time
    Work at office
    Local area
    Flexible hours
    3 days per week

    Typeface

    Palo Alto, CA
    1 day ago
  • $170k - $190k

     ...relentless focus on outcomes. ASAPP’s AI Engineering team is seeking an enterprising, talented...  ...are seeking a highly experienced Lead AI/ML Engineer to join our Core GenerativeAgent...  ...building, and deploying cutting-edge AI systems that power mission-critical enterprise applications... 
    Full time

    Asapp

    Mountain View, CA
    1 day ago
  • $160k - $225k

     ...capital will be used to expand our product and engineering teams, bringing our vision of intelligent...  ...them to reason, to the scalable serving systems that deliver their intelligence to our...  .... You'll build the robust pipelines and ML serving systems that fuel our agents with... 
    Full time

    Mai

    Mountain View, CA
    1 day ago
  •  ...starting out with understanding and building hardware; electronics systems and semiconductors where AI can design and create beyond...  ...four US presidents. What we're Looking For Strong AI/ML engineering skills from top tier CS, EECS, Math and Physics programs.... 
    Full time

    Voltai

    Palo Alto, CA
    1 day ago
  • $195k - $230k

     ...information powered by advanced AI, recommendation systems, and adtech. Recognized by  Fast...  ...looking for a Senior Machine Learning Engineer to help evolve our large-scale...  ...Experience working with large-scale data and ML systems (e.g., Spark, distributed training... 
    Full time
    Local area
    Work from home

    Newsbreak

    Mountain View, CA
    1 day ago
  • $175k - $275k

     ...raw data, curated datasets, or full-cycle data engineering, Abaka AI provides the foundation for building high-performance AI systems.     About the Role   We’re hiring...  ...of experience in applied machine learning or ML engineering, with a demonstrated ability to deliver... 
    Full time
    Immediate start
    Flexible hours

    Abaka Ai

    Palo Alto, CA
    1 day ago
  •  ...Together, we can make a meaningful impact. See more about our culture on  .   About The Job   Mistral AI is seeking a Applied AI Engineer to facilitate the adoption of its products among customers and collaborate with them to address complex technical challenges.... 
    Full time
    Work at office
    Visa sponsorship

    Mistral Ai

    Palo Alto, CA
    1 day ago
  •  ...and collaborating across teams to optimize architecture. A Bachelor's degree in Computer Science or equivalent and strong software engineering skills are required. This is a hybrid position, requiring three days on-site per week. The position offers a competitive salary... 
    Full time
    3 days per week

    MatX

    Mountain View, CA
    9 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Systems Engineer. Be the first to apply!