Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Systems Engineer

Nebius B.V.

Nebius Token Factory is building an AI training and model post-training capability for frontier model improvement. This role owns the infrastructure that makes large-scale training and RL experiments possible, reliable, reproducible, and efficient. The work sits at the intersection of distributed systems, GPU performance, model training frameworks, RL pipelines, and production engineering. Your responsibilities Build and maintain distributed training infrastructure for SFT, continued pretraining, preference optimization, and RL workloads. Integrate and extend frameworks such as Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, verl, slime, AReaL, OpenRLHF, or equivalent internal systems. Implement and debug parallelism strategies including tensor, pipeline, sequence/context, expert, and data parallelism. Build reliable rollout, reward model serving, replay/data buffer, checkpointing, evaluation, and experiment orchestration components for RL training. Profile and improve GPU utilization, communication efficiency, memory usage, and training throughput. Diagnose failures across NCCL, CUDA, PyTorch, Ray, schedulers, storage, networking, and checkpointing layers. Create reproducible training runs, launch scripts, dashboards, runbooks, and operational tooling for research users. Partner with research scientists to turn algorithmic training recipes into scalable, debuggable systems. Write clear design docs, incident reports, benchmark reports, and operating guides. Must-haves Strong Python and PyTorch engineering skills. Hands-on experience with distributed model training, large-scale ML systems, or GPU cluster workloads. Practical understanding of transformer training bottlenecks, memory pressure, gradient/optimizer state, communication overhead, and checkpointing. Experience debugging production or research training jobs across multiple GPUs or nodes. Ability to reason quantitatively about throughput, utilization, memory, reliability, cost, and research velocity. Strong communication skills and ability to collaborate with researchers, ML engineers, platform engineers, and leadership. Nice-to-haves Experience with Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, Slurm, Kubernetes, or large internal training platforms. Experience with RL infrastructure frameworks such as verl, slime, AReaL, OpenRLHF, TRL, or custom PPO/GRPO/RLHF systems. Familiarity with NCCL, CUDA, Triton, Nsight, InfiniBand, RDMA, RoCE, H100/H200/B200 clusters, or storage/network bottlenecks. Experience supporting SFT, DPO, PPO, GRPO, RLAIF, reward model serving, rollout generation, or agent training workloads. Open-source contributions to distributed training, RL infrastructure, PyTorch, Ray, Megatron, DeepSpeed, or related systems. Strong Python and PyTorch engineering skills, Hands-on experience with distributed model training, large-scale ML systems, or GPU cluster workloads, Practical understanding of transformer training bottlenecks, memory pressure, gradient/optimizer state, communication overhead, and checkpointing, Experience debugging production or research training jobs across multiple GPUs or nodes, Ability to reason quantitatively about throughput, utilization, memory, reliability, cost, and research velocity, Strong communication skills and ability to collaborate with researchers, ML engineers, platform engineers, and leadership #J-18808-Ljbffr

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the ML Systems Engineer in Palo Alto, CA vacancy
  • $300k - $400k

     ...possible. About the Role You will own the systems layer that makes our frontier model...  ...operations Profiling and benchmarking distributed ML systems to identify and eliminate...  ...team of the world’s best — the scientists, engineers, and problem-solvers who don’t just follow... 
    Suggested
    Visa sponsorship
    Flexible hours
    Shift work

    Periodic Labs

    Menlo Park, CA
    9 hours ago
  •  ...About the Role We’re looking for an Applied ML Engineer to design, evaluate, and scale recommendation and ranking systems that power how content, ads, and interactive experiences are selected and surfaced in real time. This role focuses on decision-making systems, with... 
    Suggested
    Full time

    Darwin

    Palo Alto, CA
    8 hours ago
  •  ...A growing AI technology startup is seeking an ML Engineer to design and deploy production-grade systems. The role involves using Python and collaborating with teams to optimize customer interactions through advanced AI applications. Candidates should have a degree in Computer... 
    Suggested

    Catalyst Labs, LLC

    Mountain View, CA
    3 days ago
  • $120k - $140k

     ...Deep Learning / NLP Engineer Join Avoma and work on some of the most challenging NLP problems in our mission to make every meeting...  ...looking for a Deep Learning / NLP Engineer to improve and build our systems to extract key insights and topics from conversations. For our... 
    Suggested

    Avoma Inc

    Palo Alto, CA
    5 days ago
  •  ...ML Engineer Palo Alto, California, United States About the Job Our client is a rapidly growing Tier 1 VC backed startup based...  ...long-term growth trajectory in the evolving world of intelligent systems. Location New York, NY Work Type Full Time... 
    Suggested
    Full time

    Catalyst Labs, LLC

    Palo Alto, CA
    3 days ago
  •  ...The Mission: As a Senior Machine Learning Engineer, you will be responsible for building machine learning models/systems and innovative web applications that deliver the power...  ...field Experience with building and evolving ML Training and Inferencing systems at significant... 
    Local area

    Typeface

    Palo Alto, CA
    3 days ago
  • $138.5k - $225.5k

     ...value across the platform. About the role We're hiring a ML Engineer as one of the founding engineers on Intelligence Org. You'll...  ...Responsibilities Design, train, evaluate, and ship ML systems that power governance and security capabilities, starting with... 
    Full time
    Remote work
    Home office
    Visa sponsorship
    Shift work

    Docker

    Palo Alto, CA
    4 days ago
  • $208k - $244k

     ...Join to apply for the Staff ML Engineer role at Grindr Join to apply for the Staff ML Engineer role at Grindr Get AI-powered advice on...  ...our long term ML strategy. Recommendations That Reshape: Build systems that match millions to their next big moment, adapting to a range... 
    Full time
    Casual work
    Work at office
    Immediate start
    Flexible hours

    Grindr

    Palo Alto, CA
    2 days ago
  • Job Title: ML Engineer What You Will Own End‑to‑End ML Lifecycle across real products: data ingestion, feature design, model selection,...  ..., monitoring and iteration. No handoffs. Production‑grade ML systems built with PyTorch or TensorFlow, focusing on latency, reliability... 

    MetAntz

    Palo Alto, CA
    2 days ago
  •  .... Our platform leverages cutting-edge generative AI to assist engineers in RTL design, simulation, and verification, dramatically accelerating...  ...shipped 16B chips. Position Overview We are seeking an ML Systems Engineer to optimize the performance and efficiency of large... 
    Full time

    ScOp Venture Capital

    Santa Clara, CA
    19 days ago
  • $190k - $234k

     ...Staff ML Engineer Palo Alto, CA About Typeface We help the world's biggest brands move from brief to fully personalized campaigns...  ...you will be responsible for building machine learning models/systems and innovative web applications that deliver the power of Generative... 
    Work at office
    Local area
    Flexible hours
    3 days per week

    Typeface

    Palo Alto, CA
    1 day ago
  • $230k - $260k

     ...Principal ML Engineer Palo Alto, CA About Typeface We help the world's biggest brands move from brief to fully personalized campaigns...  ...AI at Typeface. You will lead the design of large-scale ML systems and shared platforms that power all generative capabilities... 
    Work at office
    Immediate start
    Flexible hours
    3 days per week

    Typeface

    Palo Alto, CA
    1 day ago
  • $152k - $261k

     ...services for landing, tracking & attribution. ML opportunities across these portfolios...  ...?’ What You Will Do Drive end-to-end Engineering and Machine learning methods that have a...  ...maintaining highly available, distributed systems Preferred Qualifications Experience in... 
    Full time
    Temporary work
    Work experience placement
    Flexible hours

    Coupang

    Mountain View, CA
    22 hours ago
  •  ...the clean energy transition. About The Role As an AI/ML Engineer at Powerline, you will be instrumental in developing, optimizing...  ...and prediction. ~ Familiarity with MLOps and building AI/ML systems end to end. ~ Strong communication skills and ability to... 
    Full time

    Powerline

    Palo Alto, CA
    8 hours ago
  •  ...We bridge this exact gap by applying deep systems programming, software-defined networking,...  ...at UT Austin and world-renowned ML systems researcher with a pedigree spanning...  ...Seniority ~5+ years of production experience engineering ML systems, OR a PhD from a top-tier... 
    Shift work

    Success Matcher Recruitment

    Sunnyvale, CA
    29 days ago
  • $283.64k - $425.96k

     ...leading investors.About the RoleNuro is looking for a Head of Systems Engineering to own the systems backbone that enables the Nuro Driver to...  ...vehicle integration) and the validation challenges specific to ML-based and end-to-end AI systems in real-world environments.... 
    Odd job
    Part time
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    3 days ago
  • $170k - $190k

     ...relentless focus on outcomes. ASAPP’s AI Engineering team is seeking an enterprising, talented...  ...are seeking a highly experienced Lead AI/ML Engineer to join our Core GenerativeAgent...  ...building, and deploying cutting-edge AI systems that power mission-critical enterprise applications... 
    Full time

    Asapp

    Mountain View, CA
    8 hours ago
  • $205k - $235k

     ...Year.    As a  Senior Machine Learning Engineer , you will be responsible for building machine learning models/systems and innovative web applications that deliver the...  ...~ Experience with building and evolving ML Training and Inferencing systems at significant... 
    Full time
    Work at office
    Local area
    Flexible hours
    3 days per week

    Typeface

    Palo Alto, CA
    8 hours ago
  •  ...business and operations to provide insight and actionable items at real-time. ~ We are looking for streaming, data engineer ~ Understand distributed systems architecture, design and trade-off. ~ Design and develop ETL pipelines with a wide range of technologies.... 
    Permanent employment
    Work experience placement
    Local area
    3 days per week

    3B Staffing LLC

    Menlo Park, CA
    5 days ago
  • $175k - $275k

     ...raw data, curated datasets, or full-cycle data engineering, Abaka AI provides the foundation for building high-performance AI systems.     About the Role   We’re hiring...  ...of experience in applied machine learning or ML engineering, with a demonstrated ability to deliver... 
    Full time
    Immediate start
    Flexible hours

    Abaka Ai

    Palo Alto, CA
    8 hours ago
  • $160k - $225k

     ...capital will be used to expand our product and engineering teams, bringing our vision of intelligent...  ...them to reason, to the scalable serving systems that deliver their intelligence to our...  .... You'll build the robust pipelines and ML serving systems that fuel our agents with... 
    Full time

    Mai

    Mountain View, CA
    8 hours ago
  • $195k - $230k

     ...information powered by advanced AI, recommendation systems, and adtech. Recognized by  Fast...  ...looking for a Senior Machine Learning Engineer to help evolve our large-scale...  ...Experience working with large-scale data and ML systems (e.g., Spark, distributed training... 
    Full time
    Local area
    Work from home

    Newsbreak

    Mountain View, CA
    8 hours ago
  •  ...starting out with understanding and building hardware; electronics systems and semiconductors where AI can design and create beyond...  ...four US presidents. What we're Looking For Strong AI/ML engineering skills from top tier CS, EECS, Math and Physics programs.... 
    Full time

    Voltai

    Palo Alto, CA
    8 hours ago
  •  ...Together, we can make a meaningful impact. See more about our culture on  .   About The Job   Mistral AI is seeking a Applied AI Engineer to facilitate the adoption of its products among customers and collaborate with them to address complex technical challenges.... 
    Full time
    Work at office
    Visa sponsorship

    Mistral Ai

    Palo Alto, CA
    8 hours ago
  • $185k - $254k

    Who We Are Applied Materials is a global leader in materials engineering solutions used to produce virtually every new chip and advanced...  ...focus during later stages of completion. Ensures that all systems engineering projects and programs assigned are completed in accordance... 
    Full time
    Relocation

    Applied Materials

    Santa Clara, CA
    1 hour ago
  •  ...About the Role We are seeking a Senior Data / AI / ML Software Engineer with 7+ years of experience building data-intensive systems. This role is ideal for someone who enjoys designing and improving core platform components at the intersection of software engineering,... 
    Full time
    Contract work
    Internship

    Next Ventures

    Palo Alto, CA
    2 days ago
  • $78 - $83 per hour

     ...Overview Senior Systems Process Engineer — Palo Alto, CA — Duration: 5 month contract — Pay: $78-83/hr This Senior Systems Process Engineer role is responsible for defining and implementing end-to-end processes, methods, and tooling for product configuration, change,... 
    Contract work

    Hydrogen Group

    Palo Alto, CA
    1 day ago
  • $190k - $300k

     ...technology for aviation that will save lives. Automated aviation systems will enable a future where air transportation is safer, more...  ...people - move around the planet. We are a team of mission-driven engineers with experience across aerospace, robotics and self-driving... 
    Permanent employment
    Remote work

    Reliable Robotics Corporation

    Mountain View, CA
    5 days ago
  • $196k - $248k

     ...autonomously driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states. Waymo's Systems Engineering team works together to blend software and hardware systems in groundbreaking new ways. We set the high performance standards... 
    Full time
    Remote work

    Waymo

    Mountain View, CA
    5 days ago
  •  ...Job Overview Senior Systems Process Engineer Duration: 6+ Months Job Description Create process, methods and tool architecture for our product configuration, change, variant, release and sign-off management in a Software Defined Vehicle (SDV) context. Ensure a seamless... 
    Contract work

    eTeam

    Palo Alto, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Systems Engineer. Be the first to apply!