Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Systems Engineer

Nebius B.V.

Nebius Token Factory is building an AI training and model post-training capability for frontier model improvement. This role owns the infrastructure that makes large-scale training and RL experiments possible, reliable, reproducible, and efficient. The work sits at the intersection of distributed systems, GPU performance, model training frameworks, RL pipelines, and production engineering. Your responsibilities Build and maintain distributed training infrastructure for SFT, continued pretraining, preference optimization, and RL workloads. Integrate and extend frameworks such as Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, verl, slime, AReaL, OpenRLHF, or equivalent internal systems. Implement and debug parallelism strategies including tensor, pipeline, sequence/context, expert, and data parallelism. Build reliable rollout, reward model serving, replay/data buffer, checkpointing, evaluation, and experiment orchestration components for RL training. Profile and improve GPU utilization, communication efficiency, memory usage, and training throughput. Diagnose failures across NCCL, CUDA, PyTorch, Ray, schedulers, storage, networking, and checkpointing layers. Create reproducible training runs, launch scripts, dashboards, runbooks, and operational tooling for research users. Partner with research scientists to turn algorithmic training recipes into scalable, debuggable systems. Write clear design docs, incident reports, benchmark reports, and operating guides. Must-haves Strong Python and PyTorch engineering skills. Hands-on experience with distributed model training, large-scale ML systems, or GPU cluster workloads. Practical understanding of transformer training bottlenecks, memory pressure, gradient/optimizer state, communication overhead, and checkpointing. Experience debugging production or research training jobs across multiple GPUs or nodes. Ability to reason quantitatively about throughput, utilization, memory, reliability, cost, and research velocity. Strong communication skills and ability to collaborate with researchers, ML engineers, platform engineers, and leadership. Nice-to-haves Experience with Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, Slurm, Kubernetes, or large internal training platforms. Experience with RL infrastructure frameworks such as verl, slime, AReaL, OpenRLHF, TRL, or custom PPO/GRPO/RLHF systems. Familiarity with NCCL, CUDA, Triton, Nsight, InfiniBand, RDMA, RoCE, H100/H200/B200 clusters, or storage/network bottlenecks. Experience supporting SFT, DPO, PPO, GRPO, RLAIF, reward model serving, rollout generation, or agent training workloads. Open-source contributions to distributed training, RL infrastructure, PyTorch, Ray, Megatron, DeepSpeed, or related systems. Strong Python and PyTorch engineering skills, Hands-on experience with distributed model training, large-scale ML systems, or GPU cluster workloads, Practical understanding of transformer training bottlenecks, memory pressure, gradient/optimizer state, communication overhead, and checkpointing, Experience debugging production or research training jobs across multiple GPUs or nodes, Ability to reason quantitatively about throughput, utilization, memory, reliability, cost, and research velocity, Strong communication skills and ability to collaborate with researchers, ML engineers, platform engineers, and leadership #J-18808-Ljbffr

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the ML Systems Engineer in Palo Alto, CA vacancy
  • $144.7k - $261.3k

     ...challenges for autonomous vehicle development. We engineer high-performance tools that identify top-...  ...models and partner with data-intensive ML teams to drive rapid innovation. Why...  ...of next-generation autonomous systems. About the Role As a Senior Engineer... 
    Suggested
    Full time
    Local area
    Remote work
    Work from home
    Flexible hours

    General Motors

    Mountain View, CA
    2 days ago
  • $90.1k - $191.8k

     ...And Understand The World!The Data Labeling Engineering team designs, builds, and operates high-...  ...engineering, data engineering, and ML, defining labeling strategies, tooling, and...  ...technical leadership, and work directly on systems that unblock the next generation of AV models... 
    Suggested
    Work experience placement
    Flexible hours

    General Motors

    Mountain View, CA
    3 days ago
  • General Motors’ Data Labeling Engineering team is building cutting‑edge labeling tools and pipelines that power autonomous vehicle ML models. The role sits at the intersection of software...  ...and ML, focusing on scalable labeling systems and foundations for foundation‑model... 
    Suggested

    General Motors

    Mountain View, CA
    5 days ago
  • $300k - $400k

     ...possible. About the Role You will own the systems layer that makes our frontier model...  ...Profiling and benchmarking distributed ML systems to identify and eliminate bottlenecks...  ...team of the world’s best — the scientists, engineers, and problem-solvers who don’t just follow... 
    Suggested
    Visa sponsorship
    Flexible hours
    Shift work

    Periodic Labs

    Menlo Park, CA
    1 day ago
  •  ...and production-grade workflows. You will work at the intersection of distributed systems, GPU performance, and ML framework integration. The role requires strong Python and PyTorch engineering skills, hands-on experience with distributed model training, and the ability to... 
    Suggested

    Nebius B.V.

    Palo Alto, CA
    1 day ago
  • Recruiting From Scratch is seeking a Machine Learning Systems Engineer to design and operate large-scale ML training and inference infrastructure in Palo Alto. The role focuses on building high-performance, GPU-accelerated systems for model serving and deployment across... 

    Recruiting from Scratch

    Palo Alto, CA
    2 days ago
  • Arch Systems is looking for a talented individual to design and implement optimal algorithms for Wi-Fi network performance, leveraging...  ...at least three years of experience in software or systems engineering. Key skills include Wi-Fi products development, WLAN management... 
    Remote work

    Arch Systems

    Palo Alto, CA
    3 days ago
  • $204k - $259k

     ...autonomous driving technology company is looking for an experienced engineer to improve compute performance in machine learning systems. This hybrid role involves collaboration with a world-class ML team and requires strong expertise in ML software or systems. The ideal... 

    Waymo

    Mountain View, CA
    3 days ago
  • Rhoda is building the next generation of generalist robotic systems in Mountain View, CA. We are seeking a senior or staff-level Research Engineer or ML Systems Engineer to make the robot-learning pipeline reliable and measurable from end to end. You will own the supported... 

    RHODA

    Mountain View, CA
    1 day ago
  • Rhoda AI is hiring a Senior/Staff-level Research Engineer to ensure our robot-learning pipeline is reliable from data collection through...  ..., and real-robot evaluation. You will build validation systems, observability, and robust operating practices to distinguish model... 

    Socket.dev

    Mountain View, CA
    5 days ago
  •  ...Member of Technical Staff — Training to design and optimize large-scale distributed training systems for frontier AI models. Candidates should have 5+ years of experience in ML systems and be proficient in Python along with another systems language, such as C++. This role... 

    RadixArk

    Palo Alto, CA
    5 days ago
  • Rhoda AI in Mountain View is seeking a Staff / Principal ML Training Systems Engineer to lead the performance of large-scale multimodal training systems. This role involves improving training efficiency and collaborating closely with research teams to accelerate model... 

    Rhoda AI

    Mountain View, CA
    2 days ago
  • $224k - $356.5k

     ...the next phase, we are building agentic systems that can reason about, build, evaluate, and...  ...about creating the meta-layer of modern ML: the agents, tooling, pipelines, and feedback...  .... We are looking for exceptional engineers who are passionate about the idea of AI-native... 
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  • $174.9k - $261.3k

     ...and understand the world!The Data Labeling Engineering team designs, builds, and operates hybrid...  ...engineering, data engineering, and AI/ML, defining the strategies, tooling, and quality...  ...leadership, and direct impact on systems that unblock the next generation of AV capabilities... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    4 days ago
  • Rhoda AI is seeking a Staff / Principal ML Training Systems Engineer in Mountain View to enhance the training systems performance. This role focuses on large-scale multimodal training, driving efficiency, and scalability in compute use across thousands of GPUs. The ideal... 

    Rhoda AI

    Mountain View, CA
    4 days ago
  • Rhoda AI is building next-generation generalist robots and is seeking a Staff / Principal ML Training Systems Engineer to own training systems performance end-to-end. You will optimize large-scale multimodal training, define parallelism strategies, and drive efficiency... 

    Rhoda AI

    Palo Alto, CA
    4 days ago
  •  ...own the execution layer of our intelligence platform in Palo Alto. You will translate research direction into reliable, scalable ML systems deployed in production, collaborating with infra, product, and applications teams. Responsibilities include end-to-end ML pipelines... 

    A1

    Palo Alto, CA
    4 days ago
  • $200k - $300k

     ...across the United States to help them hire. Machine Learning Systems Engineer Location - Palo Alto, CA (On-site) - Five days per week...  ...Design, build, and operate infrastructure supporting large-scale ML training and inference systems Develop high-performance... 
    H1b
    Work at office
    Remote work

    Recruiting from Scratch

    Palo Alto, CA
    4 days ago
  •  ...ML Systems Engineer Role Overview: We're seeking an experienced engineer to build our ML data infrastructure platform. You'll create the systems and tools that enable efficient data preparation, feature engineering, and dataset management for machine learning. This... 
    Local area

    My3Tech Inc

    Sunnyvale, CA
    1 day ago
  •  ...ML Systems Engineer Role OverviewWe’re seeking an experienced engineer to build our ML data infrastructure platform. You’ll create the systems and tools that enable efficient data preparation, feature engineering, and dataset management for machine learning. This role... 

    My3Tech Inc

    Sunnyvale, CA
    5 days ago
  •  ...deployed on-premise inside sovereign data centers, delivering low-latency, high-throughput AI for multi-node GPU workloads. As a Senior Engineer, you will shape core infrastructure and architecture decisions, lead performance optimizations, and own the inference engine to... 

    Sanas

    Palo Alto, CA
    1 day ago
  • $227.87k

     ...developing and executing solutions that enhance the Pinterest marketplace, primarily focusing on backend systems and statistical models. Candidates need strong software engineering skills and a graduate degree in a related field. The position is unique, allowing collaboration... 
    Work at office
    Flexible hours

    Pinterest

    Palo Alto, CA
    4 days ago
  • NVIDIA is seeking an ML and Agentic Systems Engineer (Finance) to help build the meta-layer of modern ML: agents, tooling, pipelines, and feedback loops that accelerate model development. You will own large Python/PyTorch codebases and design end-to-end agentic workflows... 

    Nvidia Corporation in

    Santa Clara, CA
    3 days ago
  • JPMorgan Chase & Co. is seeking a Senior Machine Learning Engineer-Digital Intelligence in the Digital Intelligence team. You will specialize...  ..., interpretability, and related algorithms, driving end-to-end ML solutions in a fast-paced environment. Ideal candidates bring a... 

    JPMorgan Chase & Co.

    Palo Alto, CA
    3 days ago
  • SpaceXAI is seeking exceptional Applied engineers to join a high-priority project used by hundreds of millions of users monthly. You will...  ...and real-world impact, applying your skills to recommendation systems, ranking algorithms, search technologies, and more. You will design... 

    Pantera Capital

    Palo Alto, CA
    2 days ago
  •  ...hiring for a role in the Cosmos team to design and build agentic ML systems that accelerate model development. You will own large Python/...  ...spanning data generation, evaluation, and deployment. We seek engineers with deep ML system experience, strong software fundamentals,... 

    NVIDIA

    Santa Clara, CA
    1 day ago
  • NVIDIA Gruppe is seeking a Senior Engineer in Santa Clara, CA, to join the Cosmos team. This role focuses on creating AI-native systems that enhance the efficiency of machine learning workflows. Candidates should have extensive Python and PyTorch experience, along with... 

    NVIDIA Gruppe

    Santa Clara, CA
    3 days ago
  • Waymo is seeking a seasoned ML/Computer Vision Engineer to advance the Waymo Driver stack. You will own ML tasks, optimize performance, and scale...  ...in a production setting. You will analyze and monitor ML systems, build AI-aided debugging tools, and develop metrics for safety... 

    Neura Market

    Mountain View, CA
    4 days ago
  • $169k - $338k

     ...Segment: Home OfficePosition Summary...As a Distinguished AI/ML Engineer within Walmart Global Tech's Site Reliability Engineering organization...  ...lead the technical development of next-generation agentic AI systems and intelligent automation solutions that ensure mission-... 
    Full time
    Temporary work
    Part time

    Walmart

    Sunnyvale, CA
    1 day ago
  •  ...About the Role We’re looking for an Applied ML Engineer to design, evaluate, and scale recommendation and ranking systems that power how content, ads, and interactive experiences are selected and surfaced in real time. This role focuses on decision-making systems, with... 
    Full time

    Darwin

    Palo Alto, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Systems Engineer. Be the first to apply!