ML Systems Engineer
Nebius B.V.
Nebius Token Factory is building an AI training and model post-training capability for frontier model improvement. This role owns the infrastructure that makes large-scale training and RL experiments possible, reliable, reproducible, and efficient. The work sits at the intersection of distributed systems, GPU performance, model training frameworks, RL pipelines, and production engineering. Your responsibilities Build and maintain distributed training infrastructure for SFT, continued pretraining, preference optimization, and RL workloads. Integrate and extend frameworks such as Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, verl, slime, AReaL, OpenRLHF, or equivalent internal systems. Implement and debug parallelism strategies including tensor, pipeline, sequence/context, expert, and data parallelism. Build reliable rollout, reward model serving, replay/data buffer, checkpointing, evaluation, and experiment orchestration components for RL training. Profile and improve GPU utilization, communication efficiency, memory usage, and training throughput. Diagnose failures across NCCL, CUDA, PyTorch, Ray, schedulers, storage, networking, and checkpointing layers. Create reproducible training runs, launch scripts, dashboards, runbooks, and operational tooling for research users. Partner with research scientists to turn algorithmic training recipes into scalable, debuggable systems. Write clear design docs, incident reports, benchmark reports, and operating guides. Must-haves Strong Python and PyTorch engineering skills. Hands-on experience with distributed model training, large-scale ML systems, or GPU cluster workloads. Practical understanding of transformer training bottlenecks, memory pressure, gradient/optimizer state, communication overhead, and checkpointing. Experience debugging production or research training jobs across multiple GPUs or nodes. Ability to reason quantitatively about throughput, utilization, memory, reliability, cost, and research velocity. Strong communication skills and ability to collaborate with researchers, ML engineers, platform engineers, and leadership. Nice-to-haves Experience with Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, Slurm, Kubernetes, or large internal training platforms. Experience with RL infrastructure frameworks such as verl, slime, AReaL, OpenRLHF, TRL, or custom PPO/GRPO/RLHF systems. Familiarity with NCCL, CUDA, Triton, Nsight, InfiniBand, RDMA, RoCE, H100/H200/B200 clusters, or storage/network bottlenecks. Experience supporting SFT, DPO, PPO, GRPO, RLAIF, reward model serving, rollout generation, or agent training workloads. Open-source contributions to distributed training, RL infrastructure, PyTorch, Ray, Megatron, DeepSpeed, or related systems. Strong Python and PyTorch engineering skills, Hands-on experience with distributed model training, large-scale ML systems, or GPU cluster workloads, Practical understanding of transformer training bottlenecks, memory pressure, gradient/optimizer state, communication overhead, and checkpointing, Experience debugging production or research training jobs across multiple GPUs or nodes, Ability to reason quantitatively about throughput, utilization, memory, reliability, cost, and research velocity, Strong communication skills and ability to collaborate with researchers, ML engineers, platform engineers, and leadership #J-18808-Ljbffr Nebius B.V.
$300k - $400k
...possible. About the Role You will own the systems layer that makes our frontier model... ...Profiling and benchmarking distributed ML systems to identify and eliminate bottlenecks... ...team of the world’s best — the scientists, engineers, and problem-solvers who don’t just follow...SuggestedVisa sponsorshipFlexible hoursShift work$224k - $356.5k
...the next phase, we are building agentic systems that can reason about, build, evaluate, and... ...about creating the meta-layer of modern ML: the agents, tooling, pipelines, and feedback... .... We are looking for exceptional engineers who are passionate about the idea of AI-native...Suggested$150k - $300k
...Careers. Overview: We are seeking an accomplished Senior Staff ML Engineer who will serve as a technical leader for Generative AI... ...team of AI and software engineers to design, develop, and deploy systems that ensure scalability, reliability and usability of generative...SuggestedHourly payFull timeWork experience placementLocal area- ...About the Role We’re looking for an Applied ML Engineer to design, evaluate, and scale recommendation and ranking systems that power how content, ads, and interactive experiences are selected and surfaced in real time. This role focuses on decision-making systems, with...SuggestedFull time
- ...+ metadata lake Experience: 6+ years industry overall experience with 3+ in ML Infra or MLE Expertise: back end software engineering strength with recent industry exp making ML systems more reliable/scalable (with opportunities to help improve model quality in the...SuggestedFull timePart timeImmediate start
- ...Oracle is seeking a Principal AI Agent / ML Software Engineer in Santa Clara, California, to provide technical leadership in developing next-generation AI systems on Oracle Cloud Infrastructure. The ideal candidate will have extensive experience in building scalable AI...Full time
- ...electric vehicle company in Palo Alto is seeking a Sr./Staff ML Engineer to contribute to the development of advanced machine learning... ...extensive experience in working with complex datasets in safety-critical systems. Competitive compensation package offered. #J-18808-LjbffrFull time
- Grindr is seeking a Staff ML Engineer in Palo Alto to architect innovative machine learning systems aimed at enhancing user connections. You will play a pivotal role in developing scalable recommendation frameworks and cutting-edge solutions, directly impacting the LGBTQ+...Full time
- ...content intelligence platform is seeking a Staff Machine Learning Engineer in Mountain View, California. This role offers the opportunity to provide technical leadership for advanced recommendation systems and AI initiatives, guiding large-scale projects across teams....Full time
- A leading e-commerce company is seeking a Senior Staff Machine Learning Engineer in Mountain View, CA. You will be responsible for developing machine learning models and optimizing advertising features for an innovative platform. Ideal candidates have at least 8 years...Full time
$208k - $244k
...Join to apply for the Staff ML Engineer role at Grindr Join to apply for the Staff ML Engineer role at Grindr Get AI-powered advice on... ...our long term ML strategy. Recommendations That Reshape: Build systems that match millions to their next big moment, adapting to a range...Full timeCasual workWork at officeImmediate startFlexible hours- Job Title: ML Engineer What You Will Own End‑to‑End ML Lifecycle across real products: data ingestion, feature design, model selection,... ..., monitoring and iteration. No handoffs. Production‑grade ML systems built with PyTorch or TensorFlow, focusing on latency, reliability...
- A growing AI technology startup is seeking an ML Engineer to design and deploy production-grade systems. The role involves using Python and collaborating with teams to optimize customer interactions through advanced AI applications. Candidates should have a degree in Computer...
- A cutting-edge AI company is seeking a passionate Machine Learning Engineer to join their Applied Safety team in Palo Alto, California. You will drive innovative ML solutions to enhance user safety and compliance with X’s Terms of Service. Candidates should have 5+ years...Full time
- ...A financial technology company is seeking a Principal Machine Learning Engineer in Mountain View, California. This role involves leading AI strategy and deploying AI/ML solutions across financial products. Candidates should have over 10 years of experience in ML development...Full time
$283.64k - $425.96k
...investors. About The Role Nuro is looking for a Head of Systems Engineering to own the systems backbone that enables the Nuro Driver to... ...integration) and the validation challenges specific to ML‑based and end‑to‑end AI systems in real‑world environments....Odd jobFull timeImmediate startFlexible hours$197k - $266.5k
A leading financial software firm is seeking a Staff Machine Learning Engineer to join their vibrant team in Mountain View, CA. The ideal candidate will have over 6 years of experience and strong knowledge of data science tools such as Python and SQL. Responsibilities...Full time- A tech innovation company in Mountain View, California, is seeking a Machine Learning Research Engineer. You'll design sophisticated robot learning algorithms to enhance dexterous manipulation in home environments. The role requires 3+ years in machine learning for robotics...Full time
- ...the clean energy transition. About The Role As an AI/ML Engineer at Powerline, you will be instrumental in developing, optimizing... ...and prediction. ~ Familiarity with MLOps and building AI/ML systems end to end. ~ Strong communication skills and ability to...Full time
- ...A leading automotive technology company is seeking an experienced Engineering Manager in Palo Alto, California. This role will involve guiding the development of robust streaming and analytics pipelines, leading a team of data professionals, and ensuring data security...Full time
- ...A leading technology company in Mountain View is seeking a Senior Software Engineer focused on AI/ML for YouTube. The role involves writing and testing code, designing recommendation systems, and collaborating with peers. Candidates should have a Bachelor's degree, extensive...Full time
- ...We bridge this exact gap by applying deep systems programming, software-defined networking,... ...at UT Austin and world-renowned ML systems researcher with a pedigree spanning... ...Seniority ~5+ years of production experience engineering ML systems, OR a PhD from a top-tier...Shift work
$205k - $235k
...Year. As a Senior Machine Learning Engineer , you will be responsible for building machine learning models/systems and innovative web applications that deliver the... ...~ Experience with building and evolving ML Training and Inferencing systems at significant...Full timeWork at officeLocal areaFlexible hours3 days per week$170k - $190k
...relentless focus on outcomes. ASAPP’s AI Engineering team is seeking an enterprising, talented... ...are seeking a highly experienced Lead AI/ML Engineer to join our Core GenerativeAgent... ...building, and deploying cutting-edge AI systems that power mission-critical enterprise applications...Full time$160k - $225k
...capital will be used to expand our product and engineering teams, bringing our vision of intelligent... ...them to reason, to the scalable serving systems that deliver their intelligence to our... .... You'll build the robust pipelines and ML serving systems that fuel our agents with...Full time- ...starting out with understanding and building hardware; electronics systems and semiconductors where AI can design and create beyond... ...four US presidents. What we're Looking For Strong AI/ML engineering skills from top tier CS, EECS, Math and Physics programs....Full time
$195k - $230k
...information powered by advanced AI, recommendation systems, and adtech. Recognized by Fast... ...looking for a Senior Machine Learning Engineer to help evolve our large-scale... ...Experience working with large-scale data and ML systems (e.g., Spark, distributed training...Full timeLocal areaWork from home$175k - $275k
...raw data, curated datasets, or full-cycle data engineering, Abaka AI provides the foundation for building high-performance AI systems. About the Role We’re hiring... ...of experience in applied machine learning or ML engineering, with a demonstrated ability to deliver...Full timeImmediate startFlexible hours- ...Together, we can make a meaningful impact. See more about our culture on . About The Job Mistral AI is seeking a Applied AI Engineer to facilitate the adoption of its products among customers and collaborate with them to address complex technical challenges....Full timeWork at officeVisa sponsorship
- ...and collaborating across teams to optimize architecture. A Bachelor's degree in Computer Science or equivalent and strong software engineering skills are required. This is a hybrid position, requiring three days on-site per week. The position offers a competitive salary...Full time3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Systems Engineer. Be the first to apply!
- machine learning engineer Palo Alto, CA
- operating system engineer Palo Alto, CA
- computer system validation engineer Palo Alto, CA
- system performance engineer Palo Alto, CA
- senior linux systems engineer Palo Alto, CA
- sr systems engineer Palo Alto, CA
- healthcare systems engineer Palo Alto, CA
- systems engineer Palo Alto, CA
- senior windows systems engineer Palo Alto, CA
- software system engineer Palo Alto, CA















