ML Systems Engineer
Nebius B.V.
Nebius Token Factory is building an AI training and model post-training capability for frontier model improvement. This role owns the infrastructure that makes large-scale training and RL experiments possible, reliable, reproducible, and efficient. The work sits at the intersection of distributed systems, GPU performance, model training frameworks, RL pipelines, and production engineering. Your responsibilities Build and maintain distributed training infrastructure for SFT, continued pretraining, preference optimization, and RL workloads. Integrate and extend frameworks such as Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, verl, slime, AReaL, OpenRLHF, or equivalent internal systems. Implement and debug parallelism strategies including tensor, pipeline, sequence/context, expert, and data parallelism. Build reliable rollout, reward model serving, replay/data buffer, checkpointing, evaluation, and experiment orchestration components for RL training. Profile and improve GPU utilization, communication efficiency, memory usage, and training throughput. Diagnose failures across NCCL, CUDA, PyTorch, Ray, schedulers, storage, networking, and checkpointing layers. Create reproducible training runs, launch scripts, dashboards, runbooks, and operational tooling for research users. Partner with research scientists to turn algorithmic training recipes into scalable, debuggable systems. Write clear design docs, incident reports, benchmark reports, and operating guides. Must-haves Strong Python and PyTorch engineering skills. Hands-on experience with distributed model training, large-scale ML systems, or GPU cluster workloads. Practical understanding of transformer training bottlenecks, memory pressure, gradient/optimizer state, communication overhead, and checkpointing. Experience debugging production or research training jobs across multiple GPUs or nodes. Ability to reason quantitatively about throughput, utilization, memory, reliability, cost, and research velocity. Strong communication skills and ability to collaborate with researchers, ML engineers, platform engineers, and leadership. Nice-to-haves Experience with Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, Slurm, Kubernetes, or large internal training platforms. Experience with RL infrastructure frameworks such as verl, slime, AReaL, OpenRLHF, TRL, or custom PPO/GRPO/RLHF systems. Familiarity with NCCL, CUDA, Triton, Nsight, InfiniBand, RDMA, RoCE, H100/H200/B200 clusters, or storage/network bottlenecks. Experience supporting SFT, DPO, PPO, GRPO, RLAIF, reward model serving, rollout generation, or agent training workloads. Open-source contributions to distributed training, RL infrastructure, PyTorch, Ray, Megatron, DeepSpeed, or related systems. Strong Python and PyTorch engineering skills, Hands-on experience with distributed model training, large-scale ML systems, or GPU cluster workloads, Practical understanding of transformer training bottlenecks, memory pressure, gradient/optimizer state, communication overhead, and checkpointing, Experience debugging production or research training jobs across multiple GPUs or nodes, Ability to reason quantitatively about throughput, utilization, memory, reliability, cost, and research velocity, Strong communication skills and ability to collaborate with researchers, ML engineers, platform engineers, and leadership #J-18808-Ljbffr
$300k - $400k
...possible. About the Role You will own the systems layer that makes our frontier model... ...operations Profiling and benchmarking distributed ML systems to identify and eliminate... ...team of the world’s best — the scientists, engineers, and problem-solvers who don’t just follow...SuggestedVisa sponsorshipFlexible hoursShift work- ...About the Role We’re looking for an Applied ML Engineer to design, evaluate, and scale recommendation and ranking systems that power how content, ads, and interactive experiences are selected and surfaced in real time. This role focuses on decision-making systems, with...SuggestedFull time
- ...A growing AI technology startup is seeking an ML Engineer to design and deploy production-grade systems. The role involves using Python and collaborating with teams to optimize customer interactions through advanced AI applications. Candidates should have a degree in Computer...Suggested
$120k - $140k
...Deep Learning / NLP Engineer Join Avoma and work on some of the most challenging NLP problems in our mission to make every meeting... ...looking for a Deep Learning / NLP Engineer to improve and build our systems to extract key insights and topics from conversations. For our...Suggested- ...ML Engineer Palo Alto, California, United States About the Job Our client is a rapidly growing Tier 1 VC backed startup based... ...long-term growth trajectory in the evolving world of intelligent systems. Location New York, NY Work Type Full Time...SuggestedFull time
- ...The Mission: As a Senior Machine Learning Engineer, you will be responsible for building machine learning models/systems and innovative web applications that deliver the power... ...field Experience with building and evolving ML Training and Inferencing systems at significant...Local area
$138.5k - $225.5k
...value across the platform. About the role We're hiring a ML Engineer as one of the founding engineers on Intelligence Org. You'll... ...Responsibilities Design, train, evaluate, and ship ML systems that power governance and security capabilities, starting with...Full timeRemote workHome officeVisa sponsorshipShift work$208k - $244k
...Join to apply for the Staff ML Engineer role at Grindr Join to apply for the Staff ML Engineer role at Grindr Get AI-powered advice on... ...our long term ML strategy. Recommendations That Reshape: Build systems that match millions to their next big moment, adapting to a range...Full timeCasual workWork at officeImmediate startFlexible hours- Job Title: ML Engineer What You Will Own End‑to‑End ML Lifecycle across real products: data ingestion, feature design, model selection,... ..., monitoring and iteration. No handoffs. Production‑grade ML systems built with PyTorch or TensorFlow, focusing on latency, reliability...
- .... Our platform leverages cutting-edge generative AI to assist engineers in RTL design, simulation, and verification, dramatically accelerating... ...shipped 16B chips. Position Overview We are seeking an ML Systems Engineer to optimize the performance and efficiency of large...Full time
$190k - $234k
...Staff ML Engineer Palo Alto, CA About Typeface We help the world's biggest brands move from brief to fully personalized campaigns... ...you will be responsible for building machine learning models/systems and innovative web applications that deliver the power of Generative...Work at officeLocal areaFlexible hours3 days per week$230k - $260k
...Principal ML Engineer Palo Alto, CA About Typeface We help the world's biggest brands move from brief to fully personalized campaigns... ...AI at Typeface. You will lead the design of large-scale ML systems and shared platforms that power all generative capabilities...Work at officeImmediate startFlexible hours3 days per week$152k - $261k
...services for landing, tracking & attribution. ML opportunities across these portfolios... ...?’ What You Will Do Drive end-to-end Engineering and Machine learning methods that have a... ...maintaining highly available, distributed systems Preferred Qualifications Experience in...Full timeTemporary workWork experience placementFlexible hours- ...the clean energy transition. About The Role As an AI/ML Engineer at Powerline, you will be instrumental in developing, optimizing... ...and prediction. ~ Familiarity with MLOps and building AI/ML systems end to end. ~ Strong communication skills and ability to...Full time
- ...We bridge this exact gap by applying deep systems programming, software-defined networking,... ...at UT Austin and world-renowned ML systems researcher with a pedigree spanning... ...Seniority ~5+ years of production experience engineering ML systems, OR a PhD from a top-tier...Shift work
$283.64k - $425.96k
...leading investors.About the RoleNuro is looking for a Head of Systems Engineering to own the systems backbone that enables the Nuro Driver to... ...vehicle integration) and the validation challenges specific to ML-based and end-to-end AI systems in real-world environments....Odd jobPart timeImmediate startFlexible hours$170k - $190k
...relentless focus on outcomes. ASAPP’s AI Engineering team is seeking an enterprising, talented... ...are seeking a highly experienced Lead AI/ML Engineer to join our Core GenerativeAgent... ...building, and deploying cutting-edge AI systems that power mission-critical enterprise applications...Full time$205k - $235k
...Year. As a Senior Machine Learning Engineer , you will be responsible for building machine learning models/systems and innovative web applications that deliver the... ...~ Experience with building and evolving ML Training and Inferencing systems at significant...Full timeWork at officeLocal areaFlexible hours3 days per week- ...business and operations to provide insight and actionable items at real-time. ~ We are looking for streaming, data engineer ~ Understand distributed systems architecture, design and trade-off. ~ Design and develop ETL pipelines with a wide range of technologies....Permanent employmentWork experience placementLocal area3 days per week
$175k - $275k
...raw data, curated datasets, or full-cycle data engineering, Abaka AI provides the foundation for building high-performance AI systems. About the Role We’re hiring... ...of experience in applied machine learning or ML engineering, with a demonstrated ability to deliver...Full timeImmediate startFlexible hours$160k - $225k
...capital will be used to expand our product and engineering teams, bringing our vision of intelligent... ...them to reason, to the scalable serving systems that deliver their intelligence to our... .... You'll build the robust pipelines and ML serving systems that fuel our agents with...Full time$195k - $230k
...information powered by advanced AI, recommendation systems, and adtech. Recognized by Fast... ...looking for a Senior Machine Learning Engineer to help evolve our large-scale... ...Experience working with large-scale data and ML systems (e.g., Spark, distributed training...Full timeLocal areaWork from home- ...starting out with understanding and building hardware; electronics systems and semiconductors where AI can design and create beyond... ...four US presidents. What we're Looking For Strong AI/ML engineering skills from top tier CS, EECS, Math and Physics programs....Full time
- ...Together, we can make a meaningful impact. See more about our culture on . About The Job Mistral AI is seeking a Applied AI Engineer to facilitate the adoption of its products among customers and collaborate with them to address complex technical challenges....Full timeWork at officeVisa sponsorship
$185k - $254k
Who We Are Applied Materials is a global leader in materials engineering solutions used to produce virtually every new chip and advanced... ...focus during later stages of completion. Ensures that all systems engineering projects and programs assigned are completed in accordance...Full timeRelocation- ...About the Role We are seeking a Senior Data / AI / ML Software Engineer with 7+ years of experience building data-intensive systems. This role is ideal for someone who enjoys designing and improving core platform components at the intersection of software engineering,...Full timeContract workInternship
$78 - $83 per hour
...Overview Senior Systems Process Engineer — Palo Alto, CA — Duration: 5 month contract — Pay: $78-83/hr This Senior Systems Process Engineer role is responsible for defining and implementing end-to-end processes, methods, and tooling for product configuration, change,...Contract work$190k - $300k
...technology for aviation that will save lives. Automated aviation systems will enable a future where air transportation is safer, more... ...people - move around the planet. We are a team of mission-driven engineers with experience across aerospace, robotics and self-driving...Permanent employmentRemote work$196k - $248k
...autonomously driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states. Waymo's Systems Engineering team works together to blend software and hardware systems in groundbreaking new ways. We set the high performance standards...Full timeRemote work- ...Job Overview Senior Systems Process Engineer Duration: 6+ Months Job Description Create process, methods and tool architecture for our product configuration, change, variant, release and sign-off management in a Software Defined Vehicle (SDV) context. Ensure a seamless...Contract work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Systems Engineer. Be the first to apply!
- machine learning engineer Palo Alto, CA
- computer vision machine learning engineer Palo Alto, CA
- systems engineer Palo Alto, CA
- system performance engineer Palo Alto, CA
- software system engineer Palo Alto, CA
- computer system validation engineer Palo Alto, CA
- healthcare systems engineer Palo Alto, CA
- distributed systems engineer Palo Alto, CA
- sr systems engineer Palo Alto, CA
- senior windows systems engineer Palo Alto, CA



