Senior Machine Learning Engineer, LLM Inference Optimization
Nebius
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role Nebius Token Factory is building fast, reliable, and cost-efficient inference services for frontier models. As a Senior Machine Learning Engineer on our Applied AI team, you will own model and endpoint optimization from model artifacts through production deployment. Your work will span model internals, inference engines, serving architecture, and benchmarking, with a focus on improving latency, throughput, memory efficiency, GPU utilization, and cost per token while maintaining model quality and reliability. This is a hands-on role in which you will work on complex optimization projects, diagnose difficult serving problems, and deliver measurable improvements in production. Working closely with kernel and platform engineers, you will evaluate serving configurations, resolve performance and quality regressions, and optimize inference for real-world workloads, supported by reproducible benchmarks and safe production rollouts. Your responsibilities : - Own optimization work for specific model families, customer endpoints, or serving backends. - Run engine comparisons and recommend practical serving configurations for specific workloads. - Debug model quality or performance regressions during production rollouts.
- Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token. - Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or similar systems. - Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery. - Implement or integrate speculative decoding, draft-model approaches, KV -cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving. - Build reproducible benchmark harnesses for TTFT , TPOT , tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token. - Partner with GPU kernel engineers and platform engineers to diagnose bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers. - Write clear design docs, performance reports, rollout plans, and customer-facing technical explanations. Must-haves : - Strong Python and PyTorch engineering skills. - Hands-on experience deploying or optimizing LLM, VLM , or high-throughput transformer inference systems. - Practical knowledge of at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or equivalent internal systems. - Strong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving. - Ability to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs. - Strong communication skills and ability to collaborate with research, kernel, infrastructure, product, and customer teams.
$298k - $368k
...the system which learns the spatial-temporal... ...teams on the optimization and integration into... ...sensors, enabling engineers like you to (1) develop... ...: Design VLM/LLM model architecture... ...of experience in Machine Learning, with a... ...latency on-device inference techniques and a deep...SeniorFull timeRemote work- ...Community You Will Join: Machine Learning and Artificial... ...services and tools including LLM fine-tuning, alignment and optimization, RAG/Search, LLM... ...principal machine learning engineer, you will be responsible... ...optimizing models and inference run-time ~ Post-training...SuggestedRemote jobFull timeCasual workLive inWork at office
- ...explore, create, play, learn, and connect with friends... ...a billion people with optimism and civility, and... ...experiences for everyone. Our engine’s resource management... ...the application of machine learning in real-time engine... ...Design ML models that infer player and interaction...SeniorFull time
$204k - $259k
...builds the system which learns the spatial-temporal... ...downstream teams on the optimization and integration into... ...of sensors, enabling engineers like you to (1) develop... ...years of experience in Machine Learning, with a focus... ...model development (LLM, VLM, or similar foundation...SeniorFull timeRemote work- ...About the Role We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika'... ...at scale. You will design and optimize inference pipelines, implement state... ..., attention acceleration, and deep learning compiler stacks. GPU &...SuggestedFull timeWork at office3 days per week
$213k - $263k
...Waymo AI Foundations team is to develop machine learning solutions addressing open problems in... ...demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust... ...to smaller real-time models. Explore LLM/VLM distillation recipes to maximally...SeniorFull timeRemote work$213k - $263k
...Drive cross-functional collaboration to engineer robust, high-reliability training... ...Have: ~ BS or MS in Computer Vision, Machine Learning, Robotics, or a related field. ~4+ years... .... Hands-on experience managing and optimizing large-scale teacher-student training...SeniorFull timeRemote work$213k - $263k
...states. The ML Optimization team at Waymo... ...lifecycle of the machine learning workflow, including... ...are looking for engineers with ML software... ...Waymo onboard ML inference engine for Waymo... ...will report to the Senior Manager of Runtime... ...building or scaling LLM serving systems,...SeniorFull timeRemote work$195k - $230k
...We are looking for a Senior Machine Learning Engineer to help evolve our large... ...systems and apply AI / LLM technologies to real-world... ...ranking, and multi-objective optimization to balance engagement, retention... ...offline training → online inference → A/B experimentation →...SeniorFull timeLocal areaWork from home- ...Position Summary The Machine Learning Engineer will be responsible for the... ...appropriately for the chosen LLM and training pipeline... ...size, and training epochs to optimize model performance. Integration... ...pipelines for model training, inference, and deployment....SeniorFull timeH1bRemote workFlexible hours
- ...are seeking an experienced Senior Machine Learning Engineer to join our AI/ML team and... ...fine-tuning, preference optimization, and other post-training techniques... ...model-serving and inference infrastructure for open-weight... .... ~ Experience with LLM fine-tuning and post-training...SeniorRemote jobFull timeWork at office
- ...everyone else has simply learned to live with. We... ...You’ll help define how machine learning models run... ...ll work with systems engineers, product teams, hardware... ...combines applied ML, inference optimization, evaluation, and... ...SGLang, vLLM, TensorRT-LLM, ONNX Runtime, Triton...SeniorFull timeLocal area
$235.2k - $294k
...The goal of a Senior Machine Learning Engineer at Scale is to own how we apply generative... ...evaluation benchmarks, LLM judges, and verifiers -... ...infrastructure to automate and optimize our ML services Work... ...Geospatial or GEOINT experience Inference optimization experience...SeniorFull time$188.5k - $282.7k
...Semantic AI Governance Engine, which is the first... ...At its core, SAGE is "LLM-as-judge" applied to... ...fine-tuning, preference optimization (DPO/RLAIF), and... ...Performance Model Serving and Inference Infrastructure (25%... ...in Computer Science, Machine Learning, Computer Engineering...SeniorPermanent employmentFull timeLocal area$150k - $180k
.... About the Role: As our Senior Machine Learning Engineer, you’ll own the intelligence layer... ..., put physics-based, ML, and LLM-powered models into production,... ...with LLM cost/latency optimization (prompt caching, batch inference) and model governance (managing...SeniorFull timeWork at officeImmediate startShift work- ...motivated and experienced Machine Learning Engineer to join our AI &... ...production engineering (inference systems, integration, and optimization). Responsibilities... ...engineering techniques and LLM frameworks ~... ...celebrates your commitment and seniority (including paid...SeniorRemote jobFull timeTemporary work
$175k - $200k
...Senior Machine Learning Engineer Truveta is the world’s first health provider led... ...expertise in applied AI, model optimization, and agentic intelligence... ...a deep understanding of LLM fundamentals —... ...efficiency, interpretability, and inference performance. Think and...SeniorFull timeFor contractorsVisa sponsorshipWork visaFlexible hours$260k - $330k
...have: Experience building large-scale prediction or optimization systemsPubMatic is the leading AI-powered ad tech company... ...environments.About the Role:We are looking for a Senior Principal Machine Learning Engineer to help build the next generation of performance optimization...SeniorWork at officeRemote work$150.75k - $241.2k
...are seeking a seasoned Machine Learning Engineer to join a new team... ...reasoning systems. As a senior engineer on this team... ...infrastructure, the inference and serving stack,... ...multimodal workloads, optimizing latency, throughput,... ...agentic systems and LLM tool-use orchestration...SeniorWork experience placementWork at officeRemote work$232k - $310k
...and high standards. Our engineers, product leaders, and... ...boundaries of applied machine learning. We work with massive datasets... ...environments and optimize system performance.Contribute... ...(QLORA, DPO) and inference optimization (vLLM, TensorRT-LLM).Desired Skills & Experience...Work experience placementWork at officeRemote workFlexible hours3 days per week$172k - $229k
...mining framework, is the engine that powers this discovery. As a Senior Machine Learning Engineer on the Data... ...learning, retrieval optimization, and reasoning systems... ...drastically reducing inference latency and memory footprint... ...of-thought models, or LLM-based planning....SeniorFull timeWork at officeRemote work- ...Senior Machine Learning Engineer Department: Engineering Employment Type: Full... ...production Design systems that infer structured attributes and... ...indexing strategies Optimize models for inference latency... ...Experience building LLM or VLM pipelines and the evaluation...SeniorFull timeRemote workFlexible hours
$144k - $233.1k
...Senior Machine Learning Engineer Since 2003, Entrata has evolved from a visionary, student-led startup... ..., and task performance. Optimize model inference, serving, and deployment for performance... .... ~ Familiarity with modern LLM tooling, model serving, and inference...SeniorLocal areaRemote workWorldwideFlexible hours$130k - $160k
...has an opportunity for a Senior Machine Learning Engineer. Candid's Data Science and... ..., observability, and inference performance as that ownership... ...familiarity with serving optimization techniques such as quantization... ...or another managed LLM service (Anthropic API, Azure...SeniorTemporary workSummer workLocal areaRemote workMonday to FridayFlexible hours- ...and continuously learn and adapt. Moveworks... ...’ Reasoning Engine and natural language... ...are looking for a Machine Learning Engineer... ...building and serving LLM’s at Moveworks.... ...critical in building, optimizing and scaling end-to... ...training and inference pipeline for large...SeniorFull timeWork at officeRemote workFlexible hours
$201.3k - $352.3k
...entrepriseIt all started when engineer Fred Luddy wrote code... ...Emerging tech is a small senior group inside AI... ...or retrieval. Exposure to LLM fine-tuning or inference optimization in productionWhy join us... ...assigned work location. Learn more here. To determine eligibility...SeniorWork experience placementWork at officeImmediate startRemote workFlexible hours$159.18k - $295.62k
...a career- we're hiring! Senior Machine Learning Engineer Team: Data & Audience Platform... ..., model training and optimization, and the ML infrastructure... ...scalable feature and inference pipelines on Databricks (PySpark... ...search; evaluate LLM- based approaches for metadata...SeniorTemporary workLocal area$240k - $249.5k
...OpportunityGrubhub is looking for a Senior Staff Machine Learning Engineer to help lead the machine... ...the evolution of our optimization objective from short-term... ...our runtime environment: LLM-driven query and intent... ...start, and real-time inference. Assess rigorously what actually...SeniorFull timeTemporary workWork at officeFlexible hours3 days per week$141k - $249k
...with autonomy and algorithm engineers to scale safe self-driving systems... ...as TensorRT and modelopt to optimize the models running on the... ...benchmark new CUDA kernels for inference. Comprehensively profile... ...Rust. Experience in deep learning frameworks such as PyTorch....SeniorWork at officeWork from homeFlexible hours- ...Senior Machine Learning Engineer - Data Science & Analytics Contract Length: 6-18... ...integration Real-time inference Batch processing... ...and ML data pipelines Optimize latency, scalability, and... ...Experience supporting AI/LLM-enabled applications Team...SeniorContract workRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Machine Learning Engineer, LLM Inference Optimization. Be the first to apply!
- junior machine learning research engineer Remote
- data scientist machine learning engineer Remote
- graduate machine learning engineer Remote
- junior machine learning engineer Remote
- computer vision machine learning engineer Remote
- machine learning software engineer Remote
- machine learning ai engineer Remote
- senior ml engineer Remote
- machine learning engineer Remote
- ai ml engineer Remote


