Senior Machine Learning Engineer, LLM Inference Optimization
Nebius
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role Nebius Token Factory is building fast, reliable, and cost-efficient inference services for frontier models. As a Senior Machine Learning Engineer on our Applied AI team, you will own model and endpoint optimization from model artifacts through production deployment. Your work will span model internals, inference engines, serving architecture, and benchmarking, with a focus on improving latency, throughput, memory efficiency, GPU utilization, and cost per token while maintaining model quality and reliability. This is a hands-on role in which you will work on complex optimization projects, diagnose difficult serving problems, and deliver measurable improvements in production. Working closely with kernel and platform engineers, you will evaluate serving configurations, resolve performance and quality regressions, and optimize inference for real-world workloads, supported by reproducible benchmarks and safe production rollouts. Your responsibilities : - Own optimization work for specific model families, customer endpoints, or serving backends. - Run engine comparisons and recommend practical serving configurations for specific workloads. - Debug model quality or performance regressions during production rollouts.
- Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token. - Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or similar systems. - Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery. - Implement or integrate speculative decoding, draft-model approaches, KV -cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving. - Build reproducible benchmark harnesses for TTFT , TPOT , tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token. - Partner with GPU kernel engineers and platform engineers to diagnose bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers. - Write clear design docs, performance reports, rollout plans, and customer-facing technical explanations. Must-haves : - Strong Python and PyTorch engineering skills. - Hands-on experience deploying or optimizing LLM, VLM , or high-throughput transformer inference systems. - Practical knowledge of at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or equivalent internal systems. - Strong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving. - Ability to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs. - Strong communication skills and ability to collaborate with research, kernel, infrastructure, product, and customer teams.
- ...We are looking for a full-time Machine Learning Engineer, with deep knowledge and strong enthusiasm... ...and accelerating model training/inference. Our mission is to solve the autonomous... ...systems Knowledge of model optimization including quantization, pruning, etc...SeniorFull timeContract workLocal area
- ...environment that values trust, proactivity, and autonomy? Are our Engineering principles com/pennylane-engineering/our-engineering-... ...Embedded AI team owns the first layer. We build the specialized machine learning systems behind Copilot and Autopilot: invoice parsing,...SeniorFull timeRemote work
- ...software, AI, cryptography, mobile engineering, and global operations. Our... ...named to the [Time AI 100]( Learn more about the newest... ...Tools for Humanity owns the machine learning systems behind the... ...selection, loss functions, optimization, hyperparameter tuning, and...SeniorFull time
- ...humble in our pursuit of excellence, constantly learning and evolving as we pave the way for a... ...career! The role As a Staff ML Performance Engineer, you’ll play a key role in high-impact projects, optimising ML inference for edge accelerators and GPUs. The focus of...SuggestedFull timeWork at officeWork from home
- ...infrastructure. Built by engineers, for engineers. From... ...GPU orchestration to inference optimization, we own the hard... ...seek an experienced Senior ML Solutions Architect... ...workflows, build customized LLM-based solutions and... ...and reinforcement learning fine-tuning to maximize...SeniorFull timeRemote work
- ...pursuit of excellence, constantly learning and evolving as we pave the... ...Data Scientist supporting AI engineers, you will partner with one or... ...in model training and inference leading to bottlenecks in functionality... ...Practical experience with machine learning (e.g. PyTorch)....SeniorFull timeWork at officeWork from home2 days per week
- ...in our pursuit of excellence, constantly learning and evolving as we pave the way for a... ...that defines your career! ** ️ About our Engineering Teams** Wayve’s ADAS engineering teams... ...human driven vehicles to intelligent machines. Our ambition is to make autonomy universal...InternshipWork at officeRemote workVisa sponsorshipRelocation packageFlexible hoursShift work
- ...as CTO / COO (web2/web3), 3 years in AI/LLM • Serial-entrepreneur: MTK Digital (exited... ...results to identify areas of improvement, optimize campaigns for better performance, and find... ..., and strategy Humble - willing to learn, open to feedback Adaptable - comfortable...SeniorFull timeRemote workNight shift
- ..., we’re scaling fast - and we’re looking for world-class Machine Learning Engineers to help us keep pushing the boundaries of biometric technology... ...the world. What You’ll Own & Drive Innovate & Optimize - Develop and refine state-of-the-art deep learning models...Full timeFlexible hours
- ...large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute,... ...software and AI R&D. The role We're looking for a senior support engineer who can handle difficult...SeniorFull time
- ...trading firm which uses state-of-the-art machine learning technology to produce price forecasts... ...The Role XTX is seeking an experienced engineer to support our ML Performance and AI... ...the performance of XTX's training and inference platforms. The remit is wide, and you should...Full timeWork at officeWorldwide
- ...ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems... ...Nebius is looking for Mid and Senior Software Developer with a... ...compensation - Career growth and learning opportunities - Flexibility...SeniorFull time
- ...large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute,... ...of our broader HR technology roadmap. As a Senior HR Systems Analyst, you will partner closely...SeniorFull time
- ...Software Development Engineer AI/ML This position provides a seasoned AI/ML engineer... ...support. Your experience with applied machine learning, Generative AI, Large Language Models,... ...for designing, implementing, and optimizing machine learning models, AI applications...
- ...excellence, constantly learning and evolving as we pave... ...role As a software engineer for Wayve’s Simulation... ...cutting edge developments in machine learning to represent... ..., implementation, and optimization of large-scale machine learning inference systems running in...SeniorFull timeWork at officeWork from home
- ...greatest potential. Title and Summary Senior Information Security Engineer Who is Mastercard? Mastercard... .... • Provide and recommend optimal solutions to meet security and regulatory... ...web service, DevOps, cloud, GenAI, LLM models, and CI/CD efforts All About...SeniorFull timeWork experience placementWorldwide
- ...ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems... ...The role: We are hiring a Senior Technical Project Manager for... ...compensation - Career growth and learning opportunities - Flexibility...SeniorFull time
- ...you are only a 75% match. Skills can be learned, diversity cannot. __ Locations... ...as secondary projects. As a Staff Machine Learning Engineer, you will hold a hands-on technical position... ...audit trails, per-layer attribution, llm-driven analyses, or defending a model'...Full timeWork at officeRemote workRelocationHome officeFlexible hours2 days per week1 day per week
- ...oversee operations and customer support. This role involves ensuring service quality, optimizing processes, and coaching service teams. Ideal candidates will have a background in engineering or business, along with 3-5 years of relevant experience. The position offers home...SeniorContract workHome office
- ...large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute,... ...for AI agents: a system designed for machines, not humans. We are building agentic search...Full time
- ...Looking For We're looking for a Senior Data Scientist to join ARQ's... ..., Product, Finance, Data Engineering, and Engineering teams to... ...~5+ years in Data Science, Machine Learning, Applied Statistics, Analytics... ..., experiment design, causal inference, or incrementality measurement...SeniorRemote jobFull timeWork at office3 days per week
- ...founder & CTO] • 10+ years as CTO / COO (web2/web3), 3 years in AI/LLM • Serial-entrepreneur: MTK Digital (exited / 0-$20m revenue) and... ...around specs and handoffs, waiting on a designer's mockup, an engineer's sprint slot, a data analyst's report. This role is built differently...SeniorFull timeRemote work
- ...driverless cars? Join our Waymo’s Supply Optimization Infrastructure team, and help us design... ...! We are looking for a passionate Senior SWE who wants to help build the systems... ...this hybrid role, you will report to an Engineering Manager. You Will Build and evolve...SeniorFull time
- ...resolve incidents faster, and optimize telemetry at scale. Built on... ...CapitalG, and Lead Edge Capital. Learn more at grafana. com and... ...as individuals and as a team. Engineers are empowered to make decisions... .... We’re looking for a Senior Software (Database) Engineer...SeniorFull timeRemote work
- ...without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq...Full time
- ...debate, experimentation, and continuous learning, while seeking people with... ...also build VPN, tune-up, privacy, and optimization products for macOS, as well as antivirus solutions for Linux. As a Senior Principal Software Engineer, you’ll set the technical direction...SeniorRemote jobFull timeFlexible hours
- ...Panopto, we are the most customer-centric learning technology company in the world. As... ...team, we are seeking an experienced Senior Engineering Manager who brings strong people leadership... ...without lowering code quality, and optimize engineering efficiency across complex...SeniorFull time
- ...Overview Join to apply for the Senior software engineer role at Skillvue . Get AI-powered... ...handling large datasets and optimizing performance in data pipelines ~ Understanding... ..., and other stakeholders to deliver machine learning solutions to production. Capable...SeniorFull timeWork at officeRemote workShift work
- ...processes, guide, and mentor other more junior engineers within the team, take active part in the... ...challenges with respect, transparency, optimism, and confidence to succeed together. You are creative, with determination to learn continuously, break new ground, and push...SeniorFull timeRemote work
- ...businesses and governments realize their greatest potential. Title and Summary Senior ML Platform Engineer Our Mission At Mastercard Identity Verification, we build data, machine learning, and platform capabilities that help customers make safer decisions in...SeniorFull timeWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Machine Learning Engineer, LLM Inference Optimization. Be the first to apply!


