Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Machine Learning Engineer, LLM Inference Optimization

Full-time

Nebius

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role Nebius Token Factory is building fast, reliable, and cost-efficient inference services for frontier models. As a Senior Machine Learning Engineer on our Applied AI team, you will own model and endpoint optimization from model artifacts through production deployment. Your work will span model internals, inference engines, serving architecture, and benchmarking, with a focus on improving latency, throughput, memory efficiency, GPU utilization, and cost per token while maintaining model quality and reliability. This is a hands-on role in which you will work on complex optimization projects, diagnose difficult serving problems, and deliver measurable improvements in production. Working closely with kernel and platform engineers, you will evaluate serving configurations, resolve performance and quality regressions, and optimize inference for real-world workloads, supported by reproducible benchmarks and safe production rollouts. Your responsibilities : - Own optimization work for specific model families, customer endpoints, or serving backends. - Run engine comparisons and recommend practical serving configurations for specific workloads. - Debug model quality or performance regressions during production rollouts.

- Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token. - Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or similar systems. - Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery. - Implement or integrate speculative decoding, draft-model approaches, KV -cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving. - Build reproducible benchmark harnesses for TTFT , TPOT , tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token. - Partner with GPU kernel engineers and platform engineers to diagnose bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers. - Write clear design docs, performance reports, rollout plans, and customer-facing technical explanations. Must-haves : - Strong Python and PyTorch engineering skills. - Hands-on experience deploying or optimizing LLM, VLM , or high-throughput transformer inference systems. - Practical knowledge of at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or equivalent internal systems. - Strong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving. - Ability to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs. - Strong communication skills and ability to collaborate with research, kernel, infrastructure, product, and customer teams.

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Senior Machine Learning Engineer, LLM Inference Optimization in Europe vacancy
  •  ...We are looking for a full-time Machine Learning Engineer, with deep knowledge and strong enthusiasm...  ...and accelerating model training/inference. Our mission is to solve the autonomous...  ...systems Knowledge of model optimization including quantization, pruning, etc... 
    Senior
    Full time
    Contract work
    Local area

    Rivian

    United Kingdom
    13 days ago
  •  ...environment that values trust, proactivity, and autonomy? Are our Engineering principles com/pennylane-engineering/our-engineering-...  ...Embedded AI team owns the first layer. We build the specialized machine learning systems behind Copilot and Autopilot: invoice parsing,... 
    Senior
    Full time
    Remote work

    Pennylane

    France
    19 days ago
  •  ...software, AI, cryptography, mobile engineering, and global operations. Our...  ...named to the [Time AI 100]( Learn more about the newest...  ...Tools for Humanity owns the machine learning systems behind the...  ...selection, loss functions, optimization, hyperparameter tuning, and... 
    Senior
    Full time

    Tools for Humanity

    Germany
    a month ago
  •  ...humble in our pursuit of excellence, constantly learning and evolving as we pave the way for a...  ...career! The role As a Staff ML Performance Engineer, you’ll play a key role in high-impact projects, optimising ML inference for edge accelerators and GPUs. The focus of... 
    Suggested
    Full time
    Work at office
    Work from home

    Wayve

    United Kingdom
    19 days ago
  •  ...infrastructure. Built by engineers, for engineers. From...  ...GPU orchestration to inference optimization, we own the hard...  ...seek an experienced Senior ML Solutions Architect...  ...workflows, build customized LLM-based solutions and...  ...and reinforcement learning fine-tuning to maximize... 
    Senior
    Full time
    Remote work

    Nebius

    Europe
    19 days ago
  •  ...pursuit of excellence, constantly learning and evolving as we pave the...  ...Data Scientist supporting AI engineers, you will partner with one or...  ...in model training and inference leading to bottlenecks in functionality...  ...Practical experience with machine learning (e.g. PyTorch).... 
    Senior
    Full time
    Work at office
    Work from home
    2 days per week

    Wayve

    United Kingdom
    8 days ago
  •  ...in our pursuit of excellence, constantly learning and evolving as we pave the way for a...  ...that defines your career! ** ️ About our Engineering Teams** Wayve’s ADAS engineering teams...  ...human driven vehicles to intelligent machines. Our ambition is to make autonomy universal... 
    Internship
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours
    Shift work

    Wayve

    United Kingdom
    22 days ago
  •  ...as CTO / COO (web2/web3), 3 years in AI/LLM • Serial-entrepreneur: MTK Digital (exited...  ...results to identify areas of improvement, optimize campaigns for better performance, and find...  ..., and strategy Humble - willing to learn, open to feedback Adaptable - comfortable... 
    Senior
    Full time
    Remote work
    Night shift

    Everai

    Europe
    10 days ago
  •  ..., we’re scaling fast - and we’re looking for world-class Machine Learning Engineers to help us keep pushing the boundaries of biometric technology...  ...the world. What You’ll Own & Drive Innovate & Optimize - Develop and refine state-of-the-art deep learning models... 
    Full time
    Flexible hours

    Incode

    Kingdom of Spain
    a month ago
  •  ...large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute,...  ...software and AI R&D. The role We're looking for a senior support engineer who can handle difficult... 
    Senior
    Full time

    Nebius

    Europe
    19 days ago
  •  ...trading firm which uses state-of-the-art machine learning technology to produce price forecasts...  ...The Role XTX is seeking an experienced engineer to support our ML Performance and AI...  ...the performance of XTX's training and inference platforms. The remit is wide, and you should... 
    Full time
    Work at office
    Worldwide

    XTX Markets

    United Kingdom
    a month ago
  •  ...ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems...  ...Nebius is looking for Mid and Senior Software Developer with a...  ...compensation - Career growth and learning opportunities - Flexibility... 
    Senior
    Full time

    Nebius

    Europe
    19 days ago
  •  ...large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute,...  ...of our broader HR technology roadmap. As a Senior HR Systems Analyst, you will partner closely... 
    Senior
    Full time

    Nebius

    Europe
    19 days ago
  •  ...Software Development Engineer AI/ML This position provides a seasoned AI/ML engineer...  ...support. Your experience with applied machine learning, Generative AI, Large Language Models,...  ...for designing, implementing, and optimizing machine learning models, AI applications... 

    NTT DATA, Inc.

    Croatia
    10 days ago
  •  ...excellence, constantly learning and evolving as we pave...  ...role As a software engineer for Wayve’s Simulation...  ...cutting edge developments in machine learning to represent...  ..., implementation, and optimization of large-scale machine learning inference systems running in... 
    Senior
    Full time
    Work at office
    Work from home

    Wayve

    United Kingdom
    a month ago
  •  ...greatest potential. Title and Summary Senior Information Security Engineer Who is Mastercard? Mastercard...  .... • Provide and recommend optimal solutions to meet security and regulatory...  ...web service, DevOps, cloud, GenAI, LLM models, and CI/CD efforts All About... 
    Senior
    Full time
    Work experience placement
    Worldwide

    Mastercard

    Ireland
    a month ago
  •  ...ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems...  ...The role: We are hiring a Senior Technical Project Manager for...  ...compensation - Career growth and learning opportunities - Flexibility... 
    Senior
    Full time

    Nebius

    Europe
    19 days ago
  •  ...you are only a 75% match. Skills can be learned, diversity cannot. __ Locations...  ...as secondary projects. As a Staff Machine Learning Engineer, you will hold a hands-on technical position...  ...audit trails, per-layer attribution, llm-driven analyses, or defending a model'... 
    Full time
    Work at office
    Remote work
    Relocation
    Home office
    Flexible hours
    2 days per week
    1 day per week

    Moonpay

    United Kingdom
    a month ago
  •  ...oversee operations and customer support. This role involves ensuring service quality, optimizing processes, and coaching service teams. Ideal candidates will have a background in engineering or business, along with 3-5 years of relevant experience. The position offers home... 
    Senior
    Contract work
    Home office

    EPTA GROUP

    Italian Republic
    3 days ago
  •  ...large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute,...  ...for AI agents: a system designed for machines, not humans. We are building agentic search... 
    Full time

    Nebius

    Europe
    19 days ago
  •  ...Looking For We're looking for a Senior Data Scientist to join ARQ's...  ..., Product, Finance, Data Engineering, and Engineering teams to...  ...~5+ years in Data Science, Machine Learning, Applied Statistics, Analytics...  ..., experiment design, causal inference, or incrementality measurement... 
    Senior
    Remote job
    Full time
    Work at office
    3 days per week

    ARQ ( Prev DolarApp )

    United Kingdom
    a month ago
  •  ...founder & CTO] • 10+ years as CTO / COO (web2/web3), 3 years in AI/LLM • Serial-entrepreneur: MTK Digital (exited / 0-$20m revenue) and...  ...around specs and handoffs, waiting on a designer's mockup, an engineer's sprint slot, a data analyst's report. This role is built differently... 
    Senior
    Full time
    Remote work

    Everai

    Europe
    10 days ago
  •  ...driverless cars? Join our Waymo’s Supply Optimization Infrastructure team, and help us design...  ...! We are looking for a passionate Senior SWE who wants to help build the systems...  ...this hybrid role, you will report to an Engineering Manager. You Will Build and evolve... 
    Senior
    Full time

    Waymo

    Poland
    8 days ago
  •  ...resolve incidents faster, and optimize telemetry at scale. Built on...  ...CapitalG, and Lead Edge Capital. Learn more at grafana. com and...  ...as individuals and as a team. Engineers are empowered to make decisions...  .... We’re looking for a Senior Software (Database) Engineer... 
    Senior
    Full time
    Remote work

    Grafanalabs

    Ireland
    7 days ago
  •  ...without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq... 
    Full time

    Nebius

    Europe
    19 days ago
  •  ...debate, experimentation, and continuous learning, while seeking people with...  ...also build VPN, tune-up, privacy, and optimization products for macOS, as well as antivirus solutions for Linux. As a Senior Principal Software Engineer, you’ll set the technical direction... 
    Senior
    Remote job
    Full time
    Flexible hours

    MoneyLion

    Czech Republic
    a month ago
  •  ...Panopto, we are the most customer-centric learning technology company in the world. As...  ...team, we are seeking an experienced Senior Engineering Manager who brings strong people leadership...  ...without lowering code quality, and optimize engineering efficiency across complex... 
    Senior
    Full time

    Panopto

    Europe
    18 days ago
  •  ...Overview Join to apply for the Senior software engineer role at Skillvue . Get AI-powered...  ...handling large datasets and optimizing performance in data pipelines ~ Understanding...  ..., and other stakeholders to deliver machine learning solutions to production. Capable... 
    Senior
    Full time
    Work at office
    Remote work
    Shift work

    Skillvue

    Italian Republic
    5 days ago
  •  ...processes, guide, and mentor other more junior engineers within the team, take active part in the...  ...challenges with respect, transparency, optimism, and confidence to succeed together. You are creative, with determination to learn continuously, break new ground, and push... 
    Senior
    Full time
    Remote work

    Tesla

    United Kingdom
    13 days ago
  •  ...businesses and governments realize their greatest potential. Title and Summary Senior ML Platform Engineer Our Mission At Mastercard Identity Verification, we build data, machine learning, and platform capabilities that help customers make safer decisions in... 
    Senior
    Full time
    Worldwide

    Mastercard

    Hungary
    15 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Machine Learning Engineer, LLM Inference Optimization. Be the first to apply!