Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Machine Learning Engineer, LLM Inference Optimization

Full-time

Nebius

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role Nebius Token Factory is building fast, reliable, and cost-efficient inference services for frontier models. As a Senior Machine Learning Engineer on our Applied AI team, you will own model and endpoint optimization from model artifacts through production deployment. Your work will span model internals, inference engines, serving architecture, and benchmarking, with a focus on improving latency, throughput, memory efficiency, GPU utilization, and cost per token while maintaining model quality and reliability. This is a hands-on role in which you will work on complex optimization projects, diagnose difficult serving problems, and deliver measurable improvements in production. Working closely with kernel and platform engineers, you will evaluate serving configurations, resolve performance and quality regressions, and optimize inference for real-world workloads, supported by reproducible benchmarks and safe production rollouts. Your responsibilities : - Own optimization work for specific model families, customer endpoints, or serving backends. - Run engine comparisons and recommend practical serving configurations for specific workloads. - Debug model quality or performance regressions during production rollouts.

- Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token. - Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or similar systems. - Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery. - Implement or integrate speculative decoding, draft-model approaches, KV -cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving. - Build reproducible benchmark harnesses for TTFT , TPOT , tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token. - Partner with GPU kernel engineers and platform engineers to diagnose bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers. - Write clear design docs, performance reports, rollout plans, and customer-facing technical explanations. Must-haves : - Strong Python and PyTorch engineering skills. - Hands-on experience deploying or optimizing LLM, VLM , or high-throughput transformer inference systems. - Practical knowledge of at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or equivalent internal systems. - Strong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving. - Ability to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs. - Strong communication skills and ability to collaborate with research, kernel, infrastructure, product, and customer teams.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior Machine Learning Engineer, LLM Inference Optimization in Remote vacancy
  • $298k - $368k

     ...the system which learns the spatial-temporal...  ...teams on the optimization and integration into...  ...sensors, enabling engineers like you to (1) develop...  ...: Design VLM/LLM model architecture...  ...of experience in Machine Learning, with a...  ...latency on-device inference techniques and a deep... 
    Senior
    Full time
    Remote work

    Waymo

    San Francisco, CA
    more than 2 months ago
  •  ...Community You Will Join:  Machine Learning and Artificial...  ...services and tools including LLM fine-tuning, alignment and optimization, RAG/Search, LLM...  ...principal machine learning engineer, you will be responsible...  ...optimizing models and inference run-time ~ Post-training... 
    Suggested
    Remote job
    Full time
    Casual work
    Live in
    Work at office

    Airbnb, Inc.

    United States
    more than 2 months ago
  •  ...explore, create, play, learn, and connect with friends...  ...a billion people with optimism and civility, and...  ...experiences for everyone. Our engine’s resource management...  ...the application of machine learning in real-time engine...  ...Design ML models that infer player and interaction... 
    Senior
    Full time

    Roblox

    Remote
    more than 2 months ago
  • $204k - $259k

     ...builds the system which learns the spatial-temporal...  ...downstream teams on the optimization and integration into...  ...of sensors, enabling engineers like you to (1) develop...  ...years of experience in Machine Learning, with a focus...  ...model development (LLM, VLM, or similar foundation... 
    Senior
    Full time
    Remote work

    Waymo

    San Francisco, CA
    more than 2 months ago
  •  ...About the Role We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika'...  ...at scale.   You will design and optimize inference pipelines, implement state...  ..., attention acceleration, and deep learning compiler stacks. GPU &... 
    Suggested
    Full time
    Work at office
    3 days per week

    Pika

    Remote
    more than 2 months ago
  • $213k - $263k

     ...Waymo AI Foundations team is to develop machine learning solutions addressing open problems in...  ...demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust...  ...to smaller real-time models. Explore LLM/VLM distillation recipes to maximally... 
    Senior
    Full time
    Remote work

    Waymo

    Remote
    more than 2 months ago
  • $213k - $263k

     ...Drive cross-functional collaboration to engineer robust, high-reliability training...  ...Have: ~ BS or MS in Computer Vision, Machine Learning, Robotics, or a related field. ~4+ years...  .... Hands-on experience managing and optimizing large-scale teacher-student training... 
    Senior
    Full time
    Remote work

    Waymo

    Mountain View, CA
    more than 2 months ago
  • $213k - $263k

     ...states. The ML Optimization team at Waymo...  ...lifecycle of the machine learning workflow, including...  ...are looking for engineers with ML software...  ...Waymo onboard ML inference engine for Waymo...  ...will report to the Senior Manager of Runtime...  ...building or scaling LLM serving systems,... 
    Senior
    Full time
    Remote work

    Waymo

    Remote
    more than 2 months ago
  • $195k - $230k

     ...We are looking for a Senior Machine Learning Engineer to help evolve our large...  ...systems and apply AI / LLM technologies to real-world...  ...ranking, and multi-objective optimization to balance engagement, retention...  ...offline training → online inference → A/B experimentation →... 
    Senior
    Full time
    Local area
    Work from home

    News Break

    Mountain View, CA
    more than 2 months ago
  •  ...Position Summary The Machine Learning Engineer will be responsible for the...  ...appropriately for the chosen LLM and training pipeline...  ...size, and training epochs to optimize model performance. Integration...  ...pipelines for model training, inference, and deployment.... 
    Senior
    Full time
    H1b
    Remote work
    Flexible hours

    C The Signs

    Boston, MA
    more than 2 months ago
  •  ...are seeking an experienced Senior Machine Learning Engineer to join our AI/ML team and...  ...fine-tuning, preference optimization, and other post-training techniques...  ...model-serving and inference infrastructure for open-weight...  .... ~ Experience with LLM fine-tuning and post-training... 
    Senior
    Remote job
    Full time
    Work at office

    Air Company

    United States
    15 days ago
  •  ...everyone else has simply learned to live with. We...  ...You’ll help define how machine learning models run...  ...ll work with systems engineers, product teams, hardware...  ...combines applied ML, inference optimization, evaluation, and...  ...SGLang, vLLM, TensorRT-LLM, ONNX Runtime, Triton... 
    Senior
    Full time
    Local area

    Cloudflare

    Remote
    15 days ago
  • $235.2k - $294k

     ...The goal of a Senior Machine Learning Engineer at Scale is to own how we apply generative...  ...evaluation benchmarks, LLM judges, and verifiers -...  ...infrastructure to automate and optimize our ML services Work...  ...Geospatial or GEOINT experience Inference optimization experience... 
    Senior
    Full time

    Scale Ai, Inc.

    Remote
    13 days ago
  • $188.5k - $282.7k

     ...Semantic AI Governance Engine, which is the first...  ...At its core, SAGE is "LLM-as-judge" applied to...  ...fine-tuning, preference optimization (DPO/RLAIF), and...  ...Performance Model Serving and Inference Infrastructure (25%...  ...in Computer Science, Machine Learning, Computer Engineering... 
    Senior
    Permanent employment
    Full time
    Local area

    Rubrik

    Remote
    more than 2 months ago
  • $150k - $180k

     .... About the Role: As our Senior Machine Learning Engineer, you’ll own the intelligence layer...  ..., put physics-based, ML, and LLM-powered models into production,...  ...with LLM cost/latency optimization (prompt caching, batch inference) and model governance (managing... 
    Senior
    Full time
    Work at office
    Immediate start
    Shift work

    Thalo Labs

    Remote
    more than 2 months ago
  •  ...motivated and experienced Machine Learning Engineer to join our AI &...  ...production engineering (inference systems, integration, and optimization). Responsibilities...  ...engineering techniques and LLM frameworks ~...  ...celebrates your commitment and seniority (including paid... 
    Senior
    Remote job
    Full time
    Temporary work

    Keeper Security

    United States
    more than 2 months ago
  • $175k - $200k

     ...Senior Machine Learning Engineer Truveta is the world’s first health provider led...  ...expertise in applied AI, model optimization, and agentic intelligence...  ...a deep understanding of LLM fundamentals —...  ...efficiency, interpretability, and inference performance. Think and... 
    Senior
    Full time
    For contractors
    Visa sponsorship
    Work visa
    Flexible hours

    Truveta

    Remote
    23 days ago
  • $260k - $330k

     ...have: Experience building large-scale prediction or optimization systemsPubMatic is the leading AI-powered ad tech company...  ...environments.About the Role:We are looking for a Senior Principal Machine Learning Engineer to help build the next generation of performance optimization... 
    Senior
    Work at office
    Remote work

    PubMatic

    Redwood City, CA
    4 days ago
  • $150.75k - $241.2k

     ...are seeking a seasoned Machine Learning Engineer to join a new team...  ...reasoning systems. As a senior engineer on this team...  ...infrastructure, the inference and serving stack,...  ...multimodal workloads, optimizing latency, throughput,...  ...agentic systems and LLM tool-use orchestration... 
    Senior
    Work experience placement
    Work at office
    Remote work

    Axon

    Seattle, WA
    1 day ago
  • $232k - $310k

     ...and high standards. Our engineers, product leaders, and...  ...boundaries of applied machine learning. We work with massive datasets...  ...environments and optimize system performance.Contribute...  ...(QLORA, DPO) and inference optimization (vLLM, TensorRT-LLM).Desired Skills & Experience... 
    Work experience placement
    Work at office
    Remote work
    Flexible hours
    3 days per week

    Eightfold

    Santa Clara, CA
    4 days ago
  • $172k - $229k

     ...mining framework, is the engine that powers this discovery. As a Senior Machine Learning Engineer on the Data...  ...learning, retrieval optimization, and reasoning systems...  ...drastically reducing inference latency and memory footprint...  ...of-thought models, or LLM-based planning.... 
    Senior
    Full time
    Work at office
    Remote work

    Motional

    Las Vegas, NV
    more than 2 months ago
  •  ...Senior Machine Learning Engineer Department: Engineering Employment Type: Full...  ...production Design systems that infer structured attributes and...  ...indexing strategies Optimize models for inference latency...  ...Experience building LLM or VLM pipelines and the evaluation... 
    Senior
    Full time
    Remote work
    Flexible hours

    Clearview AI

    New York, NY
    5 days ago
  • $144k - $233.1k

     ...Senior Machine Learning Engineer Since 2003, Entrata has evolved from a visionary, student-led startup...  ..., and task performance. Optimize model inference, serving, and deployment for performance...  .... ~ Familiarity with modern LLM tooling, model serving, and inference... 
    Senior
    Local area
    Remote work
    Worldwide
    Flexible hours

    Entrata

    Lehi, UT
    2 days ago
  • $130k - $160k

     ...has an opportunity for a Senior Machine Learning Engineer. Candid's Data Science and...  ..., observability, and inference performance as that ownership...  ...familiarity with serving optimization techniques such as quantization...  ...or another managed LLM service (Anthropic API, Azure... 
    Senior
    Temporary work
    Summer work
    Local area
    Remote work
    Monday to Friday
    Flexible hours

    Candid

    New York, NY
    5 days ago
  •  ...and continuously learn and adapt. Moveworks...  ...’ Reasoning Engine and natural language...  ...are looking for a Machine Learning Engineer...  ...building and serving LLM’s at Moveworks....  ...critical in building, optimizing and scaling end-to...  ...training and inference pipeline for large... 
    Senior
    Full time
    Work at office
    Remote work
    Flexible hours

    Servicenow

    Remote
    a month ago
  • $201.3k - $352.3k

     ...entrepriseIt all started when engineer Fred Luddy wrote code...  ...Emerging tech is a small senior group inside AI...  ...or retrieval. Exposure to LLM fine-tuning or inference optimization in productionWhy join us...  ...assigned work location. Learn more here. To determine eligibility... 
    Senior
    Work experience placement
    Work at office
    Immediate start
    Remote work
    Flexible hours

    ServiceNow

    Santa Clara, CA
    3 days ago
  • $159.18k - $295.62k

     ...a career- we're hiring! Senior Machine Learning Engineer Team: Data & Audience Platform...  ..., model training and optimization, and the ML infrastructure...  ...scalable feature and inference pipelines on Databricks (PySpark...  ...search; evaluate LLM- based approaches for metadata... 
    Senior
    Temporary work
    Local area

    Warner Bros. Discovery

    New York, NY
    6 hours ago
  • $240k - $249.5k

     ...OpportunityGrubhub is looking for a Senior Staff Machine Learning Engineer to help lead the machine...  ...the evolution of our optimization objective from short-term...  ...our runtime environment: LLM-driven query and intent...  ...start, and real-time inference. Assess rigorously what actually... 
    Senior
    Full time
    Temporary work
    Work at office
    Flexible hours
    3 days per week

    GrubHub

    New York, NY
    5 days ago
  • $141k - $249k

     ...with autonomy and algorithm engineers to scale safe self-driving systems...  ...as TensorRT and modelopt to optimize the models running on the...  ...benchmark new CUDA kernels for inference. Comprehensively profile...  ...Rust. Experience in deep learning frameworks such as PyTorch.... 
    Senior
    Work at office
    Work from home
    Flexible hours

    Waabi

    San Francisco, CA
    1 day ago
  •  ...Senior Machine Learning Engineer - Data Science & Analytics Contract Length: 6-18...  ...integration Real-time inference Batch processing...  ...and ML data pipelines Optimize latency, scalability, and...  ...Experience supporting AI/LLM-enabled applications Team... 
    Senior
    Contract work
    Remote work

    RIT Solutions, Inc.

    Rosemont, IL
    7 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Machine Learning Engineer, LLM Inference Optimization. Be the first to apply!