Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Machine Learning Engineer (LLM inference)

GMI Cloud

GMI Cloud is a fast-growing AI infrastructure company backed by Headline VC and one of only six cloud providers worldwide to earn NVIDIA’s prestigious Reference Platform Cloud Partner designation . We operate 8 of our own GPU clusters across the U.S. and Asia, delivering a full spectrum of services from GPU compute service to AI model inference API solutions. As an NVIDIA Reference Platform Cloud Partner, our infrastructure meets the highest standards for performance, security, and scalability in AI deployments. We empower AI startups and enterprises to “build AI without limits,” providing everything they need to prototype, train, and deploy AI models quickly and reliably. About this role We are hiring a Machine Learning Engineer, LLM Optimization to build a world-leading inference optimization team and make GMI Cloud the industry benchmark for LLM serving performance. This role is for engineers who want to work at the frontier of AI systems. You will drive the research, validation, and productionization of the most advanced inference optimization techniques, and turn them into real competitive advantage across GMI’s inference platform. Our goal is to make GMI the company that leads the industry in how fast we discover, evaluate, combine, and operationalize the best optimization strategies for real customer workloads. That means not only adopting the latest advances, but also defining best practices, developing our own optimization methodologies, and building the internal framework that keeps GMI ahead of the curve. You will focus on B200-first optimization, with support for H200 evolution, across core domains including quantization, speculative decoding, KV cache and memory management, prefill/decode disaggregation, and system-level inference optimization. You will work closely with platform and infrastructure teams to transform cutting-edge ideas into measurable gains in latency, throughput, cost efficiency, and production scalability. Key Responsibilities Drive frontier research and engineering in LLM inference optimization, building GMI’s industry-leading capabilities in performance, efficiency, and scalability. Develop next-generation optimization strategies for large-scale LLM serving across model execution, runtime systems, and production inference platforms. Advance state-of-the-art techniques in quantization and precision optimization to improve throughput, latency, memory efficiency, and cost-performance across modern GPU systems. Push the frontier of speculative decoding and related acceleration methods, including both systems and model-level approaches for faster generation. Lead innovation in KV cache and memory optimization , improving long-context serving efficiency, memory utilization, and multi-tenant performance. Develop advanced architectures for prefill/decode disaggregation and other distributed inference optimization strategies for large-scale production environments. Drive system-level optimization across scheduling, batching, routing, gateway orchestration, adapter serving, and end-to-end inference efficiency. Build scalable optimization frameworks, performance methodologies, and engineering practices that allow GMI to stay ahead of the industry as models, hardware, and serving patterns evolve. Turn cutting-edge optimization ideas into production-ready capabilities that improve real-world customer workloads across latency, throughput, quality, and cost. Collaborate closely with platform, infrastructure, and product teams to make inference optimization a core technical advantage of GMI Cloud. Required Skills Strong hands-on experience with LLM inference systems and performance optimization. Solid understanding of inference metrics and tradeoffs, including TTFT, ITL, throughput, goodput, tail latency, GPU utilization, memory efficiency, and quality/cost tradeoffs . Experience with one or more modern serving stacks such as SGLang, vLLM, TensorRT-LLM, Triton, or similar systems. Deep familiarity with GPU-based inference , model serving architecture, and production bottlenecks around compute, memory bandwidth, KV-cache behavior, and scheduling. Strong experimentation skills: able to design benchmarks, interpret results, debug regressions, and produce actionable conclusions rather than isolated microbenchmark wins. Comfortable working across research-style validation and production engineering, with a bias toward measurable impact in real customer scenarios. Strong coding and systems skills in Python , with practical experience in profiling, observability, and performance debugging. Clear communication skills and the ability to explain technical tradeoffs to both engineers and cross-functional stakeholders. Preferred Qualifications 1+ years of hands-on experience in LLM inference optimization , ML systems optimization , or closely related areas. Experience working on optimization for large-scale model serving, such as latency reduction, throughput improvement, memory efficiency, or cost-performance tuning. Familiarity with one or more major areas of inference optimization, including quantization , speculative decoding , KV cache optimization , prefill/decode disaggregation , or system-level serving optimization . Experience with modern LLM serving stacks, GPU inference systems, or production ML infrastructure is a strong plus. #J-18808-Ljbffr GMI Cloud

Vacancy posted 19 hours ago
Similar jobs that could be interesting for youBased on the Machine Learning Engineer (LLM inference) in Mountain View, CA vacancy
  • $150k - $230k

     ...the RoleWe are looking for a hands-on Machine Learning Engineer to drive the post-training of our large...  ...-ready code.RequirementsHands-on LLM post-training experience. You have personally...  .../Accelerate, DeepSpeed or FSDP, and inference engines like vLLM.Solid understanding... 
    Suggested
    Full time
    Local area
    Work from home

    News Break

    Mountain View, CA
    2 days ago
  • $209k - $313k

     ...themselves, live in the moment, learn about the world, and have fun...  ...other digital services.Snap Engineering teams build fun and...  ...forefront.We’re looking for a Machine Learning Engineer to join Snap...  ...Strong understanding of causal inference and modern approaches to estimating... 
    Suggested
    Full time
    Live in
    Work at office
    Local area

    Snap

    Palo Alto, CA
    6 days ago
  • $174.72k - $295.68k

     ...through cutting-edge R&D in AI, machine learning, and smart connectivity.Our...  ...build strong foundation for LLM deployment and quality sign-...  ..., PTQ, QAT, on-vehicle inference and related fields.Key ResponsibilitiesDevelop...  ...programming and software engineering skills.Ability to work... 
    Suggested
    Full time

    XPENG Motors

    Santa Clara, CA
    2 days ago
  •  ...personalized agent experiences. Knowledge and passion in machine learning algorithms, Gen AI, LLMs, and natural language...  ...Knowledge of fine-tuning strategies (QLORA, DPO) and inference optimization (vLLM, TensorRT-LLM). Research experience in agentic AI or related fields... 
    Suggested
    Full time
    Work experience placement

    Eightfold

    Santa Clara, CA
    1 day ago
  • $229k - $343k

     ...express themselves, live in the moment, learn about the world, and have fun together...  ..., and on-device and server-side inference. Our team creates intuitive tools, platforms...  ...like Spectacles.We’re looking for a Machine Learning Engineer to join Snap Inc!What you’ll do:Develop... 
    Suggested
    Full time
    Live in
    Work at office
    Local area
    Worldwide

    Snap

    Palo Alto, CA
    6 days ago
  • $232k - $310k

     ...collaboration, and high standards. Our engineers, product leaders, and go-to-...  ...the boundaries of applied machine learning. We work with massive datasets, tackle...  ...tuning strategies (QLORA, DPO) and inference optimization (vLLM, TensorRT-LLM).Desired Skills & Experience:... 
    Work experience placement
    Work at office
    Remote work
    Flexible hours
    3 days per week

    Eightfold

    Santa Clara, CA
    3 days ago
  •  ...Distributed LLM Inference Engineer At Anyscale, we're on a mission to democratize distributed computing and make it accessible to software...  ...that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber, Spotify, Instacart, Cruise,... 
    Work at office

    Anyscale

    Palo Alto, CA
    1 day ago
  • $215k - $285k

     ...intelligence to every moving machine on the planet. Applied Intuition...  ...Bangalore; Seoul; and Tokyo. Learn more at applied.co. We are an...  ...looking for a performance engineer who specializes in making large...  ..., and high-throughput batch inference sweeping petabytes of real-... 
    Full time
    For contractors
    For subcontractor
    Casual work
    Work at office
    Remote work
    Day shift

    NLP PEOPLE

    Sunnyvale, CA
    3 days ago
  •  ...SoftwareClient: WiproContact: Meghana GorusuCompany: SRI Tech SolutionsJob Title: ML Engineer with LLMLocation: Sunnyvale, CA(onsite)Job Description:6-8 years of experience in machine learning and LLM, with a proven track record in image processing and analysis.Development and... 

    SRI Tech

    Sunnyvale, CA
    2 days ago
  • $196k - $221k

     ...alongside industry-veteran scientists and engineers. As a Machine Learning Engineer, you’ll bring your strong...  ...evolve large-scale SID / ASR / NLP / LLM systems that power mission-critical...  ...training, fine-tuning, post-training, and inference strategies for large language and... 
    Permanent employment

    Otter.ai

    Mountain View, CA
    2 days ago
  • $250k - $350k

    About the RoleWe are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly...  ..., including quantization, attention acceleration, and deep learning compiler stacks.GPU & Parallelism: Deep knowledge of GPU... 
    Work at office
    3 days per week

    Pika

    Palo Alto, CA
    3 days ago
  • GMI Cloud, a fast-growing AI infrastructure company, is hiring a Machine Learning Engineer, LLM Optimization to build a world-leading inference optimization team. You will drive research, validation, and productionization of advanced optimization techniques to boost latency... 

    GMI Cloud

    Mountain View, CA
    1 day ago
  •  ...a causal question. About The Role We are hiring a Causal Machine Learning Engineer to help build the causal ML foundation behind how DoorDash...  ...You Because You Have Deep practical experience with causal inference, econometrics, experimentation, or causal ML. Experience... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    DoorDash

    Sunnyvale, CA
    3 days ago
  • $193.3k - $261.5k

     ...Senior Software Development Engineer to bring diverse perspectives...  ...pipelines, real-time inference systems, and feature stores...  ...in distributed systems and machine learning infrastructure. You will foster...  ...Knowledge of Machine Learning and LLM fundamentals, including... 
    Internship
    Local area
    Worldwide
    Flexible hours

    Amazon

    Sunnyvale, CA
    2 days ago
  • # AI/ML Engineer, Companion Palo Alto / RemoteCAN Companion is an always-on care presence for...  ...the models, retrieval systems, and inference pipelines that make Companion feel genuinely...  ...experience, including at least one production LLM or NLP system.* Proficiency in Python and... 

    CAN-USA

    Palo Alto, CA
    3 days ago
  • $165.45k - $259.75k

     ...HP is seeking a HP is seeking a Principal Software ML Engineer to define and guide the architecture of modern...  ...software capabilities, including AI service integration, LLM workflows, model orchestration, inference API integration, evaluation pipelines, responsible AI... 
    Full time
    Temporary work
    Local area
    Flexible hours
    Shift work

    HP

    Palo Alto, CA
    1 day ago
  • $195k - $230k

     ...About the RoleWe are looking for a Senior Machine Learning Engineer to help evolve our large-scale recommendation systems and apply AI / LLM technologies to real-world production...  ...Own systems from offline training online inference A/B experimentation metric analysis.Identify... 
    Full time
    Local area
    Work from home

    News Break

    Mountain View, CA
    3 days ago
  • $188.5k - $282.7k

     ...Rubrik's Semantic AI Governance Engine, which is the first system...  ...traffic.At its core, SAGE is "LLM-as-judge" applied to AI...  ...Performance Model Serving and Inference Infrastructure (25% of time)Designing...  ...higher) in Computer Science, Machine Learning, Computer Engineering,... 
    Permanent employment
    Local area

    Rubrik

    Palo Alto, CA
    6 days ago
  •  ...workflows, and continuously learn and adapt.Moveworks is...  ...Moveworks’ Reasoning Engine and natural language...  ...RoleWe are looking for a Machine Learning Engineer to...  ...for building and serving LLM’s at Moveworks. This role...  ...training and inference pipeline for large language... 
    Work at office
    Remote work
    Flexible hours

    Moveworks

    Mountain View, CA
    5 days ago
  • $197.5k - $272k

     ...looking for a great Staff AI Engineer to join our seasoned AI...  ...(e.g., Transformers, LLM, CNN, LSTM, Trees) to...  ...and supervised learning models (e.g., Autoencoders...  ...related field.7+ years in Machine Learning Engineering, with...  ...C++ (C++14/17 for inference).Deep proficiency with... 
    Work at office
    Worldwide
    Flexible hours
    Shift work
    3 days per week

    Sonatus

    Sunnyvale, CA
    2 days ago
  • Overview As a Principal Machine Learning Engineer, you will drive the development and implementation...  ...ML models (e.g., ranking, retrieval, LLM-based systems) to optimize user experience...  ...for offline training and online inference at scaleOversee end-to-end deployment... 
    Local area
    Remote work

    Atlassian

    Mountain View, CA
    3 days ago
  • Member of Technical Staff, ML Inference Engineering Sanas is pioneering the future of human communication...  ...and CUDA Fluency in the LLM serving stack, from kernels and quantization...  ...ML projects Experience managing machine learning workloads on Kubernetes clusters Experience... 

    Sanas

    Palo Alto, CA
    5 days ago
  •  ...the world running. Our Team's Vision: Our Engineering team is shaping the future of...  ...databases at scale. AI Ops: Experience with LLM deployment optimization (e.g., vLLM, TensorRT...  ...TensorRT‑LLM) or managing proprietary model inference endpoints. This position involves access... 
    Immediate start

    Illumio

    Sunnyvale, CA
    4 days ago
  • $184.7k - $324.8k

    Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference Santa Clara, California, United States Machine Learning and AI We are the Foundation...  ...projects from end to end. Hands‑on experience with LLM inference stacks. Working knowledge of GPU or TPU... 
    Worldwide
    Relocation

    Apple

    Santa Clara, CA
    3 days ago
  •  ...job poster from GMI Cloud Focusing on LLM, AI Infrastructure, and AIGC. Opportunities...  ..., and large enterprises worldwide. Machine Learning Engineer - Video Generation We are seeking a...  ...infrastructure to support large-scale video inference. Design common workload processing... 
    Full time
    Worldwide

    GMI Cloud

    Mountain View, CA
    2 days ago
  • Bright Vision Technologies is seeking a Machine Learning Infrastructure Engineer to design, build, and operate high-performance inference platforms for serving large ML models in production. This remote U.S. role focuses on the systems engineering side of AI deployment,... 
    Remote job

    Bright Vision Technologies

    Mountain View, CA
    5 days ago
  • $150k

     ...researchers, data scientists, and engineers, tackling the most...  ...performance computing in deep learning, driving impactful discoveries...  ...pioneers. The Role As a Machine Learning Engineer at the Institute...  ...~ Hands-on experience with LLM algorithms, such as Supervised... 
    Full time
    Worldwide
    Visa sponsorship

    Institute Of Foundation Models

    Sunnyvale, CA
    1 day ago
  •  ...-latency, high-throughput AI for multi-node GPU workloads. As a Senior Engineer, you will shape core infrastructure and architecture decisions, lead performance optimizations, and own the inference engine to scale research and production workloads. This role is in the... 

    Sanas

    Palo Alto, CA
    5 days ago
  • $173k - $259k

     ...express themselves, live in the moment, learn about the world, and have fun together...  ..., and on-device and server-side inference. Our team creates intuitive tools, platforms...  ...like Spectacles.We're looking for a Machine Learning Engineer to join our Generative ML team!What you... 
    Full time
    Live in
    Work at office
    Local area
    Worldwide

    Snap

    Palo Alto, CA
    5 days ago
  • $342.7k

     ...internet is managed. As a Distinguished Machine Learning Engineer, you will have a seat at the table...  ...development, and production-level deployment of LLM-based systems serving at least 10,000+...  ...fine-tuning.Deep knowledge of inference serving optimizations, such as KV/prefix... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    Shift work

    CISCO Systems

    San Jose, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Machine Learning Engineer (LLM inference). Be the first to apply!