Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Machine Learning Engineer, LLM Inference Optimization in San Francisco

Energy Jobline ZR

Job Description About Us GMI Cloud is a fast-growing AI infrastructure company backed by Headline VC and one of only seven cloud providers worldwide to earn NVIDIA's prestigious Reference Platform Cloud Partner designation. We operate 8 of our own GPU clusters across the U.S. and Asia, delivering a full spectrum of services from GPU compute to AI model inference API solutions. As an NVIDIA Reference Platform Cloud Partner, our infrastructure meets the highest standards for performance, security, and scalability in AI deployments. We empower AI startups and enterprises to "build AI without limits," providing everything they need to prototype, train, and deploy AI models quickly and reliably. About this role GMI Cloud is building the leading inference optimization solution and the most advanced token platform in the global token market — and we are hiring world-class Machine Learning Engineers to make GMI the new industry benchmark for LLM serving performance, cost efficiency, and production reliability. Key Responsibilities Drive frontier research and engineering in LLM inference optimization across one of the four focus tracks (Speculative Decoding, Quantization, PD Disaggregation, KV Cache & Memory) while contributing across the full optimization stack. Develop next- optimization strategies for large-scale LLM serving across model execution, runtime systems, and production inference platforms — with B200 as the primary target and H200 as a continuing platform. Advance state-of-the-art techniques in quantization (NVFP4 / MXFP4 / FP8, QAT), speculative decoding (EAGLE-3, MTP, DFlash, ModelOpt, SpecForge), KV cache & memory management (LMCache / HiCache / NV KVBM, paged attention, prefix-aware routing), and PD disaggregation (NVIDIA Dynamo, KV-aware router/planner, fault recovery). Drive system-level optimization across scheduling, batching, routing, gateway orchestration, adapter serving, and end-to-end inference efficiency. Build scalable optimization frameworks, performance methodologies, and benchmark infrastructure that allow GMI to stay ahead of the industry as models, hardware, and serving patterns evolve. Productionize cutting-edge ideas into real customer workloads — measured by TTFT, ITL, throughput, goodput, tail latency, quality, and unit token cost. Engage with and contribute to the open-source community (vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo / ModelOpt, FlashInfer, LMCache, etc.) — read upstream code, file issues, send PRs, and publish tech blogs and case studies. Collaborate closely with platform, infrastructure, and product teams to make inference optimization a core technical advantage of GMI Cloud. Required Skills Strong hands-on experience with LLM inference systems and performance optimization on modern GPUs. Solid understanding of inference metrics and tradeoffs, including TTFT, ITL, throughput, goodput, tail latency, GPU utilization, memory efficiency, and quality/cost tradeoffs . Experience with one or more modern serving stacks such as SGLang, vLLM, TensorRT-LLM, NVIDIA Dynamo, or Triton . Deep familiarity with GPU-based inference, model serving architecture, and production bottlenecks around compute, memory bandwidth, KV-cache behavior, and scheduling. Demonstrable depth in at least one of the four focus areas: speculative decoding, quantization & precision, PD disaggregation, or KV cache & memory management. Strong experimentation skills: able to design benchmarks, interpret results, debug regressions, and produce actionable conclusions rather than isolated microbenchmark wins. Proficient with Claude Code at an advanced level — fluent with sub-agents, MCP servers, hooks, custom slash commands, and skills — with practical experience leveraging them for rapid iteration, profiling, observability, and performance debugging. Clear communication — able to explain technical tradeoffs to engineers and cross-functional stakeholders, and willing to publish results externally. Qualifications 2+ years of hands-on experience in LLM inference optimization , ML systems optimization , or PhD degree in related areas. Track record of large-scale model serving optimization (latency reduction, throughput improvement, memory efficiency, cost-performance tuning) in production. Specific track depth in one or more of: Speculative Decoding: EAGLE-3 / MTP / DFlash / Medusa / SpecForge / ModelOpt; experience training and shipping draft models for production. Quantization & Precision: NVFP4 / MXFP4 / FP8 / INT4-AWQ / GPTQ; QAT pipelines on Blackwell or Hopper; rigorous accuracy benchmarking. PD Disaggregation: NVIDIA Dynamo, KV-aware router/planner, large MoE serving (DeepSeek-V3/V4, Kimi, GLM, Minimax), fault recovery, autoscaling. KV Cache & Memory: LMCache / HiCache / NV KVBM, paged attention internals, prefix-aware routing, long-context and agentic workloads. Familiarity with FlashInfer, Blackwell MLA, FA4, TRT-LLM MLA, or NSA is a strong plus. Open-source contributions to vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo / ModelOpt, FlashInfer, LMCache, or related projects. Experience publishing technical blogs, case studies, or papers on inference optimization. #J-18808-Ljbffr Energy Jobline ZR

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Machine Learning Engineer, LLM Inference Optimization in San Francisco in San Francisco, CA vacancy
  • $180k - $270k

     ...is a Delaware-incorporated, San Francisco-based company pushing the boundary...  ...privacy protection. To learn more about Plaud, please...  ...throughput, ultra-low-latency inference engines for large language models or...  ...familiarity with modern LLM serving frameworks like vLLM... 
    Suggested
    Full time
    Work at office
    Worldwide

    Plaud

    San Francisco, CA
    1 day ago
  • $195k - $365k

     ...Delaware-incorporated, San Francisco-based company...  ...protection. To learn more about Plaud,...  ...of research and engineering, eager to design...  ...and edge-device optimization. Possess deep expertise...  ...: End‑to‑end inference and performance optimization...  ..., vLLM, TensorRT‑LLM, SGLang) to... 
    Suggested
    Full time
    Work at office
    Worldwide

    Plaud

    San Francisco, CA
    3 days ago
  •  ...top businesses trust to optimize billions in ad spend worldwide...  ...using optimization, machine learning, and causal inference. We are looking for...  ..., data scientists, data engineers and other MLEs to deliver...  ...distance of our offices in San Francisco, Seattle, and New York... 
    Suggested
    Full time
    Work at office
    Work from home
    Worldwide
    Flexible hours

    Haus Analytics

    San Francisco, CA
    1 day ago
  • $298k - $368k

     ...the system which learns the spatial-temporal...  ...teams on the optimization and integration into...  ...sensors, enabling engineers like you to (1) develop...  ...: Design VLM/LLM model architecture...  ...of experience in Machine Learning, with a...  ...latency on-device inference techniques and a deep... 
    Suggested
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  • $295k - $405.5k

     ...power of tech, data, and machine learning to connect this...  ...Machine Learning Platform Engineer, you will own the...  ...including training, inference, feature management,...  ...authorityExperience integrating LLM workflows into...  ...have headquarters in San Francisco and Kitchener-... 
    Suggested
    Part time
    Work experience placement
    Work at office
    Local area
    Remote work
    Monday to Friday
    Flexible hours
    3 days per week

    Faire

    San Francisco, CA
    4 days ago
  • $160k - $230k

     ...efficient and scalable inference for large language...  ...LLMs). Our mission is to optimize inference frameworks,...  ...Frameworks and Optimization Engineer to design, develop,...  ...shape the future of LLM inference infrastructure...  ...of experience in deep learning inference frameworks,... 
    Full time

    Together AI

    San Francisco, CA
    1 day ago
  • $203.5k - $299.3k

     ...RoleWe are hiring a Causal Machine Learning Engineer to help build the causal...  ...evaluation, promotion optimization, or marketplace decisioning...  ...practical experience with causal inference, econometrics,...  ...discrimination.Pursuant to the San Francisco Fair Chance Ordinance,... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    1 day ago
  • $155k - $180k

     ...100, use Roboflow’s machine learning open source and hosted...  ...not only product and engineering), so Roboflow employs...  ...of all of this is inference — one of our most important...  ..., vLLM (or other LLM/model deployment...  ...in New York City and San Francisco (and plan to open more... 
    Full time
    Second job
    Remote work
    Work from home
    Relocation package
    Flexible hours
    Night shift

    Roboflow

    San Francisco, CA
    1 day ago
  •  ...We're looking for a Machine Learning Engineer to design, build, and...  ..., monitoring, and inference Build intelligent...  ...services using modern NLP, LLM, classification,...  ...model quality Optimize model latency, scalability...  ...office presence in San Francisco and New York. R&D... 
    Full time
    Work at office
    Remote work
    Flexible hours
    2 days per week

    Plenful

    San Francisco, CA
    1 day ago
  • $160k - $230k

     ...About the Role Together AI is seeking a Machine Learning Engineer to join our Inference Engine team, focusing on optimizing and enhancing the performance of our AI inference...  ...systems such as TGI, vLLM, TensorRT-LLM, Optimum ~ Preferred: Knowledge of AI inference... 
    Full time

    Together Ai

    San Francisco, CA
    1 day ago
  • $200.8k - $251k

    A leading AI technology company in San Francisco seeks a team member to build and optimize a machine learning framework for large language models. Candidates should...  ...system optimization experience and solid software engineering skills, particularly in tools like CUDA and... 
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  • $203.5k - $299.3k

     ...Role We are hiring a Causal Machine Learning Engineer to help build the causal...  ...evaluation, promotion optimization, or marketplace decisioning...  ...experience with causal inference, econometrics, experimentation...  .... Pursuant to the San Francisco Fair Chance Ordinance, Los... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Visa Hunt

    San Francisco, CA
    4 days ago
  •  ...Role We are hiring a Causal Machine Learning Engineer to help build the causal...  ...evaluation, promotion optimization, or marketplace decisioning...  ...experience with causal inference, econometrics, experimentation...  .... Pursuant to the San Francisco Fair Chance Ordinance, Los... 
    Hourly pay
    Work at office
    Local area
    Flexible hours

    DoorDash USA

    San Francisco, CA
    5 days ago
  • $264.8k - $331k

     ...Machine Learning Systems Research Engineer, Agent Post-training - Enterprise GenAIAI...  ...technologies to optimize our ML system. Your...  ...optimize our training and inference framework.Post-...  ...least 1-3 years of LLM training in a...  ...in the locations of San Francisco, New York, Seattle... 
    Full time

    Scale AI

    San Francisco, CA
    1 day ago
  • Jaide Health is seeking an engineer for their Model Efficiency team in San Francisco. The role focuses on building reliable...  ...techniques such as GPU/CUDA optimizations and collaborate closely with...  ...Python and insights into the LLM inference ecosystem. A commitment to diversity... 
    Remote job

    Jaide Health

    San Francisco, CA
    1 day ago
  • $170k - $245k

     ...ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber,...  ...the roleAs a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of...  ...: $170K - $245KLocationSan Francisco; Palo AltoEmployment TypeFull... 
    Work at office

    Anyscale

    San Francisco, CA
    4 days ago
  • $250k

    Title : ML Inference Engineer Location : San Francisco, CA Salary : $250k base + equity An AI Unicorn startup is hiring a Senior Machine Learning Inference Engineer for a full-time role. You will be responsible for improving efficiency for AI-native infrastructure powered... 
    Full time

    Oscar Technology

    San Francisco, CA
    4 days ago
  • $200k - $400k

     ...Description Job Description Machine Learning Infrastructure Engineer San Francisco, CA · On-site (5 days/...  ...training and inference backbone for a foundation...  ...level GPU operations to optimize performance Track...  ...a scale beyond typical LLM training Well-funded... 
    Full time
    Visa sponsorship
    Relocation package

    David Joseph & Company

    San Francisco, CA
    18 days ago
  •  ...We are a small, fast-growing team of engineers in San Francisco powering Fortune 100 enterprises, YC...  ...San Francisco office ~ Eager to learn and adapt quickly ~ Prior startup...  ...curation, and active learning pipelines Optimize inference, batching, and quantization on GPU... 
    Full time
    Work at office
    Visa sponsorship
    Relocation package

    The Pulse

    San Francisco, CA
    1 day ago
  • $150k - $190k

     ...software stack for engineering and manufacturing across...  ...through AI inference across the entire engineering...  ...new levels of optimization and automation in design...  ...For As a Machine Learning Engineer in Delivery...  ...remote based in the San Francisco area. This Role... 
    Remote job
    Full time
    Flexible hours

    Physicsx

    San Francisco, CA
    1 day ago
  • $165k - $230k

     ...re looking for exceptional Machine Learning Engineers focused on Ads to help...  ...creative is generated, ranked, optimized, and ultimately performs....  ...experimentation through inference and serving. Work closely...  ...hybrid role based in the San Francisco Bay Area . Team members are... 
    Full time
    Work at office
    Remote work
    Worldwide
    3 days per week

    Higgsfield

    San Francisco, CA
    1 day ago
  •  ...person five days a week in our San Francisco, NYC, or London offices. About the Role As a Machine Learning Engineer on the Marketplace team,...  ...demand, and the need to optimize across speed, quality, and...  ...• Real-time and batch inference systems embedded in product... 
    Full time
    Work at office
    Relocation package

    Mercor

    San Francisco, CA
    1 day ago
  •  ...We are looking for a Machine Learning Engineer to join the growing AI and...  ...model building, evaluation, optimizing performance, and ensuring...  ...your time on-site in our San Francisco Office — three days per week...  ...to scaling and optimizing inference and deployment Shape AI... 
    Full time
    Work at office
    Worldwide
    Flexible hours
    3 days per week

    Strava

    San Francisco, CA
    1 day ago
  •  ...edge AI technology company based in San Francisco is seeking a specialist to design...  ...GPU systems for high-throughput inference and model performance optimization. The ideal candidate will have...  ...solid understanding of reinforcement learning technologies. Comprehensive... 

    Reflection AI

    San Francisco, CA
    3 days ago
  • $210k - $240k

     ...Machine Learning Engineer   You'll build the ML behind Firecrawl — the models...  ...across extraction quality and LLM-driven features. You'll...  ...the process. Location: San Francisco, CA (Hybrid, on-site required...  ...everything else Someone who optimizes for process over shipping... 
    Full time
    Temporary work
    For contractors
    Remote work
    Visa sponsorship
    Flexible hours

    Firecrawl

    San Francisco, CA
    1 day ago
  • Reactor in San Francisco is seeking an ML Inference Engineer to maximize performance of generative media models and push ultra-low-latency, high-throughput...  .... You will craft an in-house runtime, implement optimizations with PyTorch tools, and collaborate with partner teams... 

    Reactor.am

    San Francisco, CA
    4 days ago
  • $197.3k - $225.1k

    Lead Machine Learning Engineer At Capital One, we are creating...  ...to deliver optimized ML models at scale such...  ...large language model inference, similarity search,...  ...introduce state-of-the-art LLM optimization...  ...Learning Engineer San Francisco, CA: $215,200 - $245... 
    Full time
    Part time
    Internship
    H1b
    Local area

    Capital One Financial Corporation

    San Francisco, CA
    1 day ago
  • $227.2k - $284k

     ...model and systems optimization, and applied research...  ...the RoleAs a Staff Machine Learning Research Engineer, you will operate...  ...training/fine-tuning, inference, memory and...  ...training methods, LLM alignment, or applied...  ...in the locations of San Francisco, New York, Seattle... 
    Full time

    Scale AI

    San Francisco, CA
    1 day ago
  • $270k - $315k

     ...committed researchers, engineers, policy experts, and...  ..., implementation, and optimization of our Order-to-Cash (...  ...Experience with AI/LLM integration for financial...  ...Problems in AI Safety, and Learning from Human Preferences...  ...headquartered in San Francisco. We offer competitive... 
    Work at office
    Visa sponsorship
    Flexible hours

    Neura Market

    San Francisco, CA
    1 day ago
  •  ...Description Principal Machine Learning Engineer, Artificial...  ...operate across training, inference, evaluation, and...  ...deployment, including GPU optimization, memory efficiency,...  ...- Experience with LLM inference frameworks...  ....   Keywords:  San Francisco CA Jobs, Principal... 
    Remote work
    Work from home

    Ginas Tech Jobs

    San Francisco, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Machine Learning Engineer, LLM Inference Optimization in San Francisco. Be the first to apply!