Machine Learning Engineer, LLM Inference Optimization in San Francisco
Energy Jobline ZR
Job Description About Us GMI Cloud is a fast-growing AI infrastructure company backed by Headline VC and one of only seven cloud providers worldwide to earn NVIDIA's prestigious Reference Platform Cloud Partner designation. We operate 8 of our own GPU clusters across the U.S. and Asia, delivering a full spectrum of services from GPU compute to AI model inference API solutions. As an NVIDIA Reference Platform Cloud Partner, our infrastructure meets the highest standards for performance, security, and scalability in AI deployments. We empower AI startups and enterprises to "build AI without limits," providing everything they need to prototype, train, and deploy AI models quickly and reliably. About this role GMI Cloud is building the leading inference optimization solution and the most advanced token platform in the global token market — and we are hiring world-class Machine Learning Engineers to make GMI the new industry benchmark for LLM serving performance, cost efficiency, and production reliability. Key Responsibilities Drive frontier research and engineering in LLM inference optimization across one of the four focus tracks (Speculative Decoding, Quantization, PD Disaggregation, KV Cache & Memory) while contributing across the full optimization stack. Develop next- optimization strategies for large-scale LLM serving across model execution, runtime systems, and production inference platforms — with B200 as the primary target and H200 as a continuing platform. Advance state-of-the-art techniques in quantization (NVFP4 / MXFP4 / FP8, QAT), speculative decoding (EAGLE-3, MTP, DFlash, ModelOpt, SpecForge), KV cache & memory management (LMCache / HiCache / NV KVBM, paged attention, prefix-aware routing), and PD disaggregation (NVIDIA Dynamo, KV-aware router/planner, fault recovery). Drive system-level optimization across scheduling, batching, routing, gateway orchestration, adapter serving, and end-to-end inference efficiency. Build scalable optimization frameworks, performance methodologies, and benchmark infrastructure that allow GMI to stay ahead of the industry as models, hardware, and serving patterns evolve. Productionize cutting-edge ideas into real customer workloads — measured by TTFT, ITL, throughput, goodput, tail latency, quality, and unit token cost. Engage with and contribute to the open-source community (vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo / ModelOpt, FlashInfer, LMCache, etc.) — read upstream code, file issues, send PRs, and publish tech blogs and case studies. Collaborate closely with platform, infrastructure, and product teams to make inference optimization a core technical advantage of GMI Cloud. Required Skills Strong hands-on experience with LLM inference systems and performance optimization on modern GPUs. Solid understanding of inference metrics and tradeoffs, including TTFT, ITL, throughput, goodput, tail latency, GPU utilization, memory efficiency, and quality/cost tradeoffs . Experience with one or more modern serving stacks such as SGLang, vLLM, TensorRT-LLM, NVIDIA Dynamo, or Triton . Deep familiarity with GPU-based inference, model serving architecture, and production bottlenecks around compute, memory bandwidth, KV-cache behavior, and scheduling. Demonstrable depth in at least one of the four focus areas: speculative decoding, quantization & precision, PD disaggregation, or KV cache & memory management. Strong experimentation skills: able to design benchmarks, interpret results, debug regressions, and produce actionable conclusions rather than isolated microbenchmark wins. Proficient with Claude Code at an advanced level — fluent with sub-agents, MCP servers, hooks, custom slash commands, and skills — with practical experience leveraging them for rapid iteration, profiling, observability, and performance debugging. Clear communication — able to explain technical tradeoffs to engineers and cross-functional stakeholders, and willing to publish results externally. Qualifications 2+ years of hands-on experience in LLM inference optimization , ML systems optimization , or PhD degree in related areas. Track record of large-scale model serving optimization (latency reduction, throughput improvement, memory efficiency, cost-performance tuning) in production. Specific track depth in one or more of: Speculative Decoding: EAGLE-3 / MTP / DFlash / Medusa / SpecForge / ModelOpt; experience training and shipping draft models for production. Quantization & Precision: NVFP4 / MXFP4 / FP8 / INT4-AWQ / GPTQ; QAT pipelines on Blackwell or Hopper; rigorous accuracy benchmarking. PD Disaggregation: NVIDIA Dynamo, KV-aware router/planner, large MoE serving (DeepSeek-V3/V4, Kimi, GLM, Minimax), fault recovery, autoscaling. KV Cache & Memory: LMCache / HiCache / NV KVBM, paged attention internals, prefix-aware routing, long-context and agentic workloads. Familiarity with FlashInfer, Blackwell MLA, FA4, TRT-LLM MLA, or NSA is a strong plus. Open-source contributions to vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo / ModelOpt, FlashInfer, LMCache, or related projects. Experience publishing technical blogs, case studies, or papers on inference optimization. #J-18808-Ljbffr Energy Jobline ZR
$180k - $270k
...is a Delaware-incorporated, San Francisco-based company pushing the boundary... ...privacy protection. To learn more about Plaud, please... ...throughput, ultra-low-latency inference engines for large language models or... ...familiarity with modern LLM serving frameworks like vLLM...SuggestedFull timeWork at officeWorldwide$195k - $365k
...Delaware-incorporated, San Francisco-based company... ...protection. To learn more about Plaud,... ...of research and engineering, eager to design... ...and edge-device optimization. Possess deep expertise... ...: End‑to‑end inference and performance optimization... ..., vLLM, TensorRT‑LLM, SGLang) to...SuggestedFull timeWork at officeWorldwide- ...top businesses trust to optimize billions in ad spend worldwide... ...using optimization, machine learning, and causal inference. We are looking for... ..., data scientists, data engineers and other MLEs to deliver... ...distance of our offices in San Francisco, Seattle, and New York...SuggestedFull timeWork at officeWork from homeWorldwideFlexible hours
$298k - $368k
...the system which learns the spatial-temporal... ...teams on the optimization and integration into... ...sensors, enabling engineers like you to (1) develop... ...: Design VLM/LLM model architecture... ...of experience in Machine Learning, with a... ...latency on-device inference techniques and a deep...SuggestedFull timeRemote work$295k - $405.5k
...power of tech, data, and machine learning to connect this... ...Machine Learning Platform Engineer, you will own the... ...including training, inference, feature management,... ...authorityExperience integrating LLM workflows into... ...have headquarters in San Francisco and Kitchener-...SuggestedPart timeWork experience placementWork at officeLocal areaRemote workMonday to FridayFlexible hours3 days per week$160k - $230k
...efficient and scalable inference for large language... ...LLMs). Our mission is to optimize inference frameworks,... ...Frameworks and Optimization Engineer to design, develop,... ...shape the future of LLM inference infrastructure... ...of experience in deep learning inference frameworks,...Full time$203.5k - $299.3k
...RoleWe are hiring a Causal Machine Learning Engineer to help build the causal... ...evaluation, promotion optimization, or marketplace decisioning... ...practical experience with causal inference, econometrics,... ...discrimination.Pursuant to the San Francisco Fair Chance Ordinance,...Hourly payWork at officeLocal areaRemote workFlexible hours$155k - $180k
...100, use Roboflow’s machine learning open source and hosted... ...not only product and engineering), so Roboflow employs... ...of all of this is inference — one of our most important... ..., vLLM (or other LLM/model deployment... ...in New York City and San Francisco (and plan to open more...Full timeSecond jobRemote workWork from homeRelocation packageFlexible hoursNight shift- ...We're looking for a Machine Learning Engineer to design, build, and... ..., monitoring, and inference Build intelligent... ...services using modern NLP, LLM, classification,... ...model quality Optimize model latency, scalability... ...office presence in San Francisco and New York. R&D...Full timeWork at officeRemote workFlexible hours2 days per week
$160k - $230k
...About the Role Together AI is seeking a Machine Learning Engineer to join our Inference Engine team, focusing on optimizing and enhancing the performance of our AI inference... ...systems such as TGI, vLLM, TensorRT-LLM, Optimum ~ Preferred: Knowledge of AI inference...Full time$200.8k - $251k
A leading AI technology company in San Francisco seeks a team member to build and optimize a machine learning framework for large language models. Candidates should... ...system optimization experience and solid software engineering skills, particularly in tools like CUDA and...Full time$203.5k - $299.3k
...Role We are hiring a Causal Machine Learning Engineer to help build the causal... ...evaluation, promotion optimization, or marketplace decisioning... ...experience with causal inference, econometrics, experimentation... .... Pursuant to the San Francisco Fair Chance Ordinance, Los...Hourly payWork at officeLocal areaRemote workFlexible hours- ...Role We are hiring a Causal Machine Learning Engineer to help build the causal... ...evaluation, promotion optimization, or marketplace decisioning... ...experience with causal inference, econometrics, experimentation... .... Pursuant to the San Francisco Fair Chance Ordinance, Los...Hourly payWork at officeLocal areaFlexible hours
$264.8k - $331k
...Machine Learning Systems Research Engineer, Agent Post-training - Enterprise GenAIAI... ...technologies to optimize our ML system. Your... ...optimize our training and inference framework.Post-... ...least 1-3 years of LLM training in a... ...in the locations of San Francisco, New York, Seattle...Full time- Jaide Health is seeking an engineer for their Model Efficiency team in San Francisco. The role focuses on building reliable... ...techniques such as GPU/CUDA optimizations and collaborate closely with... ...Python and insights into the LLM inference ecosystem. A commitment to diversity...Remote job
$170k - $245k
...ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber,... ...the roleAs a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of... ...: $170K - $245KLocationSan Francisco; Palo AltoEmployment TypeFull...Work at office$250k
Title : ML Inference Engineer Location : San Francisco, CA Salary : $250k base + equity An AI Unicorn startup is hiring a Senior Machine Learning Inference Engineer for a full-time role. You will be responsible for improving efficiency for AI-native infrastructure powered...Full time$200k - $400k
...Description Job Description Machine Learning Infrastructure Engineer San Francisco, CA · On-site (5 days/... ...training and inference backbone for a foundation... ...level GPU operations to optimize performance Track... ...a scale beyond typical LLM training Well-funded...Full timeVisa sponsorshipRelocation package- ...We are a small, fast-growing team of engineers in San Francisco powering Fortune 100 enterprises, YC... ...San Francisco office ~ Eager to learn and adapt quickly ~ Prior startup... ...curation, and active learning pipelines Optimize inference, batching, and quantization on GPU...Full timeWork at officeVisa sponsorshipRelocation package
$150k - $190k
...software stack for engineering and manufacturing across... ...through AI inference across the entire engineering... ...new levels of optimization and automation in design... ...For As a Machine Learning Engineer in Delivery... ...remote based in the San Francisco area. This Role...Remote jobFull timeFlexible hours$165k - $230k
...re looking for exceptional Machine Learning Engineers focused on Ads to help... ...creative is generated, ranked, optimized, and ultimately performs.... ...experimentation through inference and serving. Work closely... ...hybrid role based in the San Francisco Bay Area . Team members are...Full timeWork at officeRemote workWorldwide3 days per week- ...person five days a week in our San Francisco, NYC, or London offices. About the Role As a Machine Learning Engineer on the Marketplace team,... ...demand, and the need to optimize across speed, quality, and... ...• Real-time and batch inference systems embedded in product...Full timeWork at officeRelocation package
- ...We are looking for a Machine Learning Engineer to join the growing AI and... ...model building, evaluation, optimizing performance, and ensuring... ...your time on-site in our San Francisco Office — three days per week... ...to scaling and optimizing inference and deployment Shape AI...Full timeWork at officeWorldwideFlexible hours3 days per week
- ...edge AI technology company based in San Francisco is seeking a specialist to design... ...GPU systems for high-throughput inference and model performance optimization. The ideal candidate will have... ...solid understanding of reinforcement learning technologies. Comprehensive...
$210k - $240k
...Machine Learning Engineer You'll build the ML behind Firecrawl — the models... ...across extraction quality and LLM-driven features. You'll... ...the process. Location: San Francisco, CA (Hybrid, on-site required... ...everything else Someone who optimizes for process over shipping...Full timeTemporary workFor contractorsRemote workVisa sponsorshipFlexible hours- Reactor in San Francisco is seeking an ML Inference Engineer to maximize performance of generative media models and push ultra-low-latency, high-throughput... .... You will craft an in-house runtime, implement optimizations with PyTorch tools, and collaborate with partner teams...
$197.3k - $225.1k
Lead Machine Learning Engineer At Capital One, we are creating... ...to deliver optimized ML models at scale such... ...large language model inference, similarity search,... ...introduce state-of-the-art LLM optimization... ...Learning Engineer San Francisco, CA: $215,200 - $245...Full timePart timeInternshipH1bLocal area$227.2k - $284k
...model and systems optimization, and applied research... ...the RoleAs a Staff Machine Learning Research Engineer, you will operate... ...training/fine-tuning, inference, memory and... ...training methods, LLM alignment, or applied... ...in the locations of San Francisco, New York, Seattle...Full time$270k - $315k
...committed researchers, engineers, policy experts, and... ..., implementation, and optimization of our Order-to-Cash (... ...Experience with AI/LLM integration for financial... ...Problems in AI Safety, and Learning from Human Preferences... ...headquartered in San Francisco. We offer competitive...Work at officeVisa sponsorshipFlexible hours- ...Description Principal Machine Learning Engineer, Artificial... ...operate across training, inference, evaluation, and... ...deployment, including GPU optimization, memory efficiency,... ...- Experience with LLM inference frameworks... .... Keywords: San Francisco CA Jobs, Principal...Remote workWork from home
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Machine Learning Engineer, LLM Inference Optimization in San Francisco. Be the first to apply!
- machine learning engineer San Francisco, CA
- machine learning software engineer San Francisco, CA
- ai ml engineer San Francisco, CA
- entry level machine learning engineer San Francisco, CA
- data scientist machine learning engineer San Francisco, CA
- junior machine learning research engineer San Francisco, CA
- machine learning ai engineer San Francisco, CA
- computer vision machine learning engineer San Francisco, CA
- graduate machine learning engineer San Francisco, CA
- senior ml engineer San Francisco, CA






