Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff ML Systems Engineer - Diffusion LLM Serving

Inception

Inception is seeking engineers and scientists to design, optimize, and scale the diffusion LLM serving systems powering production inference. Your work will help make inference faster, more cost-efficient, and more reliable. You will extend orchestration frameworks (Kubernetes, Ray, SLURM) for distributed inference, evaluation, and large-batch serving, and implement load balancing, autoscaling, and traffic routing for model endpoints. #J-18808-Ljbffr Inception

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Staff ML Systems Engineer - Diffusion LLM Serving in San Francisco, CA vacancy
  • $195k - $365k

     ...intersection of research and engineering, eager to design novel...  ...with building AI systems that natively understand...  ...Designing and training diffusion models, flow matching, or...  ...leveraging high‑throughput serving frameworks (e.g., vLLM, TensorRT‑LLM, SGLang) to minimize latency... 
    Suggested
    Full time
    Work at office
    Worldwide

    Plaud

    San Francisco, CA
    15 hours ago
  •  ...foundation in low-level operating systems concepts including multi-...  ...systems like TGI, vLLM, TensorRT-LLM, and Optimum, and comfortable...  ...inference systems for serving state-of-the‑art AI models...  ...contributions and staying current with ML infrastructure developments... 
    Suggested
    Work at office

    Reducto, Inc.

    San Francisco, CA
    15 hours ago
  • Senior ML Systems Engineer, Frameworks & Tooling at Cohere Our mission is to scale intelligence to serve humanity. We’re training and deploying frontier models for developers and enterprises...  ...framework responsible for large-scale LLM training. Design distributed training... 
    Suggested
    Full time
    Work at office
    Remote work
    Flexible hours

    Cohere

    San Francisco, CA
    2 days ago
  • $180k - $270k

     ...-low-latency inference engines for large language models...  ...between the core ML training team and the backend...  ...genuinely enjoy the systems-engineering challenge of...  ...experience with: Frontier Serving Frameworks: Deep, under...  ...with modern LLM serving frameworks like... 
    Suggested
    Full time
    Work at office
    Worldwide

    Plaud

    San Francisco, CA
    10 hours ago
  • $200.8k - $251k

     ...member to build and optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-time position offers a competitive salary... 
    Suggested
    Full time

    Scale AI

    San Francisco, CA
    15 hours ago
  • $264.8k - $331k

     ...Machine Learning Systems Research Engineer, Agent Post-training - Enterprise GenAI...  ...around the world. The Enterprise ML Research Lab works on the...  ...research and resources that serve all of our enterprise...  ...have: At least 1-3 years of LLM training in a production environment... 
    Full time
    Contract work
    For contractors
    For subcontractor
    Work at office

    SCALE INC

    San Francisco, CA
    3 days ago
  •  ...future of voice AI operating systems for clinicians, transforming...  ...Overview We are hiring two ML Engineers / Researchers to help build the...  ...real-time inference For LLM Researchers Experience fine...  ...benchmarks Experience serving and optimizing open-weight models... 
    Full time

    Knowtex

    San Francisco, CA
    10 hours ago
  •  ...Responsibilities As a senior Machine Learning Systems Engineer on the Search Platform team, you will...  ...and implement scalable search serving infrastructure, including retrieval pipelines...  ...search. Own end-to-end delivery of ML components from experimentation through... 
    Work at office
    Local area

    Atlassian

    San Francisco, CA
    3 days ago
  •  ...Responsibilities As a Principal Machine Learning Systems Engineer on the Search Platform team, you set the...  ...multi-year technical roadmap for search serving, vector infrastructure, and agentic...  ...priorities across the Search Platform, ML Platform, AI Gateway, and Rovo product... 
    Work at office
    Local area
    Shift work

    Atlassian

    San Francisco, CA
    10 hours ago
  •  ...out 1962 new Machine Learning Engineer opportunities posted on AI...  ...maintain scalable machine learning systems including data ingestion,...  ...and optimize end-to-end ML pipelines encompassing data...  ...behavior, and GPU and model-serving platforms for LLM inference. This role involves... 
    Flexible hours

    AI Chopping Block, Inc.

    San Francisco, CA
    4 days ago
  • $161.26k - $332.01k

     ...Pinterest is seeking a skilled Research Engineer to join our visual modeling team focusing on generative models, including text-to-image...  ...experience in computer vision and strong background in diffusion models. The role promotes collaboration with a small team, engaging... 

    Pinterest

    San Francisco, CA
    15 hours ago
  •  ...Perplexity is seeking experienced ML engineers to design, build, and optimize the recommendation systems that power core experiences on...  ...recommendation systems for the LLM era. Our goal is to combine the...  ..., feature stores, real-time serving). Background in user understanding... 
    Full time

    Perplexity®️

    San Francisco, CA
    10 hours ago
  • MakerMaker.AI is looking for a Senior Machine Learning Systems Engineer in San Francisco. In this role, you will build and operate production...  ...will have 3+ years of experience in production-grade serving infrastructure, be fluent in Python, and have strong knowledge... 

    MakerMaker.AI

    San Francisco, CA
    1 day ago
  • Boomtrain is seeking a talented Machine Learning Engineer for the Personalization team in San Francisco, California. You will play...  ...with responsibilities including model generation pipelines and serving systems. A degree in Computer Science or Statistics and at least 4... 

    Boomtrain

    San Francisco, CA
    10 hours ago
  •  ...with safety, privacy, and real-world responsibility in mind. Our ML team comes from a culture of academic research driven to democratize...  ...’ adoption of frontier research for their next generation of LLM products. Join us if you: Wish to work on the premier platform... 
    Local area
    Shift work

    Capitolis

    San Francisco, CA
    2 days ago
  • $198k - $230k

    CreatorIQ is the operating system for creator‑led growth trusted by...  ...individual work styles. Senior MLOps Engineer (Applied AI Focus) As a Senior...  ...Innovations team, you will serve as the technical lead for...  ...model can do the job of a massive LLM. Cloud & Infrastructure... 
    Work at office
    Remote work
    Work from home
    Worldwide
    Home office
    Flexible hours

    CreatorIQ

    San Francisco, CA
    10 hours ago
  • $124.8k - $220.8k

     ...other locations. The Machine Learning (ML) Practice team is a specialized customer-...  ...demand for Large Language Model (LLM)-based solutions. We deliver professional...  ...long-term initiatives working alongside engineering, product, and developer relations, and internal... 
    Work at office
    Remote work
    Work from home
    Home office
    Flexible hours

    Databricks, Inc.

    San Francisco, CA
    15 hours ago
  •  ...Stealth Startup is hiring a Founding Machine Learning Engineer to design, train, deploy, and monitor production ML systems that fuse LLM-powered agents with time-series models. You will shape how agents interact with multimodal data, build scalable workflows, and drive... 

    Stealth Startup

    San Francisco, CA
    15 hours ago
  •  ...Jobzhr in the San Francisco Bay Area is hiring a Founding Machine Learning Engineer to design and deploy production ML systems that fuse LLM-powered agents with time-series models. You will lead end-to-end development—from data ingestion and preprocessing to deployment... 

    Jobzhr

    San Francisco, CA
    15 hours ago
  •  ...We’re doubling down on ML as the future of Grindr...  ...building foundational systems surrounded by high‑impact...  ...systems to serve millions, balancing performance...  ...teams, collaborating with engineering, data science and product...  ...and maintaining LLM workflows for nuanced,... 
    Casual work
    Work at office
    Immediate start
    Flexible hours

    Grindr LLC

    San Francisco, CA
    15 hours ago
  • $200k - $260k

     ...voice agents and applications — serving speech-to-text and text-to-...  ...We're looking for a Senior ML Engineer to drive the model serving layer...  ...inference engines like TRT-LLM and SGLang to optimize how we...  ...(SNAC), and speech-to-speech systems. Collaborate with model partners... 
    Full time

    Together Ai

    San Francisco, CA
    10 hours ago
  •  ...across hospital and health systems, pharmacies and payors...  ..., and we’re proud to serve 90+ leading health systems...  ...for a Machine Learning Engineer to design, build, and deploy production-grade ML systems that power the next...  ...using modern NLP, LLM, classification, recommendation... 
    Full time
    Work at office
    Remote work
    Flexible hours
    2 days per week

    Plenful

    San Francisco, CA
    10 hours ago
  •  ...Description As a Machine Learning Engineer at Advex, you will play a...  ...you are well versed with the ML life-cycle process, collect...  ...distributions # Controllability of Diffusion Models # Evaluation of...  ...engineering, model tuning, and model serving ~ Technical expertise... 
    Full time

    Openreq

    San Francisco, CA
    10 hours ago
  •  ...proprietary, high-efficiency serving platform. Backed by multi-million...  ...hands-on support from AMD engineers the team is scaling rapidly to...  ...applications. About the role As an ML Engineer at Sciforium, you...  ...end-to-end multimodal GenAI systems. In this role, you will build... 
    Full time
    Flexible hours

    Sciforium

    San Francisco, CA
    3 hours ago
  • $213k - $263k

     ...seeking visionary machine learning engineers and researchers to architect the scalable deep learning systems, novel data workflows, and...  ...-scale generative models (diffusion models, flow matching, vision...  ...clusters for efficient production serving. Professional experience... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    10 hours ago
  •  ...page! Machine Learning Engineer @ Clay Clay's...  ...uses it. This means data, ML, and AI are at the heart...  ...heart of the product: systems that learn a customer's...  ...data lake foundations and serving infrastructure....  ...designing eval frameworks for LLM or ML systems Familiarity... 
    Full time

    Clay Labs

    San Francisco, CA
    10 hours ago
  • $225k - $300k

     ...meaningful ROI for health systems across the country....  ...Senior Machine Learning Engineer at Ambience , you will...  ...evaluation pipelines for LLM and agentic systems,...  ...evaluation, orchestration, serving, and observability,...  ...5+ years in production ML, research engineering,... 
    Full time
    Work at office
    Immediate start
    Remote work
    Flexible hours
    3 days per week

    Ambience Healthcare

    San Francisco, CA
    10 hours ago
  • $204k - $259k

     ...states. The Perception team builds the system which learns the spatial-temporal...  ...from a diverse set of sensors, enabling engineers like you to (1) develop methods for efficiently...  ...focus on large-scale model development (LLM, VLM, or similar foundation models). ~... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    10 hours ago
  • $308k - $423.5k

     ...role: We are seeking a  Principal ML / AI Engineer to be a  company-level technical...  ...unblock – Build and lead deployment of AI systems (LLM fine-tuning, RLHF, agent frameworks, etc...  ...warehouses, feature stores, and model serving — to ensure our infrastructure is AI-... 
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    Monday to Friday
    Flexible hours
    3 days per week

    Faire Faire

    San Francisco, CA
    10 hours ago
  • $174.5k - $240k

     ...come join ours.About this roleGTM Engineering builds and operates the intelligent systems, integrations, and automations that...  ...Working fluency with AI tooling — LLM APIs, agent/orchestration...  ...You'll own meaningful problems that serve customers around the globe with the... 
    Work experience placement
    Work at office
    Local area
    Remote work
    Monday to Friday
    Flexible hours
    3 days per week

    Faire

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff ML Systems Engineer - Diffusion LLM Serving. Be the first to apply!