Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff ML Systems Engineer - Diffusion LLM Serving

Inception

Inception is seeking engineers and scientists to design, optimize, and scale the diffusion LLM serving systems powering production inference. Your work will help make inference faster, more cost-efficient, and more reliable. You will extend orchestration frameworks (Kubernetes, Ray, SLURM) for distributed inference, evaluation, and large-batch serving, and implement load balancing, autoscaling, and traffic routing for model endpoints. #J-18808-Ljbffr Inception

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Staff ML Systems Engineer - Diffusion LLM Serving in San Francisco, CA vacancy
  • Gravity Engineering Services Pvt Ltd. is seeking candidates to enhance their ML platform (RLXF). In this role, you will build and optimize frameworks...  ...position requires excitement for system optimization and experience with multi-node LLM training. If you are passionate... 
    Suggested

    Gravity Engineering Services Pvt Ltd.

    San Francisco, CA
    13 hours ago
  • TRM Labs, an AI-powered intelligence company, seeks a Senior or Staff ML Systems Engineer - LLM to scale AI infrastructure and production-grade ML tooling. You will build modular stacks, embed models into real-time workflows, and evaluate state-of-the-art AI tools to ensure... 
    Suggested

    Trm-Labs

    San Francisco, CA
    4 days ago
  • Senior ML Systems Engineer, Frameworks & Tooling at Cohere Our mission is to scale intelligence to serve humanity. We’re training and deploying frontier models for developers and enterprises...  ...framework responsible for large-scale LLM training. Design distributed training... 
    Suggested
    Full time
    Work at office
    Remote work
    Flexible hours

    Cohere

    San Francisco, CA
    13 hours ago
  •  ...in low-level operating systems concepts including multi...  ...TGI , vLLM , TensorRT-LLM , and Optimum , and comfortable...  ...inference systems for serving state-of-the-art AI...  ...staying current with ML infrastructure...  ...usually requires a large engineering effort dedicated to building... 
    Suggested
    Work at office

    Gravity Engineering Services Pvt Ltd.

    San Francisco, CA
    3 days ago
  • $200k - $275k

     ...and more secure. The AI Engineering Team is chartered with...  ...(LLMs) and agentic systems. Our mission is to build...  ...petabyte-scale pipelines, serve models with millisecond...  ...-edge tools in the LLM and agent space — including...  .... As a Senior or Staff ML Systems Engineer - LLM... 
    Suggested
    Remote work
    Worldwide

    Trm-Labs

    San Francisco, CA
    4 hours ago
  • $200.8k - $251k

     ...member to build and optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-time position offers a competitive salary... 
    Full time

    Scale AI

    San Francisco, CA
    1 day ago
  • $218.4k - $273k

     ...research in Physical AI and developing ML pipelines for processing, training, and...  ...evaluation for Physical AI. The Role As an ML Systems Engineer on the Physical AI team, you will design...  ...for scalable, reliable, and efficient serving of foundation models specifically... 
    Full time

    Scale AI

    San Francisco, CA
    13 hours ago
  •  ...research in Physical AI and developing ML pipelines for processing, training, and...  ...model evaluation for Physical AI. As an ML Systems Engineer on the Physical AI team, you will design...  ...for scalable, reliable, and efficient serving of foundation models specifically tailored... 

    Gravity Engineering Services Pvt Ltd.

    San Francisco, CA
    13 hours ago
  • Gravity Engineering Services Pvt Ltd. is looking for a skilled ML Systems Engineer to drive research and development in Physical AI. You'll design and build platforms for scalable and efficient model serving, enhancing both research and production systems for autonomous... 

    Gravity Engineering Services Pvt Ltd.

    San Francisco, CA
    13 hours ago
  • $150k - $300k

     ...specialist to build and optimize machine learning systems as part of a hybrid team. This role...  ...on developing efficient architecture for serving LLMs and optimizing performance using...  ...candidates will have significant experience with ML systems, ensuring robust performance and... 
    Remote job

    Prime-Intellect

    San Francisco, CA
    13 hours ago
  • $200.8k - $251k

    Overview Scale's ML platform (RLXF) team builds our...  ...our next generation of LLM training, inference and...  ...to optimize our ML system Ideally you'd have: Strong...  ...Strong software engineering skills, proficient in frameworks...  ..." means a veteran who served on active duty in the U... 
    Full time
    For contractors
    Shift work

    Scale AI

    San Francisco, CA
    13 hours ago
  •  ...and real‑world responsibility in mind. Our ML team comes from a culture of academic...  ...frontier research for their next generation of LLM products. Join us if you: Wish to work...  ...inference attacks. Collaborate with our engineering team to deliver real‑world applications... 
    Full time
    Shift work

    DynamoFL

    San Francisco, CA
    2 days ago
  • $100k - $130k

    ML Engineer - LLM Evaluation at Dynamo AI (W22) Compliant-Ready AI for the Enterprise San Francisco, CA Full-time US citizen/visa only 3+ years The enterprise platform for enabling private, secure, and regulation-compliant Gen AI models. About the role At Dynamo AI,... 
    Full time
    Local area
    Shift work

    DynamoFL

    San Francisco, CA
    2 days ago
  • $161.26k - $332.01k

    Pinterest is seeking a skilled Research Engineer to join our visual modeling team focusing on generative models, including text-to-image...  ...experience in computer vision and strong background in diffusion models. The role promotes collaboration with a small team, engaging... 

    Pinterest

    San Francisco, CA
    13 hours ago
  •  ...out 1962 new Machine Learning Engineer opportunities posted on AI...  ...maintain scalable machine learning systems including data ingestion,...  ...and optimize end-to-end ML pipelines encompassing data...  ...behavior, and GPU and model-serving platforms for LLM inference. This role involves... 
    Flexible hours

    AI Chopping Block, Inc.

    San Francisco, CA
    13 hours ago
  • Boomtrain is seeking a talented Machine Learning Engineer for the Personalization team in San Francisco, California. You will play...  ...with responsibilities including model generation pipelines and serving systems. A degree in Computer Science or Statistics and at least 4... 

    Boomtrain

    San Francisco, CA
    3 days ago
  • $198k - $230k

    CreatorIQ is the operating system for creator‑led growth trusted by...  ...individual work styles. Senior MLOps Engineer (Applied AI Focus) As a Senior...  ...Innovations team, you will serve as the technical lead for...  ...model can do the job of a massive LLM. Cloud & Infrastructure... 
    Work at office
    Remote work
    Work from home
    Worldwide
    Home office
    Flexible hours

    CreatorIQ

    San Francisco, CA
    3 days ago
  • $204k - $259k

     ...technology company in San Francisco is seeking a Machine Learning Engineer to optimize and develop large-scale model training pipelines for...  ...of experience in Machine Learning, with strong expertise in LLM/VLM pre-training. This position offers a competitive salary range... 

    Waymo

    San Francisco, CA
    13 hours ago
  •  ...Principal / Distinguished AI/ML Researcher and/or Engineer with deep experience in...  ...planning, and decision-making systems . This role is ideal for...  ...Interfaces . Advance techniques in LLM/LRM post-training,...  ...AI/ML platform strategy . Serve as the technical conscience... 
    Local area

    Gravity Engineering Services Pvt Ltd.

    San Francisco, CA
    1 day ago
  • $275k

     ...SF) We’re hiring a Senior ML Infrastructure / Backend Engineer to join a well-funded AI company...  ...is developing production systems that bring AI characters...  ...systems that serve ML-powered functionality at...  ...Exposure to generative models (diffusion or transformer-based systems... 
    Work at office

    Acceler8 Talent

    San Francisco, CA
    13 hours ago
  •  ...We’re doubling down on ML as the future of Grindr...  ...building foundational systems surrounded by high‑impact...  ...systems to serve millions, balancing performance...  ...teams, collaborating with engineering, data science and product...  ...and maintaining LLM workflows for nuanced,... 
    Casual work
    Work at office
    Immediate start
    Flexible hours

    Grindr LLC

    San Francisco, CA
    13 hours ago
  • Inception in San Francisco is seeking experienced engineers and scientists to develop the evaluation metrics and systems that drive frontier LLM performance. You will design the frameworks that tell us whether our models are improving and ensure they perform reliably at... 

    Inception LLC

    San Francisco, CA
    2 days ago
  •  ...the lookout for a Senior Machine Learning Engineer in San Francisco, California. The ideal...  ...and optimizing production-grade ML systems. This role encompasses end-to-end ML system...  ...ownership, from data pipelines to model serving, requiring expertise in MLOps, Python, and... 

    Dormont Manufacturing Co

    San Francisco, CA
    2 days ago
  •  ...company is seeking experienced ML engineers to design, build, and optimize the recommendation systems that power core experiences on...  ...recommendation systems for the LLM era. Our goal is to combine the...  ...ranking, feature stores, real‑time serving). Background in user... 

    United States Digital Space LLC

    San Francisco, CA
    2 days ago
  • $102k - $182.71k

     ...Overview We’re seeking a Senior Search Systems Engineer to build the intelligence and automation...  ...driven discovery environments. This role serves as the bridge between SEO and AEO strategy...  ...tools, APIs, and frameworks related to LLM analysis, AI agents, search intelligence... 
    Shift work

    Autodesk

    San Francisco, CA
    4 days ago
  • $245k - $272k

     ...reflects the people we serve. All full‑time...  ...About the Role As a Staff MLE, you’ll design,...  ...decisions, and mentor engineers across the...  ...high‑impact group of ML engineers and platform...  ...deploying end‑to‑end AI/ML systems, with experience...  ...building LLM‑based applications... 
    Full time
    Work at office
    Local area
    Remote work
    2 days per week
    3 days per week

    Gusto

    San Francisco, CA
    4 days ago
  • $159.18k - $295.62k

     ...Senior Machine Learning Engineer – Data & Audience...  ...between our Senior MLE and Staff MLE levels. The Engineer...  ...delivery of production ML systems end‑to‑end, drives technical...  ...decisions, and serves as a US‑based technical...  ...LangChain/LangGraph, evaluate LLM‑based approaches for... 
    Temporary work

    Warner Bros. Discovery

    San Francisco, CA
    1 day ago
  • $208k - $263.5k

     ..., mining, and transport. Our systems are designed to understand, predict...  ...scale. We are roboticists, engineers, operators, and builders. We...  ...VLA models, or generative Diffusion models. Strong background in...  ...Our work only matters if it serves others, and we know that meaningful... 
    Full time
    Internship
    Work at office
    Flexible hours

    ATOMS Careers page

    San Francisco, CA
    4 days ago
  • $96.8k - $306.4k

     ...Senior Principal AI Agent / ML Software Engineer is a Senior Staff-level, hands‑on technical...  ...next-generation AI systems on Oracle Cloud Infrastructure...  ...organization. Responsibilities Serve as a senior technical...  ...ecosystems. Deep understanding of LLM application patterns,... 
    Temporary work
    Flexible hours

    Oracle

    San Francisco, CA
    4 days ago
  • $300k - $400k

    About the Role Principal AI/ML Engineer - AdTech Drive the development...  ...and broader ad‑tech stack. System Architecture: Build end‑to‑end...  ...seamless integration with ad‑serving architecture. Strategic Road...  ...the cutting edge. Agentic & LLM Applications: Develop AI... 

    Zeta Global

    San Francisco, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff ML Systems Engineer - Diffusion LLM Serving. Be the first to apply!