Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Systems Engineer: Distributed LLM Training & Inference

$200.8k - $251k

Scale AI

A leading AI technology company in San Francisco seeks a team member to build and optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-time position offers a competitive salary range of $200,800 - $251,000, along with comprehensive benefits.#J-18808-Ljbffr

Vacancy posted 12 hours ago
Similar jobs that could be interesting for youBased on the ML Systems Engineer: Distributed LLM Training & Inference in San Francisco, CA vacancy
  • $180k - $270k

     ...throughput, ultra-low-latency inference engines for large language...  ...between the core ML training team and the backend...  ...genuinely enjoy the systems-engineering challenge...  ...familiarity with modern LLM serving frameworks...  ...accuracy. Large-Scale Distributed Systems: Deploying... 
    Training
    Full time
    Work at office
    Worldwide

    Plaud

    San Francisco, CA
    7 hours ago
  •  ...looking for a Senior Software Engineer to build scalable infrastructure for large‑scale training and fine-tuning of foundation models. You will design distributed training systems and optimize GPU utilization...  ...over 5 years of experience in ML infrastructure and a strong... 
    Training

    Baseten

    San Francisco, CA
    12 hours ago
  • Genesis AI in San Francisco is seeking a senior ML infrastructure engineer to design and optimize distributed training systems and performance-critical components. You will profile bottlenecks, implement low‑level code (CUDA, Triton) and ensure efficient hardware utilization... 
    Training

    Genesis AI

    San Francisco, CA
    5 days ago
  •  ...foundation in low-level operating systems concepts including multi-...  ...experienced with modern inference systems like TGI, vLLM, TensorRT-LLM, and Optimum, and...  ...and staying current with ML infrastructure developments...  ...computing and distributed systems Have worked in... 
    Suggested
    Work at office

    Reducto, Inc.

    San Francisco, CA
    12 hours ago
  • Inception is seeking engineers and scientists to design, optimize, and scale the diffusion LLM serving systems powering production inference. Your work will help make inference faster, more...  ...(Kubernetes, Ray, SLURM) for distributed inference, evaluation, and large-batch... 
    Suggested

    Inception

    San Francisco, CA
    4 days ago
  • $189.6k - $237k

    Scale’s ML platform (RLXF) team builds our internal distributed framework for large language model training and inference. The platform has been powering...  ...evaluation of LLM's, as well as...  ...about system optimizationExperience...  ...systemsStrong software engineering skills,... 
    Training
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  • Senior ML Systems Engineer, Frameworks & Tooling at Cohere Our mission is to...  ...to serve humanity. We’re training and deploying frontier models...  ...intersection of large‑scale training, distributed systems, and HPC...  ...responsible for large-scale LLM training. Design distributed... 
    Training
    Full time
    Work at office
    Remote work
    Flexible hours

    Cohere

    San Francisco, CA
    2 days ago
  • $195k - $365k

     ...of building and training large-scale audio...  ...of research and engineering, eager to design...  ...one day and debug distributed training clusters...  ...with building AI systems that natively understand...  ...: End‑to‑end inference and performance...  ..., vLLM, TensorRT‑LLM, SGLang) to minimize... 
    Training
    Full time
    Work at office
    Worldwide

    Plaud

    San Francisco, CA
    12 hours ago
  • $200k - $240k

     ...world for all. The AI Engineering Team is chartered...  ...LLMs) and agentic systems. Our mission is to...  ...edge tools in the LLM and agent space —...  ...a Senior or Staff ML Systems Engineer -...  ...for model training, evaluation, and deployment...  ...TRM operates as a distributed-first company with... 
    Training
    Remote work
    Worldwide

    TRM Labs

    San Francisco, CA
    6 days ago
  • $124.8k - $220.8k

     ...The Machine Learning (ML) Practice team is a specialized...  ...Large Language Model (LLM)-based solutions. We...  ...working alongside engineering, product, and developer...  ...to process large-scale distributed datasets Benefits...  ...relevant certifications and training, and specific work... 
    Training
    Work at office
    Remote work
    Work from home
    Home office
    Flexible hours

    Databricks, Inc.

    San Francisco, CA
    12 hours ago
  •  ...voice AI operating systems for clinicians,...  ...We are hiring two ML Engineers / Researchers to help...  ...systems, train and fine-tune models...  ...latency, real-time inference at production scale...  ...scale datasets and distributed training environments...  ...inference For LLM Researchers Experience... 
    Training
    Full time

    Knowtex

    San Francisco, CA
    7 hours ago
  •  ...practical constraints of robotic platforms. About the Role As a Research Engineer, Distributed Data Systems, you will design and scale the infrastructure that powers large-scale multimodal training and evaluation at OpenAI. You’ll manage distributed data pipelines,... 
    Training
    Full time
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    7 hours ago
  •  ...of Technical Staff to design and operate distributed systems for serving models in production and driving large-scale post-training workflows. You will work where model execution...  ...will own the infrastructure enabling fast inference and scalable RL iteration, balancing KV-... 
    Training

    Magic AI, Inc

    San Francisco, CA
    4 days ago
  • $298k - $368k

     ...Perception team builds the system which learns the...  ...set of sensors, enabling engineers like you to (1) develop...  ...develop models and model training at scale, to (3)...  ...will: Design VLM/LLM model architecture and...  ...low-latency on-device inference techniques and a deep understanding... 
    Training
    Full time
    Remote work

    Waymo

    San Francisco, CA
    7 hours ago
  •  ...Machine Learning Engineer opportunities...  ...machine learning systems including data...  ...preprocessing, training, testing, and...  ...optimize end-to-end ML pipelines...  ...tuning training and inference end-to-end for...  ..., PyTorch, distributed systems, GPUs,...  ...platforms for LLM inference. This... 
    Training
    Flexible hours

    AI Chopping Block, Inc.

    San Francisco, CA
    4 days ago
  • $110 per hour

     ...Dorsey . Position: MLOps Engineer (JAX, PyTorch, Pallas/...  ...performance in MLOps , training infrastructure, and ML framework-level topics ....  ...solutions to MLOps and ML systems problems . Evaluate...  ...training pipeline design, distributed systems reasoning, and kernel... 
    Training
    Remote job
    Contract work
    Summer work
    Weekday work

    Mercor

    San Francisco, CA
    3 days ago
  •  ...leading AI research firm located in San Francisco is seeking a Senior ML Systems Engineer to build and maintain the training framework for large-scale language models. The role involves designing distributed training solutions and improving training throughput across multi-... 
    Training
    Flexible hours

    Cohere

    San Francisco, CA
    2 days ago
  •  ...a Member of Technical Staff to design and optimize inference systems. The role involves managing KV cache allocation and...  ...components. Ideal candidates should have strong software engineering skills and experience with ML inference systems, particularly in Python and C++.... 

    Gimlet Labs

    San Francisco, CA
    2 days ago
  • $264.8k - $331k

     ...Machine Learning Systems Research Engineer, Agent Post-training - Enterprise GenAI AI is becoming...  ...the world. The Enterprise ML Research Lab works on the...  ...our training and inference framework. Post-train state...  ...have: At least 1-3 years of LLM training in a production... 
    Training
    Full time
    Contract work
    For contractors
    For subcontractor
    Work at office

    SCALE INC

    San Francisco, CA
    3 days ago
  •  ...seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate will have hands-on experience with modern inference... 
    Training

    Reflection AI

    San Francisco, CA
    4 days ago
  • Reflection is seeking a role to build and operate distributed training systems powering frontier models in SF. You will work with research teams to...  ...training pipelines. This role emphasizes collaboration with ML researchers, debugging across GPU stacks, and improving... 
    Training

    Visa Hunt

    San Francisco, CA
    6 days ago
  • $227.2k - $417k

     ...Role:As a Software Engineer on the ML Infrastructure...  ...machine learning inference platforms. These platforms...  ...ML model serving systems that support Deep Learning, LLM, and Search models...  ..., and low latency distributed systems using...  ...ElastiCache, model training orchestration, etc... 
    Training
    Full time
    Temporary work
    Local area
    Flexible hours

    Tubi TV

    San Francisco, CA
    2 days ago
  •  ...boundaries of what our ML systems can do. We're hiring a Founding ML Engineer to own the...  ...will be researching, training, and shipping models...  ...raw people data, infer the org chart —...  ...transfer Scaled LLM inference pipelines...  ...Experience with distributed training on GPU clusters... 
    Training

    Crustdata (YC F24)

    San Francisco, CA
    2 days ago
  •  ...OpenAI is seeking a Research Engineer, Distributed Data Systems, to design and scale infrastructure powering large-scale multimodal training and evaluation. You will manage distributed data pipelines and partner with researchers to translate requirements into robust systems... 
    Training
    Relocation package

    Neura Market

    San Francisco, CA
    12 hours ago
  •  ...About the role: As a ML Engineer, you’ll build and operate...  ...engineering. You’ll work on the systems that support training and evaluating large...  ...work on: Optimize inference and training throughput for...  ...high-performance distributed training infrastructure... 
    Training
    Full time
    Internship

    Tilde Research

    San Francisco, CA
    7 hours ago
  •  ...research and infra to prototype, train, and deploy state-of-the-...  ...— scale training and inference for LLM-class workloads; chase latency...  .... Proven software engineer who loves ML; comfortable writing production...  ...especially user-facing, online ML systems—despite shifting... 
    Training
    Full time
    Contract work
    Flexible hours
    Shift work

    Sesame, L.l.c.

    San Francisco, CA
    7 hours ago
  •  ...support from AMD engineers the team is scaling...  ...the role As an ML Engineer at...  ...multimodal GenAI systems. In this role, you...  ...across Serving, Post-Training and Agentic frameworks...  ...with one or more distributed ML training frameworks...  ...JAX, or Ray and inference engines like... 
    Training
    Full time
    Flexible hours

    Sciforium

    San Francisco, CA
    19 minutes ago
  •  ...hospital and health systems, pharmacies and...  ...Learning Engineer to design, build...  ...production-grade ML systems that...  ...pipelines for training, evaluation, monitoring, and inference Build intelligent...  ...modern NLP, LLM, classification...  ...using SQL and distributed data processing... 
    Training
    Full time
    Work at office
    Remote work
    Flexible hours
    2 days per week

    Plenful

    San Francisco, CA
    7 hours ago
  • $204k - $259k

     ...The Perception team builds the system which learns the spatial-...  ...diverse set of sensors, enabling engineers like you to (1) develop methods...  ...(2) develop models and model training at scale, to (3) analyze real-...  ...large-scale model development (LLM, VLM, or similar foundation models... 
    Training
    Full time
    Remote work

    Waymo

    San Francisco, CA
    7 hours ago
  •  ...Machine Learning Engineer to build and ship...  ...consumer-facing AI systems that power...  ...Build and deploy ML models that improve...  ...product workflows (LLM + tools/RAG, multimodal...  ...models: scalable training/inference pipelines, model...  ...tooling (SQL, distributed compute such as... 
    Training
    Full time
    Immediate start
    Worldwide
    Night shift

    Eight Sleep

    San Francisco, CA
    7 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Systems Engineer: Distributed LLM Training & Inference. Be the first to apply!