ML Systems Engineer: Distributed LLM Training & Inference
$200.8k - $251kScale AI
A leading AI technology company in San Francisco seeks a team member to build and optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-time position offers a competitive salary range of $200,800 - $251,000, along with comprehensive benefits.#J-18808-Ljbffr
Vacancy posted 12 hours ago
Similar jobs that could be interesting for youBased on the ML Systems Engineer: Distributed LLM Training & Inference in San Francisco, CA vacancy
$180k - $270k
...throughput, ultra-low-latency inference engines for large language... ...between the core ML training team and the backend... ...genuinely enjoy the systems-engineering challenge... ...familiarity with modern LLM serving frameworks... ...accuracy. Large-Scale Distributed Systems: Deploying...TrainingFull timeWork at officeWorldwide- ...looking for a Senior Software Engineer to build scalable infrastructure for large‑scale training and fine-tuning of foundation models. You will design distributed training systems and optimize GPU utilization... ...over 5 years of experience in ML infrastructure and a strong...Training
- Genesis AI in San Francisco is seeking a senior ML infrastructure engineer to design and optimize distributed training systems and performance-critical components. You will profile bottlenecks, implement low‑level code (CUDA, Triton) and ensure efficient hardware utilization...Training
- ...foundation in low-level operating systems concepts including multi-... ...experienced with modern inference systems like TGI, vLLM, TensorRT-LLM, and Optimum, and... ...and staying current with ML infrastructure developments... ...computing and distributed systems Have worked in...SuggestedWork at office
- Inception is seeking engineers and scientists to design, optimize, and scale the diffusion LLM serving systems powering production inference. Your work will help make inference faster, more... ...(Kubernetes, Ray, SLURM) for distributed inference, evaluation, and large-batch...Suggested
$189.6k - $237k
Scale’s ML platform (RLXF) team builds our internal distributed framework for large language model training and inference. The platform has been powering... ...evaluation of LLM's, as well as... ...about system optimizationExperience... ...systemsStrong software engineering skills,...TrainingFull time- Senior ML Systems Engineer, Frameworks & Tooling at Cohere Our mission is to... ...to serve humanity. We’re training and deploying frontier models... ...intersection of large‑scale training, distributed systems, and HPC... ...responsible for large-scale LLM training. Design distributed...TrainingFull timeWork at officeRemote workFlexible hours
$195k - $365k
...of building and training large-scale audio... ...of research and engineering, eager to design... ...one day and debug distributed training clusters... ...with building AI systems that natively understand... ...: End‑to‑end inference and performance... ..., vLLM, TensorRT‑LLM, SGLang) to minimize...TrainingFull timeWork at officeWorldwide$200k - $240k
...world for all. The AI Engineering Team is chartered... ...LLMs) and agentic systems. Our mission is to... ...edge tools in the LLM and agent space —... ...a Senior or Staff ML Systems Engineer -... ...for model training, evaluation, and deployment... ...TRM operates as a distributed-first company with...TrainingRemote workWorldwide$124.8k - $220.8k
...The Machine Learning (ML) Practice team is a specialized... ...Large Language Model (LLM)-based solutions. We... ...working alongside engineering, product, and developer... ...to process large-scale distributed datasets Benefits... ...relevant certifications and training, and specific work...TrainingWork at officeRemote workWork from homeHome officeFlexible hours- ...voice AI operating systems for clinicians,... ...We are hiring two ML Engineers / Researchers to help... ...systems, train and fine-tune models... ...latency, real-time inference at production scale... ...scale datasets and distributed training environments... ...inference For LLM Researchers Experience...TrainingFull time
- ...practical constraints of robotic platforms. About the Role As a Research Engineer, Distributed Data Systems, you will design and scale the infrastructure that powers large-scale multimodal training and evaluation at OpenAI. You’ll manage distributed data pipelines,...TrainingFull timeWork at officeRelocation package
- ...of Technical Staff to design and operate distributed systems for serving models in production and driving large-scale post-training workflows. You will work where model execution... ...will own the infrastructure enabling fast inference and scalable RL iteration, balancing KV-...Training
$298k - $368k
...Perception team builds the system which learns the... ...set of sensors, enabling engineers like you to (1) develop... ...develop models and model training at scale, to (3)... ...will: Design VLM/LLM model architecture and... ...low-latency on-device inference techniques and a deep understanding...TrainingFull timeRemote work- ...Machine Learning Engineer opportunities... ...machine learning systems including data... ...preprocessing, training, testing, and... ...optimize end-to-end ML pipelines... ...tuning training and inference end-to-end for... ..., PyTorch, distributed systems, GPUs,... ...platforms for LLM inference. This...TrainingFlexible hours
$110 per hour
...Dorsey . Position: MLOps Engineer (JAX, PyTorch, Pallas/... ...performance in MLOps , training infrastructure, and ML framework-level topics .... ...solutions to MLOps and ML systems problems . Evaluate... ...training pipeline design, distributed systems reasoning, and kernel...TrainingRemote jobContract workSummer workWeekday work- ...leading AI research firm located in San Francisco is seeking a Senior ML Systems Engineer to build and maintain the training framework for large-scale language models. The role involves designing distributed training solutions and improving training throughput across multi-...TrainingFlexible hours
- ...a Member of Technical Staff to design and optimize inference systems. The role involves managing KV cache allocation and... ...components. Ideal candidates should have strong software engineering skills and experience with ML inference systems, particularly in Python and C++....
$264.8k - $331k
...Machine Learning Systems Research Engineer, Agent Post-training - Enterprise GenAI AI is becoming... ...the world. The Enterprise ML Research Lab works on the... ...our training and inference framework. Post-train state... ...have: At least 1-3 years of LLM training in a production...TrainingFull timeContract workFor contractorsFor subcontractorWork at office- ...seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate will have hands-on experience with modern inference...Training
- Reflection is seeking a role to build and operate distributed training systems powering frontier models in SF. You will work with research teams to... ...training pipelines. This role emphasizes collaboration with ML researchers, debugging across GPU stacks, and improving...Training
$227.2k - $417k
...Role:As a Software Engineer on the ML Infrastructure... ...machine learning inference platforms. These platforms... ...ML model serving systems that support Deep Learning, LLM, and Search models... ..., and low latency distributed systems using... ...ElastiCache, model training orchestration, etc...TrainingFull timeTemporary workLocal areaFlexible hours- ...boundaries of what our ML systems can do. We're hiring a Founding ML Engineer to own the... ...will be researching, training, and shipping models... ...raw people data, infer the org chart —... ...transfer Scaled LLM inference pipelines... ...Experience with distributed training on GPU clusters...Training
- ...OpenAI is seeking a Research Engineer, Distributed Data Systems, to design and scale infrastructure powering large-scale multimodal training and evaluation. You will manage distributed data pipelines and partner with researchers to translate requirements into robust systems...TrainingRelocation package
- ...About the role: As a ML Engineer, you’ll build and operate... ...engineering. You’ll work on the systems that support training and evaluating large... ...work on: Optimize inference and training throughput for... ...high-performance distributed training infrastructure...TrainingFull timeInternship
- ...research and infra to prototype, train, and deploy state-of-the-... ...— scale training and inference for LLM-class workloads; chase latency... .... Proven software engineer who loves ML; comfortable writing production... ...especially user-facing, online ML systems—despite shifting...TrainingFull timeContract workFlexible hoursShift work
- ...support from AMD engineers the team is scaling... ...the role As an ML Engineer at... ...multimodal GenAI systems. In this role, you... ...across Serving, Post-Training and Agentic frameworks... ...with one or more distributed ML training frameworks... ...JAX, or Ray and inference engines like...TrainingFull timeFlexible hours
- ...hospital and health systems, pharmacies and... ...Learning Engineer to design, build... ...production-grade ML systems that... ...pipelines for training, evaluation, monitoring, and inference Build intelligent... ...modern NLP, LLM, classification... ...using SQL and distributed data processing...TrainingFull timeWork at officeRemote workFlexible hours2 days per week
$204k - $259k
...The Perception team builds the system which learns the spatial-... ...diverse set of sensors, enabling engineers like you to (1) develop methods... ...(2) develop models and model training at scale, to (3) analyze real-... ...large-scale model development (LLM, VLM, or similar foundation models...TrainingFull timeRemote work- ...Machine Learning Engineer to build and ship... ...consumer-facing AI systems that power... ...Build and deploy ML models that improve... ...product workflows (LLM + tools/RAG, multimodal... ...models: scalable training/inference pipelines, model... ...tooling (SQL, distributed compute such as...TrainingFull timeImmediate startWorldwideNight shift
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Systems Engineer: Distributed LLM Training & Inference. Be the first to apply!
Related searches
- ai ml engineer San Francisco, CA
- senior ml engineer San Francisco, CA
- computer vision machine learning engineer San Francisco, CA
- machine learning engineer San Francisco, CA
- machine learning ai engineer San Francisco, CA
- junior machine learning research engineer San Francisco, CA
- data scientist machine learning engineer San Francisco, CA
- graduate machine learning engineer San Francisco, CA
- machine learning software engineer San Francisco, CA
- senior windows systems engineer San Francisco, CA




