ML Systems Engineer: Distributed LLM Training & Inference
$200.8k - $251kScale AI
A leading AI technology company in San Francisco seeks a team member to build and optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-time position offers a competitive salary range of $200,800 - $251,000, along with comprehensive benefits. #J-18808-Ljbffr Scale AI
- Fintal Partners in New York is hiring engineers to build the distributed training and inference systems that take models from research into production. You’ll work side-by-side with researchers and traders on a trading floor-like environment, owning critical, IC-heavy infrastructure...Training
- General Information Job Title ML Staff Engineer - LLM & Production Systems Job ID 107242 Work Areas... ...strategies and optimize inference workloads through capacity planningOwn... ...education, licensure/certifications, training, and skill level.In NYC, NY, the good...TrainingPermanent employmentFull timeWork at officeLocal area1 day per week
- ...business problems.We’re training and deploying... ...are building AI systems. We believe that our... ...of researchers, engineers, designers, and more... ...-scale training, distributed systems, and HPC infrastructure... ...the full stack of ML systems, this role... ...for large-scale LLM training.Design...TrainingFull timeWork at officeLocal areaRemote workHome office
$189.6k - $237k
Scale’s ML platform (RLXF) team builds our internal distributed framework for large language model training and inference. The platform has been powering... ...evaluation of LLM's, as well as... ...about system optimizationExperience... ...systemsStrong software engineering skills,...TrainingFull time- ...strong junior Machine Learning Engineer to support the ML development cycle from data... ...of production ML systems in real customer environments... ...build data pipelines, manage distributed multi-GPU training, and optimize training and inference code, while collaborating across...TrainingWork at office
- ...LLC in New York seeks an experienced ML Research Engineer to join the Enterprise ML Research Lab... ...will build, profile and optimize our training and inference framework and post-train state-of-the... ...-scale GPU clusters and multi-node LLM workloads. #J-18808-Ljbffr United States...Training
$227.2k - $324.5k
...a Staff Software Engineer on the ML Infrastructure team... ...machine learning inference platforms. These platforms... ...ML model serving systems that support Deep Learning, LLM, and Search models... ..., and low latency distributed systems using... ..., model training orchestration, etc...TrainingFull timeTemporary workLocal areaFlexible hours$264.8k - $331k
...state of the art post-training algorithms to reach... .... The Enterprise ML Research Lab works... ...an ML Sys Research Engineer, you’ll work on building... ...to optimize our ML system. Your customer will... ...our training and inference framework. Post-... ...least 1-3 years of LLM training in a production...TrainingFull time$264.8k - $331k
...Machine Learning Systems Research Engineer, Agent Post-training - Enterprise GenAI AI is becoming... ...the world. The Enterprise ML Research Lab works on the... ...our training and inference framework. Post-train state... ...have: At least 1-3 years of LLM training in a production...TrainingFull timeContract workFor contractorsFor subcontractorWork at office$127.4k - $191.2k
...the machine learning systems that decide which... ...world's leading game engine. Recommendation and... ...using causal inference, A/B testing, and offline... ...full pipeline from training data to deployed... ...reinforcement learning, LLM post-training or... ...-scale data and ML systems, whether through...TrainingFull timeInternshipWork at officeWorldwideShift work- ...most valuable training data in the world... ...context distributed across all of them... ...improve how well our system understands and... ...-facing ML role. You will... ...model sweeps, LLM-assisted review... ...rollout Optimize inference cost, latency,... ...Data and Product Engineering on pipeline and...TrainingFull timeShift work
$157.95k - $259.48k
...platforms are on the training ground and in the... ...every sport — a system that connects the... ...highest-leverage engineering position in Phase... ...against the full distribution of real-world... ...Experience with causal inference or counterfactual... ...Familiarity with LLM evaluation...Training$213k - $263k
...demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust... ...Conduct comprehensive experimentation to train and deploy state-of-the-art Multimodal LLMs... ...model training flows in a scalable, distributed and performant manner such as Data parallel...TrainingFull timeTemporary workRemote work$61k - $101k
...Science or a related Engineering field with 10+ years... ...AWQ for accelerating LLM inference on specific GPU architectures... ...transformer-based systems. We require solid... ...of machine learning training, especially... ...robust pipelines for distributed training on GPU-enabled...TrainingFull time$235k - $260k
...The Principal AI/ML Engineer, Semantic Data will... ...across MLS systems.This role combines... ...systems with applied LLM engineering to... ...environmentsOptimize inference workflows for... ...understanding of distributed systems and data... ...offering on-the-job training, feedback, and ongoing...TrainingWork at officeLocal areaRemote work1 day per week$147.6k - $274k
...enable researchers and engineers to move models from... ...scope extends beyond LLM serving. You will work... ...real-time and batch inference workloads, GPU-backed... ...infrastructure, Kubernetes, and distributed systems. Prior inference-... ..., distributed training, or large-scale data...TrainingFull timeLocal areaImmediate startWorldwideRelocation package$175k - $215k
...generative modeling, Bayesian inference, hierarchical learning... ...via data, eval, and systems Implement and... ...developing recipes for ML models We prefer:... ...implement, and extend large distributed pipelines... ...data for ML eval and training The expected base salary...TrainingFull timeRemote work- ...fully remote and globally distributed. In 2026 we opened a... ...working in tech: AI Systems Engineering, AI & Machine Learning... ..., Senior/Staff ML Engineer, Solutions Architect... ...and guardrails, MCP; LLM evals — eval harnesses... ...design, model serving and inference cost. AI Systems...Hourly payFull timeWork at officeRemote workWorldwideAfternoon shift
$130k - $160k
...team building the systems these markets... ...An AI Engineer at Octaura, is... ...by integrating trained machine learning... ...shape Octaura’s ML architecture and... ...ML training and inference. Collaborate... ...for Agentic AI, LLM orchestration,... ...Strong grasp of distributed systems, REST API...TrainingFull timeWork at officeRemote workFlexible hours$207k - $300k
...automation, and alerting systems to ensure model... ...machine learning model training and inference performance, and programming... ...’s degree or PhD in Engineering, Computer Science, or... ...retrieval, distributed computing, large-scale... ...models using AI and ML techniques to predict...Training$192.5k - $357.5k
...Machine Learning Engineer to join our Foundation... ...) and agentic systems, enabling them to... ...useful, to the distributed infrastructure and... ...sources.Scalable ML Systems &... ...scale distributed training and inference systems for foundation... ...Deep expertise in LLM serving, test-time...TrainingFull timeLocal areaWorldwideRelocation package$192k - $260k
...their business. Founded by engineers — and customer obsessed... ...bring enterprise-grade ML and AI personalization... ...have:Evaluate ML and LLM approaches for... ...evaluating ML models and/or LLM systems for real product or... ...relevant certifications and training, and specific work...TrainingWork at officeLocal areaWorldwide$229k - $343k
...services.Snap’s Generative ML Platform team builds... ...device and server-side inference. Our team creates... ...platforms, and agentic systems that empower creators,... ...for a Machine Learning Engineer to join Snap Inc!What you... ...related fieldExperience training large-scale diffusion models...TrainingFull timeLive inWork at officeLocal areaWorldwide- ...world of intelligent systems. Location : New York,... ...deploy production‑grade ML systems with end‑to‑... ...drive innovation in LLM and audio ML... ...preprocessing, model training, deployment, inference, and monitoring in production... ...professional experience in ML engineering. Strong programming...TrainingFull time
$350k
...judgment. We are training frontier models with... ...AI Infrastructure Engineer to keep our post-... ...reinforcement learning (RL) systems fast, reliable,... ...large-scale distributed systems in... ...and mixed training/inference workloadsExperience... ...purpose-built for ML training, not just...TrainingPermanent employmentVisa sponsorshipWork visaRelocation packageShift work- Scale AI is hiring a Machine Learning Systems Research Engineer, Agent Post-training for our Enterprise GenAI team in New York. You will build, optimize, and... ...cybersecurity to healthtech. You will collaborate with ML engineers to run large-scale experiments, iterate on post...Training
- ...business problems.We’re training and deploying frontier... ...who are building AI systems. We believe that our work... ...team of researchers, engineers, designers, and more,... ..., highly available distributed systems with Kubernetes... ...latency and throughput of inference.Strong understanding...TrainingFull timeWork experience placementWork at officeLocal areaRemote workHome office
- ...Data Scientist / ML EngineerWe're looking for a Data Scientist / ML Engineer to help Themis turn data into intelligence that... ...on machine learning models and LLM-powered featuresDesign... ...maintain data and ML pipelines for training, inference, and monitoringDeploy models and...TrainingRemote workFlexible hoursShift work
$135k - $150k
...Machine Learning Engineer. In this role,... ...generation, to training models, to... ...machine learning systems in real customer... ...environments. At Pangram, ML engineers are... ...models Manage distributed infrastructure for multi-GPU LLM training... ...optimizing training and inference code Deploy...TrainingInternshipWork at office$200k - $300k
...on expanding our ML research platform... ...across our entire distributed ML stack. By... ...you will push our systems to their limits,... ...comprehensively test both the training and inference environments.... ...to assist engineering teams with platform... ...Familiarity with LLM tooling, agentic...Training
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Systems Engineer: Distributed LLM Training & Inference. Be the first to apply!
- ai ml engineer New York, NY
- junior machine learning research engineer New York, NY
- senior ml engineer New York, NY
- machine learning ai engineer New York, NY
- computer vision machine learning engineer New York, NY
- data scientist machine learning engineer New York, NY
- machine learning engineer New York, NY
- machine learning software engineer New York, NY
- senior staff systems engineer New York, NY
- application system engineer New York, NY



