Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Infra Engineer — Scalable Training Systems

Monograph

A leading tech company in San Francisco seeks a Machine Learning Engineer to build and maintain infrastructure for large-scale model training. In this hands-on role, you will design systems, work closely with researchers, and optimize training processes. Candidates should have strong software engineering skills and experience with JAX or PyTorch. Join a dynamic team at the forefront of machine learning and contribute to core training code and systems. #J-18808-Ljbffr Monograph

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the ML Infra Engineer — Scalable Training Systems in San Francisco, CA vacancy
  •  ...is looking for a Senior Software Engineer to build scalable infrastructure for large‑scale training and fine-tuning of foundation...  ...will design distributed training systems and optimize GPU utilization while...  ...over 5 years of experience in ML infrastructure and a strong background... 
    Training

    Baseten

    San Francisco, CA
    3 days ago
  • Physical Intelligence seeks a systems-focused ML infrastructure engineer to own scheduling, placement, and cluster...  ...management for large-scale model training. You will design multi-tenant...  ...to translate workloads into scalable infra, optimize utilization, and build... 
    Training

    Physical Intelligence

    San Francisco, CA
    2 days ago
  •  ...scale and optimize our training systems and core model code. You...  ...with researchers and model engineers to translate ideas into...  ...at the intersection of ML, software engineering, and scalable infrastructure. The Team...  ...research needs into infra capabilities and guide best... 
    Training
    Full time

    Monograph

    San Francisco, CA
    13 hours ago
  •  ...the physical world. Training our models...  ...end: the scheduling systems, the placement logic...  ...seamless. The Team The ML Infrastructure...  ...work closely with ML Infra (training systems)...  ...Strong software engineering fundamentals...  ...engineering, and scalable infrastructure. The... 
    Training
    Flexible hours

    Physical Intelligence

    San Francisco, CA
    2 days ago
  • A leading AI research firm located in San Francisco is seeking a Senior ML Systems Engineer to build and maintain the training framework for large-scale language models. The role involves designing distributed training solutions and improving training throughput across... 
    Training
    Flexible hours

    Cohere

    San Francisco, CA
    1 day ago
  • Senior ML Systems Engineer, Frameworks & Tooling at Cohere Our mission is to...  ...to serve humanity. We’re training and deploying frontier models...  ...enable fast, reliable, and scalable model training and build the...  .... Collaborate closely with infra teams to ensure Slurm setups... 
    Training
    Full time
    Work at office
    Remote work
    Flexible hours

    Cohere

    San Francisco, CA
    1 day ago
  • Ginas Tech Jobs is seeking a Principal Machine Learning Engineer to set the technical standard for ML systems across training, inference, evaluation and deployment. This 100% remote role balances hands‑on engineering with architectural leadership, spanning data, models,... 
    Training
    Remote job

    Ginas Tech Jobs

    San Francisco, CA
    2 days ago
  • AI Chopping Block, Inc. is seeking a Machine Learning Engineer to design and build scalable machine learning systems. Responsibilities involve developing end-to-end ML pipelines, optimizing AI models for mobile environments, and integrating AI-driven solutions into applications... 

    AI Chopping Block, Inc.

    San Francisco, CA
    1 day ago
  • $227.2k - $284k

     ...research in Physical AI and developing ML pipelines for processing, training, and fine-tuning on data collected...  .... The Role As an ML Systems Engineer on the Physical AI team, you will design and build platforms for scalable, reliable, and efficient serving of... 
    Training
    Full time

    Scale AI

    San Francisco, CA
    1 day ago
  • Genesis AI in San Francisco is seeking a senior ML infrastructure engineer to design and optimize distributed training systems and performance-critical components. You will...  ...‑node GPU clusters. Join a team focused on scalable AI foundations, monitoring tools, and robust... 
    Training

    GenesisAI

    San Francisco, CA
    4 days ago
  • Scale AI is hiring a Machine Learning Systems Research Engineer, Agent Post-training for our Enterprise GenAI team in...  ...state-of-the-art models, enabling scalable agent training across enterprise use...  ...healthtech. You will collaborate with ML engineers to run large-scale... 
    Training

    Scale

    San Francisco, CA
    2 days ago
  • $250k

     ...platforms for large-scale AI training and inference workloads....  ...enables AI teams to access scalable compute environments without...  ...limitations. As a Senior ML Infrastructure Engineer, the successful candidate...  ...scale training and inference systems. The role focuses on... 
    Training
    Full time
    San Francisco, CA
    a month ago
  •  ...Technical Staff to design and operate distributed systems for serving models in production and driving large-scale post-training workflows. You will work where model...  ...the infrastructure enabling fast inference and scalable RL iteration, balancing KV-cache strategies,... 
    Training

    Magic AI Corp.

    San Francisco, CA
    3 days ago
  •  ...About the Role As a Research Engineer, Distributed Data Systems, you will design and scale the infrastructure...  ...that powers large-scale multimodal training and evaluation at OpenAI. You’ll...  ...infrastructure while ensuring scalability, reliability, and security. Ensure... 
    Training
    Full time
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    13 hours ago
  •  ...Francisco is seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate will have hands-on experience with modern... 
    Training

    Reflection AI

    San Francisco, CA
    3 days ago
  •  ...don't believe culture can be engineered - but when it falls into place...  ...We're looking for an ML infrastructure engineer to help...  ..., and scale the foundational systems we need to realize our ambitious...  ...supports every stage of the ML training flywheel and be an important... 
    Training
    Local area

    Humble Robotics

    San Francisco, CA
    3 days ago
  • $189.6k - $237k

    Scale’s ML platform (RLXF) team builds our internal distributed...  ...for large language model training and inference. The platform has...  ...have:Strong excitement about system optimizationExperience with multi...  ...ML systemsStrong software engineering skills, proficient in... 
    Training
    Full time

    Scale AI

    San Francisco, CA
    1 day ago
  • Shipt is seeking a Staff Machine Learning Engineer on the Personalization Platform team to drive AI initiatives...  ...engagement. You will architect and implement scalable ML infrastructure, design data pipelines for training and serving models, and lead the team to deliver... 
    Training

    Shipt

    San Francisco, CA
    4 days ago
  •  ...decks — partner with research and infra to prototype, train, and deploy state-of-the-art voice...  ...-level PyTorch. Proven software engineer who loves ML; comfortable writing production code...  ...—especially user-facing, online ML systems—despite shifting requirements and surprise... 
    Training
    Full time
    Contract work
    Flexible hours
    Shift work

    Sesame, L.l.c.

    San Francisco, CA
    13 hours ago
  • $200.8k - $251k

     ...member to build and optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-time position offers a competitive salary... 
    Training
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  • $129k - $198.4k

     ...DescriptionRole: As an AI/ML Engineer on the Metrics...  ...monitoring, data mining and training, and simulation...  ...high-performing, and scalable driverless technology....  ...Evaluation, Embodied AI, and System and Test Engineering...  ...other frameworks and data infra teams to build and... 
    Training
    Full time
    Local area
    Work from home

    General Motors

    San Francisco, CA
    13 hours ago
  • $197.3k - $313.7k

     ...for a Staff Machine Learning Engineer with deep expertise in model training and finetuning to join our ML team. You'll design, train, and...  ....Build and maintain scalable finetuning training pipelines...  ...Expertise with recommendation systems or search.Familiarity with model... 
    Training
    Full time

    Salesforce

    San Francisco, CA
    13 hours ago
  • $171k

     ...the Core Security Engineering organization, is building...  ...-driven security systems. We're evolving...  ...toward real-time, ML-driven access...  ...ensuring reliability and scalability. # Collaborate...  ...engineering, training, and evaluation....  ...large-scale data/infra systems (Kafka, Hive... 
    Training
    Full time
    Work at office
    Remote work

    Uber Technologies Inc

    San Francisco, CA
    5 days ago
  •  ...AI/ML Engineer (RL & Physical Systems) FLUIX is building the AI Operating System for data centers. We deploy autonomous AI that optimizes, predicts...  ...digital twin and simulation environments to accelerate training, testing, and Sim2Real deployment. Conduct lab-... 
    Training
    Weekend work

    Fluix AI

    San Francisco, CA
    5 days ago
  • $250k - $400k

     ...Define how large-scale AI systems for scientific discovery are actually built, trained, and run in production....  .... It's building the engine that research runs on....  ...Experience building and scaling ML systems in production...  ...available: ML Engineer, ML Infra, Research Engineers &... 
    Training
    Remote work

    techire ai

    San Francisco, CA
    3 days ago
  • $190k - $205k

     ...automated hazard detection systems Work with multimodal...  ...visual signals Production Engineering Write clean, scalable, well-tested Python code that...  .... Build end-to-end ML pipelines including data processing...  ..., feature extraction, training, evaluation, and... 
    Training
    Full time
    Live in

    Gridware

    San Francisco, CA
    13 hours ago
  • $213k - $263k

     ...as the foundation for training and validating the AV...  ...stack. We are an advanced ML and engineering team that leverages...  ...Design and implement a scalable AI agent framework...  ...continuously improving the system's captioning and...  ...Collaborate closely with the ML Infra, Perception, Behavior,... 
    Training
    Full time
    Remote work

    Waymo

    San Francisco, CA
    3 days ago
  •  ...build the shared ML and AI infrastructure...  ...the foundational systems, models, and data...  ...network data into scalable, general-purpose representations...  ...Machine Learning Engineer, you will lead the...  ..., distributed training, serving...  ...capabilities (serving infra, feature stores)... 
    Training
    Full time
    Work experience placement
    Local area
    Immediate start

    Plaid Inc.

    San Francisco, CA
    13 hours ago
  • $170k - $200k

     ...valuable for robotics and physical AI training. We’re working with leading...  .... The Instawork Robotics ML Engineer will help build and scale the technology...  ...with expertise in distributed systems, cloud computing on AWS infrastructure, and scalable data processing for large-scale... 
    Training
    Hourly pay
    Internship
    Local area
    Shift work

    Instawork

    San Francisco, CA
    3 days ago
  • $227.33k - $312.58k

    We’re looking for a Staff ML Data Engineer to join Procore’s AI & Frontier...  ...and building the data systems that power frontier‑scale machine...  ...ambiguous research needs into scalable, production‑ready data...  ...that support machine learning training, evaluation, or inference workflows... 
    Training
    Full time
    Work at office
    Local area
    Immediate start
    3 days per week

    Procore Technologies

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Infra Engineer — Scalable Training Systems. Be the first to apply!