Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Distributed ML Systems Engineer — Large-Scale Pretraining

Dormont Manufacturing Co

Dormont Manufacturing Co is looking for a Software Engineer for their Pre-training Systems team in San Francisco. Your primary role will be to design and maintain the distributed infrastructure that trains long-context models at scale, tackling challenges related to memory pressure, communication, and job recovery. The ideal candidate should possess strong software engineering fundamentals, experience with large models, and a proactive attitude towards maintaining critical systems in a high-performance environment. #J-18808-Ljbffr Dormont Manufacturing Co

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Distributed ML Systems Engineer — Large-Scale Pretraining in San Francisco, CA vacancy
  • Responsibilities Design, deploy, and maintain large distributed ML training and inference clusters...  ...end-to-end pipelines to manage petabyte-scale datasets and model training throughout...  ...working on distributed task management systems and scalable model serving & deployment... 
    Suggested

    Kindredventures

    San Francisco, CA
    4 days ago
  •  ...looking for a Senior Software Engineer to build scalable infrastructure for large‑scale training and fine-tuning of foundation...  ...models. You will design distributed training systems and optimize GPU utilization...  ...over 5 years of experience in ML infrastructure and a strong background... 
    Suggested

    Baseten

    San Francisco, CA
    2 days ago
  • Genesis AI in San Francisco is seeking a senior ML infrastructure engineer to design and optimize distributed training systems and performance-critical components. You...  ...monitoring tools, and robust performance improvements for large‑scale runs. #J-18808-Ljbffr Genesis AI
    Suggested

    Genesis AI

    San Francisco, CA
    3 days ago
  • $245k - $385k

     ...the Team The Platform ML team builds the ML...  ...edge models. We work on distributed model execution as well...  ...Role As a Distributed Systems/ML engineer, you will work on...  ...framework is used for large runs with massive numbers...  ...you: Have run small scale ML experiments Love figuring... 
    Suggested
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...bridge research with production systems. You’ll work across research and engineering to push performance and scale using cutting-edge...  ...software engineers, Python and ML framework expertise, and experience with large-scale distributed training, aiming to push frontier... 
    Suggested
    Remote work

    Cohere

    San Francisco, CA
    3 days ago
  • $237k - $337k

     ...qualifications: Master’s degree or PhD in Engineering, Computer Science, or a related...  ...to handle information at massive scale, and extend well beyond web search....  ..., including information retrieval, distributed computing, large-scale system design, networking and data storage... 
    Full time

    Google

    San Francisco, CA
    18 hours ago
  •  ...company in San Francisco is looking for a Founding ML Research Engineer to develop the infrastructure for training large speech models. This entry-level position...  ..., building scalable data pipelines, and owning distributed training processes. Ideal candidates will have... 
    Full time

    Kalpa Labs (YC F25)

    San Francisco, CA
    18 hours ago
  • $200.8k - $251k

     ...optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA...  ...of $200,800 - $251,000, along with comprehensive benefits. J-18808-Ljbffr Scale AI
    Full time

    Scale AI

    San Francisco, CA
    1 day ago
  • $190k - $265k

     ...San Francisco is looking for a Research Engineer to enhance their AI platform for law,...  ...low-latency search and retrieval systems and managing large datasets of complex documents. The ideal...  ...ideal candidate will have over 5 years of ML engineering experience and proven... 

    Harvey

    San Francisco, CA
    18 hours ago
  • $245k - $385k

    Dormont Manufacturing Co is seeking a Distributed Systems/ML engineer in San Francisco, CA. You'll improve training throughput for our internal framework and enable researchers to innovate. Strong Python skills are essential. The position offers a hybrid work model, robust... 

    Dormont Manufacturing Co

    San Francisco, CA
    2 days ago
  •  ...seeking a research scientist to design, implement, and optimize the large-scale training infrastructure powering our reinforcement learning...  ...office. You’ll work with researchers to ensure the RL system is fast, reliable, and capable of days-long runs with minimal... 
    Work at office

    Applied Compute

    San Francisco, CA
    3 days ago
  • $160k - $180k

    Discord is seeking a backend systems engineer to join the Realtime Infrastructure team in San Francisco...  ...role, you will build and maintain high-scale services that are essential for our...  ...of experience and strong knowledge of distributed systems. The position offers a... 

    Discord

    San Francisco, CA
    18 hours ago
  • Senior ML Systems Engineer, Frameworks & Tooling at Cohere Our mission is to scale intelligence to serve humanity. We’re training and deploying frontier models...  ...This role sits at the intersection of large‑scale training, distributed systems, and HPC infrastructure. You... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Cohere

    San Francisco, CA
    18 hours ago
  •  ...leading AI research firm located in San Francisco is seeking a Senior ML Systems Engineer to build and maintain the training framework for large-scale language models. The role involves designing distributed training solutions and improving training throughput across multi-... 
    Flexible hours

    Cohere

    San Francisco, CA
    18 hours ago
  • Inception is seeking engineers and scientists to design, optimize, and scale the diffusion LLM serving systems powering production inference. Your work will...  ...frameworks (Kubernetes, Ray, SLURM) for distributed inference, evaluation, and large-batch serving, and implement load... 

    Inception

    San Francisco, CA
    2 days ago
  • $200k - $275k

     ...more secure. The AI Engineering Team is chartered...  ...special focus on Large Language Models (LLMs) and agentic systems. Our mission is to...  ..., safety, and scale. We manage petabyte...  ...a Senior or Staff ML Systems Engineer -...  ...TRM operates as a distributed-first company with... 
    Remote work
    Worldwide

    Trm-Labs

    San Francisco, CA
    8 hours ago
  • About TrueFoundry Every production AI system, whether it's powering customer support, writing code, analyzing financial...  ...TrueFoundry, and we're building it. We're looking for a Staff ML Platform Engineer - Large Scale Training (LLMOps/MLOps) to join the team. The Problem We'... 
    Flexible hours

    TrueFoundry

    San Francisco, CA
    4 days ago
  • Dormont Manufacturing Co is seeking a Staff ML Platform Engineer to join their team. This role focuses on building infrastructure for large-scale training of ML models, requiring deep knowledge of PyTorch and multi-node systems. Join a dynamic environment poised for... 

    Dormont Manufacturing Co

    San Francisco, CA
    18 hours ago
  • $149k - $186k

     ...are looking for brilliant software engineers who have a passion for consumer...  ...years experience in the realm of large-scale data processing, distributed systems, and/or machine learning infrastructure...  ...with operational tooling for ML services (monitoring, alerting, etc... 
    Full time
    Work at office
    Local area

    Turo Inc.

    San Francisco, CA
    18 hours ago
  •  ...seeking a qualified Data Infrastructure Engineer based in New York. This role involves designing, maintaining, and optimizing data systems with a focus on correctness and performance...  ...with Ray Data, and hands-on ownership of large data pipelines. Responsibilities include... 
    Flexible hours

    cursor

    San Francisco, CA
    3 days ago
  • $350k

     ...looking for a Research Engineer in San Francisco to work with the ML Performance and Scaling team. The role...  ...critical management of pretraining pipelines and collaboration...  ...debugging complex ML systems. Successful...  ...will have a passion for large-scale ML models and be... 

    Anthropic

    San Francisco, CA
    18 hours ago
  • EngineersOfAI is seeking a talented backend engineer to join our Realtime Infrastructure...  ...role, you will build and operate large-scale, reliable systems while collaborating with product...  ...strong capability in solving complex distributed system problems. Join us in shaping... 

    EngineersOfAI

    San Francisco, CA
    18 hours ago
  •  ...the most advanced Large Language Models. Overview...  ...talented MLOps Engineers with deep, hands-...  ...for frontier AI systems. This is a W-2...  ...infrastructure, and ML framework-level topics...  ...pipeline design, distributed systems reasoning,...  ...with PyTorch at scale. Experience writing... 
    Full time
    Weekday work

    Obsidian

    San Francisco, CA
    4 days ago
  • $350k

     ...Francisco, is searching for an engineer to work at the intersection of research and systems on their pretraining stack. The role involves...  ...model architectures and scaling distributed training jobs across thousands...  ...'re driven to explore what large models can achieve, we want... 

    Mirendil

    San Francisco, CA
    2 days ago
  •  ...research and inference. Distributed training, 1000+ K8...  ..., petabyte scale data pipelines, etc...  ...job orchestration systems, and streaming pipelines...  ...and scaling for large-scale GPU job...  ...systems for large-scale pretraining Collaborate with...  ...Applied ML pipelines Find clean... 

    krea.ai

    San Francisco, CA
    18 hours ago
  •  ...technology firm in San Francisco is seeking systems-oriented candidates to enhance their...  ...in Python and experience managing distributed systems, particularly within GPU environments...  ...on cutting-edge projects involving large-scale data processing and custom infrastructure... 

    krea.ai

    San Francisco, CA
    18 hours ago
  • Systems Engineers at Inngest build the core of our product: the durable execution layer, queueing system, state stores, and the distributed systems that connect them together. It's extremely fun, rewarding...  ...able to build something that scales and functions well, a Systems... 
    Live in

    Inngest

    San Francisco, CA
    8 hours ago
  • Harrison Clarke is looking for engineers to join their rapidly growing...  ...machine learning, building systems that directly impact model quality...  ...experience. You’ll focus on large-scale user interaction data,...  ...collaborating with product and ML teams. This position offers the... 

    Harrison Clarke

    San Francisco, CA
    18 hours ago
  • $218.4k - $273k

    Scale's Physical AI business unit is dedicated to solving the data...  ...Physical AI and developing ML pipelines for processing, training...  ...AI. The Role As an ML Systems Engineer on the Physical AI team, you...  ...years of experience building large‑scale, high‑performance backend... 
    Full time

    Scale AI

    San Francisco, CA
    18 hours ago
  • About the Role Scale's Physical AI business unit is dedicated to...  ...in Physical AI and developing ML pipelines for processing,...  ...evaluation for Physical AI. As an ML Systems Engineer on the Physical AI team, you...  ...years of experience building large‑scale, high‑performance... 

    Gravity Engineering Services Pvt Ltd.

    San Francisco, CA
    18 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Distributed ML Systems Engineer — Large-Scale Pretraining. Be the first to apply!