Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Distributed ML Systems Engineer — Large-Scale Pretraining

Dormont Manufacturing Co

Dormont Manufacturing Co is looking for a Software Engineer for their Pre-training Systems team in San Francisco. Your primary role will be to design and maintain the distributed infrastructure that trains long-context models at scale, tackling challenges related to memory pressure, communication, and job recovery. The ideal candidate should possess strong software engineering fundamentals, experience with large models, and a proactive attitude towards maintaining critical systems in a high-performance environment. #J-18808-Ljbffr Dormont Manufacturing Co

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Distributed ML Systems Engineer — Large-Scale Pretraining in San Francisco, CA vacancy
  •  ...looking for a Senior Software Engineer to build scalable infrastructure for large‑scale training and fine-tuning of foundation...  ...models. You will design distributed training systems and optimize GPU utilization...  ...over 5 years of experience in ML infrastructure and a strong background... 
    Suggested

    Baseten

    San Francisco, CA
    20 hours ago
  • Genesis AI in San Francisco is seeking a senior ML infrastructure engineer to design and optimize distributed training systems and performance-critical components. You...  ...monitoring tools, and robust performance improvements for large‑scale runs. #J-18808-Ljbffr Genesis AI
    Suggested

    Genesis AI

    San Francisco, CA
    16 hours ago
  •  ...bridge research with production systems. You’ll work across research and engineering to push performance and scale using cutting-edge...  ...software engineers, Python and ML framework expertise, and experience with large-scale distributed training, aiming to push frontier... 
    Suggested
    Remote work

    Cohere

    San Francisco, CA
    6 hours ago
  • $200.8k - $251k

     ...optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA...  ...of $200,800 - $251,000, along with comprehensive benefits. #J-18808-Ljbffr Scale AI
    Suggested
    Full time

    Scale AI

    San Francisco, CA
    3 days ago
  •  ...an Applied Machine Learning Engineer in San Francisco, CA to...  ...researchers, product managers, and systems engineers to optimize, deploy, and scale models like GPT-4 and...  ...engineering with at least 2 years in ML systems, plus deep knowledge of distributed training and modern... 
    Suggested

    AI Breaking Wire

    San Francisco, CA
    20 hours ago
  •  ...under the practical constraints of robotic platforms. About the Role As a Research Engineer, Distributed Data Systems, you will design and scale the infrastructure that powers large-scale multimodal training and evaluation at OpenAI. You’ll manage distributed data... 
    Full time
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    16 hours ago
  •  ...seeking a research scientist to design, implement, and optimize the large-scale training infrastructure powering our reinforcement learning...  ...office. You’ll work with researchers to ensure the RL system is fast, reliable, and capable of days-long runs with minimal... 
    Work at office

    Applied Compute

    San Francisco, CA
    16 hours ago
  • $190k - $265k

     ...San Francisco is looking for a Research Engineer to enhance their AI platform for law,...  ...low-latency search and retrieval systems and managing large datasets of complex documents. The ideal...  ...ideal candidate will have over 5 years of ML engineering experience and proven... 

    Harvey

    San Francisco, CA
    2 days ago
  •  ...leading AI research firm located in San Francisco is seeking a Senior ML Systems Engineer to build and maintain the training framework for large-scale language models. The role involves designing distributed training solutions and improving training throughput across multi-... 
    Flexible hours

    Cohere

    San Francisco, CA
    2 days ago
  •  ...seeking a role to build and operate distributed training systems powering frontier models in SF. You will...  ...role emphasizes collaboration with ML researchers, debugging across GPU stacks...  ...and memory efficiency in large-scale training environments. #J-18808-Ljbffr... 

    Visa Hunt

    San Francisco, CA
    1 day ago
  • Senior ML Systems Engineer, Frameworks & Tooling at Cohere Our mission is to scale intelligence to serve humanity. We’re training and deploying frontier models...  ...This role sits at the intersection of large‑scale training, distributed systems, and HPC infrastructure. You... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Cohere

    San Francisco, CA
    2 days ago
  • Inception is seeking engineers and scientists to design, optimize, and scale the diffusion LLM serving systems powering production inference. Your work will...  ...frameworks (Kubernetes, Ray, SLURM) for distributed inference, evaluation, and large-batch serving, and implement load... 

    Inception

    San Francisco, CA
    4 days ago
  •  ...architectures, and pretraining science. We build...  ...role: As a ML Engineer, you’ll build and...  ...to rapidly test, scale, and iterate on them...  ...’ll work on the systems that support...  ...training and evaluating large models, scaling...  ...high-performance distributed training... 
    Full time
    Internship

    Tilde Research

    San Francisco, CA
    16 hours ago
  •  ...OpenAI is seeking a Research Engineer, Distributed Data Systems, to design and scale infrastructure powering large-scale multimodal training and evaluation. You will manage distributed data pipelines and partner with researchers to translate requirements into robust systems... 
    Relocation package

    Neura Market

    San Francisco, CA
    21 hours ago
  •  ...Whatnot is hiring Senior Software Engineers across multiple pillars to...  ...product and platform systems powering one of the fastest-...  ...sellers. Our teams rebuild large distributed systems for rapid growth, increasing...  ..., and real-time commerce at scale. #J-18808-Ljbffr... 
    Remote work

    Whatnot

    San Francisco, CA
    20 hours ago
  •  ...the most advanced Large Language Models. Overview...  ...talented MLOps Engineers with deep, hands-...  ...for frontier AI systems. This is a W-2...  ...infrastructure, and ML framework-level topics...  ...pipeline design, distributed systems reasoning,...  ...with PyTorch at scale. Experience writing... 
    Full time
    Weekday work

    Obsidian

    San Francisco, CA
    1 day ago
  • $189.6k - $237k

    Scale’s ML platform (RLXF) team builds our internal distributed framework for large language model training and inference. The platform has...  ...:Strong excitement about system optimizationExperience with...  ...ML systemsStrong software engineering skills, proficient in frameworks... 
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  • $110 per hour

     .... Position: MLOps Engineer (JAX, PyTorch, Pallas/Triton...  ...infrastructure, and ML framework-level topics...  ...to MLOps and ML systems problems . Evaluate...  ...training pipeline design, distributed systems reasoning, and...  ...JAX and/or PyTorch at scale. ~ Experience writing... 
    Remote job
    Contract work
    Summer work
    Weekday work

    Mercor

    San Francisco, CA
    3 days ago
  • $350k

     ...Francisco, is searching for an engineer to work at the intersection of research and systems on their pretraining stack. The role involves...  ...model architectures and scaling distributed training jobs across thousands...  ...'re driven to explore what large models can achieve, we want... 

    Mirendil

    San Francisco, CA
    4 days ago
  • Design Arena hosts cutting-edge ML evaluation platforms and data engines. This role focuses on building systems that learn from human preference signals to evaluate...  ...will collaborate with researchers and scale data pipelines for large-scale experiments. Ideal candidates... 

    Design Arena

    San Francisco, CA
    4 days ago
  •  ...Scale AI is hiring a Machine Learning Systems Research Engineer, Agent Post-training for our Enterprise GenAI team in New York. You will build, optimize,...  ...to healthtech. You will collaborate with ML engineers to run large-scale experiments, iterate on post-training recipes... 

    Scale

    San Francisco, CA
    20 hours ago
  •  ...Anthropic is seeking an engineer to own the infrastructure behind safeguards research. You will build the tooling researchers rely...  .... Successful candidates will have a track record of solving large-scale systems and data problems and will grow deep machine learning... 

    Jobzhr

    San Francisco, CA
    20 hours ago
  •  ...-on support from AMD engineers the team is scaling rapidly to build the...  ...About the role As an ML Engineer at...  ...end multimodal GenAI systems. In this role, you will...  ...AI systems including Large Language Models, Transformers...  ...with one or more distributed ML training frameworks... 
    Full time
    Flexible hours

    Sciforium

    San Francisco, CA
    8 hours ago
  •  ...support from AMD engineers the team is scaling rapidly to build...  ...-model stack: pretraining and scaling ,...  ...Scaling Train large byte-native foundation...  ...and maintain distributed training...  ...robust, performant systems) Experience with...  ...-end production ML systems with monitoring... 
    Full time
    Flexible hours

    Sciforium

    San Francisco, CA
    16 hours ago
  •  ...of voice AI operating systems for clinicians, transforming...  ...platform scaling to thousands of clinicians...  ...We are hiring two ML Engineers / Researchers to help...  ...text systems using our large proprietary dataset of...  ...large-scale datasets and distributed training environments... 
    Full time

    Knowtex

    San Francisco, CA
    16 hours ago
  •  ...Francisco seeks a Machine Learning Engineer to work with the full ML stack, implementing advanced model architectures...  ...extensive data pipelines for large datasets. The ideal candidate will...  ...models, and familiarity with distributed training processes. A bonus is knowledge... 

    Kindredventures

    San Francisco, CA
    2 days ago
  • $170k - $280k

     ...give leaders clarity and engineers time. We help...  ...looking for an Applied ML Engineer to help build...  ...the machine learning systems that power Macroscope'...  ...role where you'll own large parts of the model development...  ...Experience with large-scale distributed training, preference... 
    Odd job
    Full time

    Macroscope Inc

    San Francisco, CA
    16 hours ago
  •  ...interpretability, new architectures, and pretraining science. As an ML Engineer, you will build and operate the...  ...research in training and evaluating large models. You will optimize inference and training throughput, scale distributed pipelines, and collaborate with researchers... 

    Tilde Research

    San Francisco, CA
    1 day ago
  •  ...Machine Learning Infrastructure Engineer to help architect the...  ...technology enterprises, this team is scaling rapidly to solve complex...  ...ownership over scaling distributed systems across massive hardware clusters...  ...reliability while running large-scale workloads across extensive... 
    Full time
    Work at office
    Flexible hours

    Objective Paradigm

    San Francisco, CA
    3 days ago
  • $166k - $210.25k

     ...SummaryAs a Senior Applied ML Engineer on the Applied AI team...  ...Growth: Drive the scaling and efficiency of...  ...optimization techniques.Build Systems: Design end-to-end ML4...  ...computer systems and distributed environments.Minimum...  ...record of optimizing large-scale distributed systems... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Distributed ML Systems Engineer — Large-Scale Pretraining. Be the first to apply!