Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Machine Learning Engineer - Distributed ML Systems

Pluralis Research

Overview

Pluralis Research carries out foundational research on Protocol Learning : multi-participant training of foundation models where no single participant has, or can ever obtain, a full copy of the model. The purpose of Protocol Learning is to facilitate the creation of community-trained and community-owned frontier models with self-sustaining economics.

We’re looking for Senior/Staff engineers with 5+ years of experience in distributed systems and ML large‑scale training. You’ll be implementing a novel substrate for training distributed ML models that work under consumer grade internet connection.

Responsibilities

Distributed Training Architecture & Optimization

  • Design and implement large‑scale distributed training systems optimized for heterogeneous hardware operating under low‑bandwidth, high‑latency conditions.

  • Develop and optimize model‑parallel training strategies (data, tensor, pipeline parallelism) with custom sharding techniques that minimize communication overhead.

  • Optimize GPU utilization, memory efficiency, and compute performance across distributed nodes.

  • Implement robust checkpointing, state synchronization, and recovery mechanisms for long‑running, fault‑prone training jobs.

  • Build monitoring and metrics systems to track training progress, model quality, and system bottlenecks.

Decentralized Networking & Resilience

  • Architect resilient training systems where nodes can fail, networks can partition, and participants can dynamically join or leave.

  • Design and optimize peer‑to‑peer topologies for decentralized coordination across non‑co‑located nodes.

  • Implement NAT traversal, peer discovery, dynamic routing, and connection lifecycle management.

  • Profile and optimize communication patterns to reduce latency and bandwidth overhead in multi‑participant environments.

What You’ll Bring

  • Strong experience building and operating distributed systems in production.

  • Hands‑on expertise with distributed training frameworks (FSDP, DeepSpeed, Megatron, or similar).

  • Deep understanding of model parallelism (data, tensor, pipeline parallelism).

  • Expert‑level Python with production experience (concurrency, error handling, retry logic, clean architecture).

  • Strong networking fundamentals: P2P systems, gRPC, routing, NAT traversal, distributed coordination.

  • Experience optimizing GPU workloads, memory management, and large‑scale compute efficiency.

What We Offer

  • Equity‑heavy compensation with meaningful ownership in a mission‑driven company

  • Competitive base salary for senior engineering roles in Australia

  • Visa sponsorship available for exceptional candidates

  • Remote‑first with optional access to our Melbourne hub

  • World‑class team — team mates were previously at Google, Amazon, Microsoft, and leading startups

Backed by Union Square Ventures and other tier‑1 investors, we’re a world‑class, deeply technical team of ML researchers and engineers. Pluralis is unapologetically ideological. We view the world as a better place if we are able to implement what we are attempting, and Protocol Learning as the only plausible approach to preventing a handful of massive corporations monopolising model development, access and release, and achieving massive economic capture. If this resonates, please apply.

#J-18808-Ljbffr
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Machine Learning Engineer - Distributed ML Systems in San Francisco, CA vacancy
  •  ...robotic platforms. About the Role As a Research Engineer, Distributed Data Systems, you will design and scale the infrastructure that powers...  ..., distributed storage, streaming infrastructure, machine learning infrastructure while ensuring scalability, reliability... 
    Suggested
    Full time
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    11 hours ago
  •  ...Francisco is looking for a Senior Software Engineer to build scalable infrastructure for...  ...of foundation models. You will design distributed training systems and optimize GPU utilization while...  ...candidates have over 5 years of experience in ML infrastructure and a strong background... 
    Suggested

    Baseten

    San Francisco, CA
    3 days ago
  • Genesis AI in San Francisco is seeking a senior ML infrastructure engineer to design and optimize distributed training systems and performance-critical components. You will profile bottlenecks, implement low‑level code (CUDA, Triton) and ensure efficient hardware utilization... 
    Suggested

    GenesisAI

    San Francisco, CA
    4 days ago
  • $200.8k - $251k

     ...company in San Francisco seeks a team member to build and optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-... 
    Suggested
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  • $229.9k - $262.4k

     ...Senior Lead Software Engineer, Distributed Systems (Golang + Python on Kubernetes) Do you love building...  ...within Capital One. The Machine Learning Experience Team (MLX Tech) is committed...  ...pioneering and responsibly implementing AI/ML across Capital One . We achieve... 
    Suggested
    Full time
    Part time
    Internship
    Local area

    Capital One

    San Francisco, CA
    4 days ago
  • $90k

    Distributed Systems Software Engineer, Python / Go Join to apply for the Distributed Systems Software Engineer...  ...capabilities to new clouds and developing AI/ML pipelines for automatic analysis of...  ...remotely since 2004! Personal learning and development budget of USD 2,000... 
    Full time
    Freelance
    Internship
    Local area
    Remote work
    Worldwide

    Canonical

    San Francisco, CA
    1 day ago
  •  ...have a legal entity. Responsibilities As a senior Machine Learning Systems Engineer on the Search Platform team, you will own and drive the design...  ...-based semantic search. Own end-to-end delivery of ML components from experimentation through production rollout... 
    Work at office
    Local area

    Atlassian

    San Francisco, CA
    2 days ago
  • An innovative company is seeking a Distributed Systems/ML Engineer to enhance the training throughput of its internal framework. This role involves collaborating with researchers to develop efficient video models and applying cutting-edge techniques to optimize training... 

    Jobleads-US

    San Francisco, CA
    11 hours ago
  •  ...where we have a legal entity.ResponsibilitiesAs a Principal ML System Engineer in the Rovo & AI Engineering org, you will contribute to...  ...breakthrough quality and reliabilityDeep understanding of Machine Learning projects lifecycleExperience leading and supporting senior... 
    Work at office
    Local area

    Atlassian

    San Francisco, CA
    3 days ago
  •  ...have a legal entity. Responsibilities As a Principal Machine Learning Systems Engineer on the Search Platform team, you set the technical...  ...ambiguity and competing priorities across the Search Platform, ML Platform, AI Gateway, and Rovo product teams to keep critical... 
    Work at office
    Local area
    Shift work

    Atlassian

    San Francisco, CA
    3 days ago
  • $189.6k - $237k

    Scale’s ML platform (RLXF) team builds our internal distributed framework for large language model training...  ...excitement about system optimizationExperience...  ...systemsStrong software engineering skills, proficient in frameworks...  ...retirement benefits, a learning and development stipend... 
    Full time

    Scale AI

    San Francisco, CA
    1 day ago
  • $160k - $250k

     ...the future of AI! Senior Machine Learning Engineer In order to execute...  ...Everything involved in applying a ML model to a production use...  ...product and core backend systems by suggesting and executing...  ...code and training across distributed systems You have an ability... 
    Full time

    Hive

    San Francisco, CA
    11 hours ago
  •  ...problems, care about elegant systems design, and want to build...  ...Role We’re looking for a Machine Learning Engineer who loves getting close to...  ...of running large-scale ML systems, and thrives in fast...  ...preferred). ~ Familiarity with distributed training frameworks (e.g.,... 
    Full time
    Work at office

    Relace

    San Francisco, CA
    11 hours ago
  •  ...reinvent the way people learn, starting with...  ...Ventures, and more, with a distributed team across San Francisco...  ...We’re hiring an ML Engineer, Assessments to help...  ...best-in-class assessment systems across multiple products...  .../Learning Design) , Machine Learning, Product, and... 
    Full time
    Live in
    Immediate start

    Speak

    San Francisco, CA
    11 hours ago
  • $244k - $320k

     ...AI-powered personalization engine delivers bespoke experiences...  ...customers.   With a distributed global workforce and employee...  ...! About the Role Our Machine Learning Engineering team powers personalized...  ...operating production-grade ML systems that drive real-time... 
    Full time

    Attentive

    San Francisco, CA
    11 hours ago
  •  ...The role We’re looking for a Machine Learning Engineer to build and ship consumer-facing AI systems that power personalization,...  ...contribute Build and deploy ML models that improve sleep experiences...  ...with data tooling (SQL, distributed compute such as Spark/Ray, and... 
    Full time
    Immediate start
    Worldwide
    Night shift

    Eight Sleep

    San Francisco, CA
    11 hours ago
  • $150k - $190k

     ...simulation software stack for engineering and manufacturing across...  ...We're Looking For As a Machine Learning Engineer in Delivery, you...  ...and used. You’ve shipped ML systems end-to-end and at scale: you...  ...g., TensorFlow, MLFlow) ~ Distributed computing frameworks (e.g.,... 
    Remote job
    Full time
    Flexible hours

    Physicsx

    San Francisco, CA
    11 hours ago
  •  ...hands-on support from AMD engineers the team is scaling...  ...and Reinforcement Learning , sandbox environments...  ...Build and maintain distributed training infrastructure...  ...writing robust, performant systems) Experience with...  ...end-to-end production ML systems with monitoring... 
    Full time
    Flexible hours

    Sciforium

    San Francisco, CA
    11 hours ago
  •  ...inspires future generations. Senior Machine Learning Engineer Primary: Bay Area (San Francisco...  ...constraints, so you get to stand up ML systems the right way from day one. You're...  ...Thinker: You've designed scalable distributed systems and data-intensive applications... 
    Full time
    Work at office
    Relocation package
    Flexible hours
    2 days per week

    Mrbeast

    San Francisco, CA
    11 hours ago
  • $200k - $400k

     ...understanding and designing AI systems that people can trust....  ...researchers and engineers from organizations...  ...We’re looking for Machine Learning Engineers to help build...  ...years of experience in ML infra, research engineering...  ..., PyTorch or Jax, and distributed systems. ~... 
    Full time

    Goodfire

    San Francisco, CA
    11 hours ago
  • $225k - $300k

     ...Machine Learning Engineer About Latent Health Healthcare today is only...  ...history is scattered across systems that don’t communicate. Physicians...  ...extraordinary depth. ML at Latent Health The...  ...Experience building distributed systems or large-scale data... 
    Full time
    Work at office
    Immediate start

    Latent

    San Francisco, CA
    11 hours ago
  • $180k - $270k

     ...privacy protection. To learn more about Plaud,...  ...-low-latency inference engines for large language models...  ...intersection between the core ML training team and the...  ...genuinely enjoy the systems-engineering challenge...  ...accuracy. Large-Scale Distributed Systems: Deploying... 
    Full time
    Work at office
    Worldwide

    Plaud

    San Francisco, CA
    11 hours ago
  •  ...We’re a team of engineers, neuroscientists,...  ...intense, fast-paced learning. You will:...  ...implement the best machine learning approaches...  ...design end-to-end systems Explore new model...  ...years of applied ML research or development...  ...Experience with distributed or large-scale... 
    Full time

    Orbit Neuro Co.

    San Francisco, CA
    11 hours ago
  •  ...operators across hospital and health systems, pharmacies and payors to...  ...Role We're looking for a Machine Learning Engineer to design, build, and deploy production-grade ML systems that power the next...  ...data pipelines using SQL and distributed data processing tools You'... 
    Full time
    Work at office
    Remote work
    Flexible hours
    2 days per week

    Plenful

    San Francisco, CA
    11 hours ago
  • $55 per hour

     ...rely heavily on legacy ERP systems that don't talk to each...  ...of the Acquired podcast to learn how credit card networks did...  ...how we do modeling, data engineering, and production ML at Wholesail for years to...  ...APIs, and reasoning about distributed systems that consume your... 
    Full time
    Temporary work
    Work at office
    Local area

    Wholesail

    San Francisco, CA
    11 hours ago
  • Senior ML Systems Engineer, Frameworks & Tooling at Cohere Our mission is to scale intelligence to serve humanity. We’re training and deploying...  ...role sits at the intersection of large‑scale training, distributed systems, and HPC infrastructure. You will design and... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Cohere

    San Francisco, CA
    1 day ago
  • $148.7k - $199.4k

    Job Posting Title:Senior Machine Learning Engineer - ESPNReq ID:10150610Job Description...  ...worldwide advertising and distribution to maximize flexibility and...  ...distributed data and ML infrastructure that supports...  ..., feature computation systems, and ML‑adjacent services that... 
    Full time
    Worldwide

    Hulu

    San Francisco, CA
    11 hours ago
  • $117k - $152k

     ...Vector builds an offline ML platform that powers...  ...the company.Our systems operate at scale across...  ...product intelligence, machine learning pipelines, and business...  ...for a Machine Learning Engineer to join our Offline Infrastructure...  ..., ML workflows, and distributed model training.... 
    Full time
    Work at office
    Remote work
    Worldwide

    Unity Technologies

    San Francisco, CA
    4 days ago
  •  ...leading AI research firm located in San Francisco is seeking a Senior ML Systems Engineer to build and maintain the training framework for large-scale language models. The role involves designing distributed training solutions and improving training throughput across multi-... 
    Flexible hours

    Cohere

    San Francisco, CA
    1 day ago
  • $151.8k - $265.35k

     ...seeking SeniorMachine Learning Engineers for our GenAI Services...  ...performance generative AI systems—powering features...  ...direction, and mentor other ML engineers.Job...  ...in Computer Science, Machine Learning, or a related...  ...expertise in Kubernetes, distributed systems, and MLOps platforms... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Francisco, CA
    7 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Machine Learning Engineer - Distributed ML Systems. Be the first to apply!