Machine Learning Engineer - Distributed ML Systems
Pluralis Research
Overview
Pluralis Research carries out foundational research on Protocol Learning : multi-participant training of foundation models where no single participant has, or can ever obtain, a full copy of the model. The purpose of Protocol Learning is to facilitate the creation of community-trained and community-owned frontier models with self-sustaining economics.
We’re looking for Senior/Staff engineers with 5+ years of experience in distributed systems and ML large‑scale training. You’ll be implementing a novel substrate for training distributed ML models that work under consumer grade internet connection.
Responsibilities
Distributed Training Architecture & Optimization
Design and implement large‑scale distributed training systems optimized for heterogeneous hardware operating under low‑bandwidth, high‑latency conditions.
Develop and optimize model‑parallel training strategies (data, tensor, pipeline parallelism) with custom sharding techniques that minimize communication overhead.
Optimize GPU utilization, memory efficiency, and compute performance across distributed nodes.
Implement robust checkpointing, state synchronization, and recovery mechanisms for long‑running, fault‑prone training jobs.
Build monitoring and metrics systems to track training progress, model quality, and system bottlenecks.
Decentralized Networking & Resilience
Architect resilient training systems where nodes can fail, networks can partition, and participants can dynamically join or leave.
Design and optimize peer‑to‑peer topologies for decentralized coordination across non‑co‑located nodes.
Implement NAT traversal, peer discovery, dynamic routing, and connection lifecycle management.
Profile and optimize communication patterns to reduce latency and bandwidth overhead in multi‑participant environments.
What You’ll Bring
Strong experience building and operating distributed systems in production.
Hands‑on expertise with distributed training frameworks (FSDP, DeepSpeed, Megatron, or similar).
Deep understanding of model parallelism (data, tensor, pipeline parallelism).
Expert‑level Python with production experience (concurrency, error handling, retry logic, clean architecture).
Strong networking fundamentals: P2P systems, gRPC, routing, NAT traversal, distributed coordination.
Experience optimizing GPU workloads, memory management, and large‑scale compute efficiency.
What We Offer
Equity‑heavy compensation with meaningful ownership in a mission‑driven company
Competitive base salary for senior engineering roles in Australia
Visa sponsorship available for exceptional candidates
Remote‑first with optional access to our Melbourne hub
World‑class team — team mates were previously at Google, Amazon, Microsoft, and leading startups
Backed by Union Square Ventures and other tier‑1 investors, we’re a world‑class, deeply technical team of ML researchers and engineers. Pluralis is unapologetically ideological. We view the world as a better place if we are able to implement what we are attempting, and Protocol Learning as the only plausible approach to preventing a handful of massive corporations monopolising model development, access and release, and achieving massive economic capture. If this resonates, please apply.
#J-18808-Ljbffr- ...robotic platforms. About the Role As a Research Engineer, Distributed Data Systems, you will design and scale the infrastructure that powers... ..., distributed storage, streaming infrastructure, machine learning infrastructure while ensuring scalability, reliability...SuggestedFull timeWork at officeRelocation package
- ...Francisco is looking for a Senior Software Engineer to build scalable infrastructure for... ...of foundation models. You will design distributed training systems and optimize GPU utilization while... ...candidates have over 5 years of experience in ML infrastructure and a strong background...Suggested
- Genesis AI in San Francisco is seeking a senior ML infrastructure engineer to design and optimize distributed training systems and performance-critical components. You will profile bottlenecks, implement low‑level code (CUDA, Triton) and ensure efficient hardware utilization...Suggested
$200.8k - $251k
...company in San Francisco seeks a team member to build and optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-...SuggestedFull time$229.9k - $262.4k
...Senior Lead Software Engineer, Distributed Systems (Golang + Python on Kubernetes) Do you love building... ...within Capital One. The Machine Learning Experience Team (MLX Tech) is committed... ...pioneering and responsibly implementing AI/ML across Capital One . We achieve...SuggestedFull timePart timeInternshipLocal area$90k
Distributed Systems Software Engineer, Python / Go Join to apply for the Distributed Systems Software Engineer... ...capabilities to new clouds and developing AI/ML pipelines for automatic analysis of... ...remotely since 2004! Personal learning and development budget of USD 2,000...Full timeFreelanceInternshipLocal areaRemote workWorldwide- ...have a legal entity. Responsibilities As a senior Machine Learning Systems Engineer on the Search Platform team, you will own and drive the design... ...-based semantic search. Own end-to-end delivery of ML components from experimentation through production rollout...Work at officeLocal area
- An innovative company is seeking a Distributed Systems/ML Engineer to enhance the training throughput of its internal framework. This role involves collaborating with researchers to develop efficient video models and applying cutting-edge techniques to optimize training...
- ...where we have a legal entity.ResponsibilitiesAs a Principal ML System Engineer in the Rovo & AI Engineering org, you will contribute to... ...breakthrough quality and reliabilityDeep understanding of Machine Learning projects lifecycleExperience leading and supporting senior...Work at officeLocal area
- ...have a legal entity. Responsibilities As a Principal Machine Learning Systems Engineer on the Search Platform team, you set the technical... ...ambiguity and competing priorities across the Search Platform, ML Platform, AI Gateway, and Rovo product teams to keep critical...Work at officeLocal areaShift work
$189.6k - $237k
Scale’s ML platform (RLXF) team builds our internal distributed framework for large language model training... ...excitement about system optimizationExperience... ...systemsStrong software engineering skills, proficient in frameworks... ...retirement benefits, a learning and development stipend...Full time$160k - $250k
...the future of AI! Senior Machine Learning Engineer In order to execute... ...Everything involved in applying a ML model to a production use... ...product and core backend systems by suggesting and executing... ...code and training across distributed systems You have an ability...Full time- ...problems, care about elegant systems design, and want to build... ...Role We’re looking for a Machine Learning Engineer who loves getting close to... ...of running large-scale ML systems, and thrives in fast... ...preferred). ~ Familiarity with distributed training frameworks (e.g.,...Full timeWork at office
- ...reinvent the way people learn, starting with... ...Ventures, and more, with a distributed team across San Francisco... ...We’re hiring an ML Engineer, Assessments to help... ...best-in-class assessment systems across multiple products... .../Learning Design) , Machine Learning, Product, and...Full timeLive inImmediate start
$244k - $320k
...AI-powered personalization engine delivers bespoke experiences... ...customers. With a distributed global workforce and employee... ...! About the Role Our Machine Learning Engineering team powers personalized... ...operating production-grade ML systems that drive real-time...Full time- ...The role We’re looking for a Machine Learning Engineer to build and ship consumer-facing AI systems that power personalization,... ...contribute Build and deploy ML models that improve sleep experiences... ...with data tooling (SQL, distributed compute such as Spark/Ray, and...Full timeImmediate startWorldwideNight shift
$150k - $190k
...simulation software stack for engineering and manufacturing across... ...We're Looking For As a Machine Learning Engineer in Delivery, you... ...and used. You’ve shipped ML systems end-to-end and at scale: you... ...g., TensorFlow, MLFlow) ~ Distributed computing frameworks (e.g.,...Remote jobFull timeFlexible hours- ...hands-on support from AMD engineers the team is scaling... ...and Reinforcement Learning , sandbox environments... ...Build and maintain distributed training infrastructure... ...writing robust, performant systems) Experience with... ...end-to-end production ML systems with monitoring...Full timeFlexible hours
- ...inspires future generations. Senior Machine Learning Engineer Primary: Bay Area (San Francisco... ...constraints, so you get to stand up ML systems the right way from day one. You're... ...Thinker: You've designed scalable distributed systems and data-intensive applications...Full timeWork at officeRelocation packageFlexible hours2 days per week
$200k - $400k
...understanding and designing AI systems that people can trust.... ...researchers and engineers from organizations... ...We’re looking for Machine Learning Engineers to help build... ...years of experience in ML infra, research engineering... ..., PyTorch or Jax, and distributed systems. ~...Full time$225k - $300k
...Machine Learning Engineer About Latent Health Healthcare today is only... ...history is scattered across systems that don’t communicate. Physicians... ...extraordinary depth. ML at Latent Health The... ...Experience building distributed systems or large-scale data...Full timeWork at officeImmediate start$180k - $270k
...privacy protection. To learn more about Plaud,... ...-low-latency inference engines for large language models... ...intersection between the core ML training team and the... ...genuinely enjoy the systems-engineering challenge... ...accuracy. Large-Scale Distributed Systems: Deploying...Full timeWork at officeWorldwide- ...We’re a team of engineers, neuroscientists,... ...intense, fast-paced learning. You will:... ...implement the best machine learning approaches... ...design end-to-end systems Explore new model... ...years of applied ML research or development... ...Experience with distributed or large-scale...Full time
- ...operators across hospital and health systems, pharmacies and payors to... ...Role We're looking for a Machine Learning Engineer to design, build, and deploy production-grade ML systems that power the next... ...data pipelines using SQL and distributed data processing tools You'...Full timeWork at officeRemote workFlexible hours2 days per week
$55 per hour
...rely heavily on legacy ERP systems that don't talk to each... ...of the Acquired podcast to learn how credit card networks did... ...how we do modeling, data engineering, and production ML at Wholesail for years to... ...APIs, and reasoning about distributed systems that consume your...Full timeTemporary workWork at officeLocal area- Senior ML Systems Engineer, Frameworks & Tooling at Cohere Our mission is to scale intelligence to serve humanity. We’re training and deploying... ...role sits at the intersection of large‑scale training, distributed systems, and HPC infrastructure. You will design and...Full timeWork at officeRemote workFlexible hours
$148.7k - $199.4k
Job Posting Title:Senior Machine Learning Engineer - ESPNReq ID:10150610Job Description... ...worldwide advertising and distribution to maximize flexibility and... ...distributed data and ML infrastructure that supports... ..., feature computation systems, and ML‑adjacent services that...Full timeWorldwide$117k - $152k
...Vector builds an offline ML platform that powers... ...the company.Our systems operate at scale across... ...product intelligence, machine learning pipelines, and business... ...for a Machine Learning Engineer to join our Offline Infrastructure... ..., ML workflows, and distributed model training....Full timeWork at officeRemote workWorldwide- ...leading AI research firm located in San Francisco is seeking a Senior ML Systems Engineer to build and maintain the training framework for large-scale language models. The role involves designing distributed training solutions and improving training throughput across multi-...Flexible hours
$151.8k - $265.35k
...seeking SeniorMachine Learning Engineers for our GenAI Services... ...performance generative AI systems—powering features... ...direction, and mentor other ML engineers.Job... ...in Computer Science, Machine Learning, or a related... ...expertise in Kubernetes, distributed systems, and MLOps platforms...Full timeTemporary workLocal areaWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Machine Learning Engineer - Distributed ML Systems. Be the first to apply!
- graduate machine learning engineer San Francisco, CA
- senior ml engineer San Francisco, CA
- machine learning engineer San Francisco, CA
- machine learning ai engineer San Francisco, CA
- entry level machine learning engineer San Francisco, CA
- machine learning software engineer San Francisco, CA
- ai ml engineer San Francisco, CA
- computer vision machine learning engineer San Francisco, CA
- junior machine learning research engineer San Francisco, CA
- lead system engineer San Francisco, CA


