Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Infra Engineer (Supercomputing)

Physical Intelligence

Physical Intelligence builds general-purpose AI for the physical world. Training our models requires orchestrating thousands of accelerators across a heterogeneous fleet of GPU and TPU clusters — spanning different hardware generations, cloud providers, and cluster topologies. Today, researchers often need to know which cluster to target, what resources are available, and how to configure their jobs accordingly. That doesn't scale. We need a scheduling and compute layer that makes the right placement decision automatically — routing jobs to the best cluster based on availability, hardware fit, cost, and priority — so researchers can focus entirely on the science. This role owns that problem end-to-end: the scheduling systems, the placement logic, the cluster management layer, and the operational tooling that keeps it all running. This is not cloud DevOps. It's not about standing up clusters and walking away. It's a systems role for people who care about intelligent resource allocation, utilization, fault tolerance, and making large-scale distributed training seamless. The Team The ML Infrastructure team supports and accelerates PI’s core modeling efforts by building the systems that make large-scale training reliable, reproducible, and fast. You will work closely with ML Infra (training systems), data platform, and research teams to ensure compute scheduling is never the bottleneck. In This Role You Will Own Intelligent Job Scheduling and Placement : Design and build multi-tenant scheduling systems that automatically place training jobs on the best available cluster based on hardware requirements, topology, availability, cost, and priority. Support fair resource sharing across teams and projects with quota management, priority tiers, and preemption policies. Abstract away cluster differences so researchers submit jobs without needing to know where they will land. Scale Multi-cluster Orchestration : Build the control plane that manages the job lifecycle across diverse clusters (mixed GPU/TPU, multi-generation hardware, on-prem/cloud) and enables seamless job migration, failover, and re-scheduling. Optimize Accelerator Utilization and Efficiency : Monitor and optimize GPU/TPU utilization across the entire fleet. Implement priority, preemption, queueing, and fairness policies that balance research velocity with cost efficiency. Ensure Scaling and Stability : Implement fault detection, automatic recovery, and resilience for long-running multi-node training jobs. Manage health checking, node management, and scaling to thousands of accelerators. Support Inference and Robot Deployment : Extend scheduling and orchestration to inference workloads, including deploying models to edge devices on physical robots. Enhance Observability and Developer Experience : Build the dashboards, alerting, SLOs, and debugging tools necessary for researchers to understand job status and for the team to ensure high scheduling quality and cluster reliability. What We Hope You’ll Bring We’re intentionally flexible on exact background, but strong candidates usually have: Strong software engineering fundamentals Experience building or operating job scheduling / resource management systems at scale Experience with large-scale compute clusters (GPU and/or TPU) Familiarity with schedulers and orchestration systems (SLURM, Kubernetes, GKE, K3S, or internal equivalents) Comfort reasoning about resource allocation, bin-packing, priority scheduling, and multi-tenancy Understanding of how ML training workloads behave — long-running, multi-node, sensitive to stragglers, topology-dependent A bias toward owning systems end-to-end, from design to operation Enjoy working closely with researchers and unblocking fast-moving projects Bonus Points If You Have Experience building multi-cluster or federated scheduling systems Experience with TPU infrastructure (GCP TPU slices, Multislice, GKE) Background in cluster resource managers (Borg, YARN, Mesos, or custom schedulers) Linux systems engineering, networking, and infrastructure-as-code NCCL/collective communication and topology-aware placement Experience with capacity planning and cloud cost optimization at scale Familiarity with JAX, PyTorch, or similar ML frameworks at the runtime/systems level In this role you will help scale and optimize our training systems and core model code. You’ll own critical infrastructure for large-scale training, from managing GPU/TPU compute and job orchestration to building reusable and efficient JAX training pipelines. You’ll work closely with researchers and model engineers to translate ideas into experiments—and those experiments into production training runs. This is a hands-on, high-leverage role at the intersection of ML, software engineering, and scalable infrastructure. The Team The ML Infrastructure team supports and accelerates PI’s core modeling efforts by building the systems that make large-scale training reliable, reproducible, and fast. The team works closely with research, data, and platform engineers to ensure models can scale from prototype to production-grade training runs. In This Role You Will Own training/inference infrastructure: Design, implement, and maintain systems for large-scale model training, including scheduling, job management, checkpointing, and metrics/logging. Scale distributed training: Work with researchers to scale JAX-based training across TPU and GPU clusters with minimal friction. Optimize performance: Profile and improve memory usage, device utilization, throughput, and distributed synchronization. Enable rapid iteration: Build abstractions for launching, monitoring, debugging, and reproducing experiments. Manage compute resources: Ensure efficient allocation and utilization of cloud-based GPU/TPU compute while controlling cost. Partner with researchers: Translate research needs into infra capabilities and guide best practices for training at scale. Contribute to core training code: Evolve JAX model and training code to support new architectures, modalities, and evaluation metrics. What We Hope You’ll Bring Strong software engineering fundamentals and experience building ML training infrastructure or internal platforms. Hands-on large-scale training experience in JAX (preferred), PyTorch. Familiarity with distributed training, multi-host setups, data loaders, and evaluation pipelines. Experience managing training workloads on cloud platforms (e.g., SLURM, Kubernetes, GCP TPU/GKE, AWS). Ability to debug and optimize performance bottlenecks across the training stack. Strong cross-functional communication and ownership mindset. Bonus Points If You Have Deep ML systems background (e.g., training compilers, runtime optimization, custom kernels). Experience operating close to hardware (GPU/TPU performance tuning). Background in robotics, multimodal models, or large-scale foundation models. Experience designing abstractions that balance researcher flexibility with system reliability. Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. #J-18808-Ljbffr Physical Intelligence

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the ML Infra Engineer (Supercomputing) in San Francisco, CA vacancy
  • Physical Intelligence seeks a systems-focused ML infrastructure engineer to own scheduling, placement, and cluster management for large-scale model...  ...and ML engineers to translate workloads into scalable infra, optimize utilization, and build dashboards, tests, and tooling... 
    Suggested

    Physical Intelligence

    San Francisco, CA
    2 days ago
  •  ...work closely with researchers and model engineers to translate ideas into experiments—and those...  ...‑leverage role at the intersection of ML, software engineering, and scalable infrastructure...  ...: Translate research needs into infra capabilities and guide best practices for... 
    Suggested
    Full time

    Monograph

    San Francisco, CA
    12 hours ago
  •  ...Ship models, not slide decks — partner with research and infra to prototype, train, and deploy state-of-the-art voice models...  ...Qualifications: Expert-level PyTorch. Proven software engineer who loves ML; comfortable writing production code across the stack. Hands... 
    Suggested
    Full time
    Contract work
    Flexible hours
    Shift work

    Sesame, L.l.c.

    San Francisco, CA
    12 hours ago
  • A leading AI evaluation platform in San Francisco is looking for a Senior Software Engineer specializing in ML infrastructure. The successful candidate will design and develop robust real-time data and API systems, enabling insights for researchers and developers. Ideal... 
    Suggested

    LMArena

    San Francisco, CA
    2 hours ago
  • A cutting-edge AI technology company based in San Francisco is seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate...
    Suggested

    Reflection AI

    San Francisco, CA
    3 days ago
  •  ...technology company in San Francisco is looking for a Senior Software Engineer to build scalable infrastructure for large‑scale training and...  .... Ideal candidates have over 5 years of experience in ML infrastructure and a strong background in distributed training frameworks... 

    Baseten

    San Francisco, CA
    3 days ago
  • A leading tech company in San Francisco seeks a Machine Learning Engineer to build and maintain infrastructure for large-scale model training. In this hands-on role, you will design systems, work closely with researchers, and optimize training processes. Candidates should... 

    Monograph

    San Francisco, CA
    3 days ago
  • $250k - $400k

     ...research in isolation. It's building the engine that research runs on. You'll work closely...  ...Experience building and scaling ML systems in production Strong background...  ...Principal Roles available: ML Engineer, ML Infra, Research Engineers & Research Scientists... 
    Remote work

    techire ai

    San Francisco, CA
    3 days ago
  • $250k

     ...enables AI teams to access scalable compute environments without traditional infrastructure limitations. As a Senior ML Infrastructure Engineer, the successful candidate will help build and scale Kubernetes-based machine learning platforms supporting large-scale training... 
    Full time
    San Francisco, CA
    a month ago
  •  ...heals, evolves, and documents codebases autonomously. The Founding ML Engineer will architect intelligence powering autonomous pull requests, long-term memory, and self-improving workflows across product, infra, and full-stack teams. This in-person role is based in San... 
    Remote work

    Kodezi Inc.

    San Francisco, CA
    4 days ago
  •  ..., evolves, and documents codebases autonomously. As a Founding ML Engineer , you’ll architect the intelligence powering autonomous pull requests...  ..., and self-improving workflows. You’ll work across product, infra, and full-stack teams to embed real-time decision-making into... 
    Remote work
    Flexible hours

    Kodezi Inc.

    San Francisco, CA
    2 days ago
  •  ...We’re a team of AI researchers, designers, growth experts, and engineers rethinking human-computer interaction from the ground up. We value...  ...A, this is just the beginning. About the Role As a ML engineer at Wispr, you’ll play a crucial role in building the first... 
    Full time

    Wispr Flow

    San Francisco, CA
    12 hours ago
  • $137.1k - $201.6k

     ...leverage artificial intelligence and advanced ML, deep learning techniques to power...  ...We’re looking for a Machine Learning Engineer to help design, build, optimize and scale...  ...pipelines in partnership with Platform and Infra teams. Write high-quality, maintainable... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash Usa

    San Francisco, CA
    12 hours ago
  • $200k - $400k

     ...with a team of the world’s top interpretability researchers and engineers from organizations like OpenAI and DeepMind. We’ve raised $59M...  ...Required experience ~5+ years of experience in ML infra, research engineering, or systems programming. ~ Comfort working... 
    Full time

    Goodfire

    San Francisco, CA
    12 hours ago
  • Reddit is hiring Machine Learning Engineers (IC4) to build and evolve the auction, bidding and budgeting systems...  ...Ads Optimization, Product, Data Science and Infra. This senior role requires 3-5+ years in production ML, strong Python/Java/Go skills, and experience with... 

    Reddit, Inc.

    San Francisco, CA
    3 days ago
  •  ...leverage artificial intelligence and advanced ML, deep learning techniques to power...  ...RoleWe’re looking for a Machine Learning Engineer to help design, build, optimize and scale...  ...pipelines in partnership with Platform and Infra teams.Write high-quality, maintainable code... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    1 day ago
  • Senior ML Systems Engineer, Frameworks & Tooling at Cohere Our mission is to scale intelligence to serve humanity. We’re training and deploying...  ..., and developer ergonomics. Collaborate closely with infra teams to ensure Slurm setups, container environments, and hardware... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Cohere

    San Francisco, CA
    1 day ago
  •  ...training software, bridging research and production while writing production-grade code alongside researchers. You’ll work with Python, ML frameworks, and distributed training tools. Join a world-class team with extensive compute resources, remote-friendly policies, and... 
    Remote work
    Worldwide

    Cohere

    San Francisco, CA
    4 days ago
  •  ...building production-ready systems from research to real-world use, backed by UC Berkeley/UCSF origins. The role focuses on scalable ML infrastructure for fast, secure clinical workflows. Join a team that values clarity, ownership, and judgment as you deploy GPU inference... 

    Voio

    Berkeley, CA
    4 days ago
  •  ...first commercially available AI Co-Scientist. It is a discovery engine that transforms messy biological data into insights in minutes. Scientists...  ...to patient outcomes. ABOUT THE ROLE We are hiring an ML Engineer, Discovery Applications to build the high level, end-to-... 
    Full time
    Work at office

    Mithrl

    San Francisco, CA
    12 hours ago
  •  ...pretraining science. We build foundational understanding of models to advance the frontier of intelligence. About the role: As a ML Engineer, you’ll build and operate the infrastructure that makes cutting-edge machine learning research possible. At Tilde, we believe... 
    Full time
    Internship

    Tilde Research

    San Francisco, CA
    12 hours ago
  •  ...governance, maintain auditability, and deliver reliable outcomes at scale.    About the Role We are looking for a visionary Senior ML Engineer who will bridge the gap between high-level architecture and hands-on execution, specifically focusing on simplifying enterprise... 
    Full time
    Shift work

    Palm Venture Studios

    San Francisco, CA
    12 hours ago
  •  ...first commercially available AI Co-Scientist. It is a discovery engine that transforms messy biological data into insights in minutes. Scientists...  ...to patient outcomes. ABOUT THE ROLE We are hiring an ML Engineer, Analysis and Simulation to build the core analytical... 
    Full time
    Work at office

    Mithrl

    San Francisco, CA
    12 hours ago
  • $171k

     ...Security team, part of the Core Security Engineering organization, is building the foundation...  ...rules and manual approvals toward real-time, ML-driven access decisions that secure both...  ...# Familiarity with large-scale data/infra systems (Kafka, Hive, Spark, Flink, Pinot... 
    Full time
    Work at office
    Remote work

    Uber Technologies Inc

    San Francisco, CA
    5 days ago
  • $129k - $198.4k

    Job DescriptionRole: As an AI/ML Engineer on the Metrics Frameworks team, part of the Simulation, Evaluation, and Data organization, you will...  ...the organization. Collaborate with other frameworks and data infra teams to build and deploy tools to improve productivity. Work... 
    Full time
    Local area
    Work from home

    General Motors

    San Francisco, CA
    12 hours ago
  •  ...Machine Learning Engineer, Applied AI As a machine learning engineer you'll join our product innovations team and work across the full applied ML stack - deploying models, building the evaluation systems that tell us whether they actually work, and making the data and... 
    Work at office
    Work from home
    Home office

    CreatorIQ

    San Francisco, CA
    3 days ago
  •  ...customers. How we think → The team is still small enough that every person shapes what gets built. About the Role As a ML engineer at Wispr, you'll play a crucial role in building the first capable, habit forming voice interface that scales to a billion users.... 
    H1b
    Work at office
    Remote work
    Relocation
    Visa sponsorship
    Flexible hours

    Wispr Flow

    San Francisco, CA
    3 days ago
  •  ...Founding Ml Engineer Skills: Python, PyTorch, NLP, LLMs, Information Retrieval, Entity Resolution, Text Classification We're building the gateway to the internet for AI agents. Our APIs already power hundreds of customers — and we went from 0 to $7M ARR in our first... 

    Crustdata (YC F24)

    San Francisco, CA
    2 days ago
  •  ...transformers and spatial models run efficiently on both cloud and edge compute resources. Learn more at   About the Role As an ML / DevOps Engineer, you will play a pivotal role in advancing our infrastructure, scaling enterprise deployment workflows, and refining... 
    Work at office

    Zensors

    San Francisco, CA
    more than 2 months ago
  •  ...Platform Raindrop is the monitoring platform for AI agents. Engineering teams at Fortune 100s and the fastest-growing AI companies use...  ...of requests a day + growing. Architect, implement, and scale ML pipelines Quick iteration without compromising on quality... 
    Temporary work

    Raindrop

    San Francisco, CA
    12 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Infra Engineer (Supercomputing). Be the first to apply!