Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Infra Engineer, Modeling

physicalintelligence

Physical Intelligence is bringing general-purpose AI into the physical world. We are a group of engineers, scientists, roboticists, and company builders developing foundation models and learning algorithms to power the robots of today and the physically-actuated devices of the future. In this role you will help scale and optimize our training systems and core model code. You'll own critical infrastructure for large-scale training, from managing GPU/TPU compute and job orchestration to building reusable and efficient JAX training pipelines. You'll work closely with researchers and model engineers to translate ideas into experiments-and those experiments into production training runs. This is a hands-on, high-leverage role at the intersection of ML, software engineering, and scalable infrastructure. The Team The ML Infrastructure team supports and accelerates PI's core modeling efforts by building the systems that make large-scale training reliable, reproducible, and fast. The team works closely with research, data, and platform engineers to ensure models can scale from prototype to production-grade training runs. In This Role You Will Own training/inference infrastructure: Design, implement, and maintain systems for large-scale model training, including scheduling, job management, checkpointing, and metrics/logging. Scale distributed training: Work with researchers to scale JAX-based training across TPU and GPU clusters with minimal friction. Optimize performance: Profile and improve memory usage, device utilization, throughput, and distributed synchronization. Enable rapid iteration: Build abstractions for launching, monitoring, debugging, and reproducing experiments. Partner with researchers: Translate research needs into infra capabilities and guide best practices for training at scale. Contribute to core training code: Evolve JAX model and training code to support new architectures, modalities, and evaluation metrics. What We Hope You'll Bring Strong software engineering fundamentals and experience building ML training infrastructure or internal platforms. Hands-on large-scale training experience in JAX (preferred), PyTorch. Familiarity with distributed training, multi-host setups, data loaders, and evaluation pipelines. Experience managing training workloads on cloud platforms (e.g., SLURM, Kubernetes, GCP TPU/GKE, AWS). Ability to debug and optimize performance bottlenecks across the training stack. Strong cross-functional communication and ownership mindset. Bonus Points If You Have Deep ML systems background (e.g., training compilers, runtime optimization, custom kernels). Experience operating close to hardware (GPU/TPU performance tuning). Background in robotics, multimodal models, or large-scale foundation models. Experience designing abstractions that balance researcher flexibility with system reliability. Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. #J-18808-Ljbffr physicalintelligence

Vacancy posted 13 hours ago
Similar jobs that could be interesting for youBased on the ML Infra Engineer, Modeling in San Francisco, CA vacancy
  •  ...thinkers. We don't believe culture can be engineered - but when it falls into place, it's a...  ...Position Overview We're looking for an ML infrastructure engineer to help design, build...  ...collection to dataset curation to large-scale model training and deployment, help us build... 
    Suggested
    Local area

    Humble Robotics

    San Francisco, CA
    23 hours ago
  • $190k - $205k

     ...rates and noise characteristics. Design models that are robust, interpretable, and...  ...electrical, and visual signals Production Engineering Write clean, scalable, well-tested Python...  ...shared codebase. Build end-to-end ML pipelines including data processing, feature... 
    Suggested
    Full time
    Live in

    Gridware

    San Francisco, CA
    23 hours ago
  •  ...Platform, an agentic operating system for sales. You will own models end-to-end—from training and fine-tuning to production deployment...  ...as needed, and collaborate closely with the founders and engineering teams to shape what we build next. #J-18808-Ljbffr Hyperbound
    Suggested

    Hyperbound

    San Francisco, CA
    23 hours ago
  •  ...and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate will have hands-on experience with modern inference frameworks and a solid understanding... 
    Suggested

    Reflection AI

    San Francisco, CA
    23 hours ago
  • Harrison Clarke is seeking a technically talented engineer to join a well-funded AI startup in San Francisco, building cutting-edge ML systems at the intersection of data pipelines, model training, and production inference. You will own the full ML lifecycle—from data generation... 
    Suggested

    Harrison Clarke

    San Francisco, CA
    23 hours ago
  • Responsibilities Design, deploy, and maintain large distributed ML training and inference clusters Develop efficient, scalable end-to-end pipelines to manage petabyte-scale datasets and model training throughout the entire ML lifecycle Research and test various training... 

    Kindredventures

    San Francisco, CA
    2 days ago
  •  ...is seeking a passionate Machine Learning Engineer to design, implement, and validate algorithmic...  ...systems. Join a growing New Verticals ML team and help build ML-powered products...  ...business. You will work on production ML models using Python (PyTorch, TensorFlow or similar... 

    DoorDash

    San Francisco, CA
    23 hours ago
  • Physical Intelligence is seeking an hands-on ML Infrastructure Engineer to scale and optimize our training systems and core model code. You will own critical infrastructure for large-scale training, managing GPU/TPU compute, job orchestration, and building efficient JAX... 

    physicalintelligence

    San Francisco, CA
    13 hours ago
  •  ...thinkers. We don’t believe culture can be engineered - but when it falls into place, it’s a once...  .... Position Overview We’re looking for an ML engineer to design, train, and ship the vision-language-action (VLA) foundation model at the core of Humble’s autonomous driving... 
    Local area

    Humble Robotics

    San Francisco, CA
    3 days ago
  • David Joseph & Company seeks a talented ML infrastructure engineer to own the distributed training and inference backbone for a large foundation model. You will stand up clusters, build data pipelines for petabyte-scale datasets, and squeeze GPU performance across model... 
    Relocation package

    David Joseph & Company

    San Francisco, CA
    4 days ago
  • $161.26k - $332.01k

    Pinterest is seeking a skilled Research Engineer to join our visual modeling team focusing on generative models, including text-to-image development. Candidates should have significant experience in computer vision and strong background in diffusion models. The role promotes... 

    Pinterest

    San Francisco, CA
    3 days ago
  •  ...evals that keep our speech and multimodal models honest in production. Harness the...  ...slide decks — partner with research and infra to prototype, train, and deploy state-...  ...Expert-level PyTorch. Proven software engineer who loves ML; comfortable writing production code... 
    Full time
    Contract work
    Flexible hours
    Shift work

    Sesame, L.l.c.

    San Francisco, CA
    23 hours ago
  • $250k - $295k

     ...protection. We use advanced physics and AI to model catastrophic risk at the asset level, then...  .... The real product is a scalable risk engine, our Stand World Model . We stay when traditional...  ...outcomes Build on and extend scalable ML infrastructure Partner with Stand’s... 
    Full time
    Temporary work
    H1b
    Work at office
    Remote work
    Visa sponsorship
    Work visa
    Flexible hours

    Stand Insurance

    San Francisco, CA
    23 hours ago
  •  ...Perplexity is seeking experienced ML engineers to design, build, and optimize the recommendation systems that power core experiences on...  ...drive impact on core user and business metrics. Build user modeling that captures intent, preference, and propensity, and powers more... 
    Full time

    Perplexity®️

    San Francisco, CA
    23 hours ago
  • A leading tech company in San Francisco seeks a Machine Learning Engineer to build and maintain infrastructure for large-scale model training. In this hands-on role, you will design systems, work closely with researchers, and optimize training processes. Candidates should... 

    Monograph

    San Francisco, CA
    23 hours ago
  •  ...Francisco is looking for a Senior Software Engineer to build scalable infrastructure for...  ...scale training and fine-tuning of foundation models. You will design distributed training...  ...candidates have over 5 years of experience in ML infrastructure and a strong background in... 

    Baseten

    San Francisco, CA
    23 hours ago
  • OpenAI is seeking an Applied Machine Learning Engineer in San Francisco, CA to bridge cutting-...  ...engineers to optimize, deploy, and scale models like GPT-4 and custom architectures. The...  ...engineering with at least 2 years in ML systems, plus deep knowledge of distributed... 

    AI Breaking Wire

    San Francisco, CA
    23 hours ago
  • Parallel Bio in San Francisco is looking for a candidate to own the training pipeline behind models essential for both their search stack and agents. You will be responsible for building pathways from real product usage to high-quality training data while rigorously fine... 

    Parallel Bio

    San Francisco, CA
    3 days ago
  • SF Tensor in San Francisco is building the fastest GPU compiler and enterprise model-foundry to accelerate AI research. This role owns the post-training pipeline end-to-end, from data curation through RL, distillation and deployment, working with customers and domain experts... 
    Relocation package

    San Francisco Tensor Company

    San Francisco, CA
    4 days ago
  • Parallel is seeking a professional who will own the training pipeline behind models that support both the search stack and agents. Responsibilities include building pathways from product usage to high-quality training data, rigorously fine-tuning models, and shipping them... 

    Parallel

    San Francisco, CA
    3 days ago
  • Jaide Health is seeking an engineer for their Model Efficiency team in San Francisco. The role focuses on building reliable ML systems while enhancing core performance metrics across model execution. You'll work with advanced performance techniques such as GPU/CUDA optimizations... 
    Remote job

    Jaide Health

    San Francisco, CA
    3 days ago
  • $170k - $216k

     ...from a diverse set of sensors, enabling engineers like you to (1) develop methods for efficiently...  ...scale real-world data, to (2) develop models and model training at scale, to (3)...  ...Experience with Python ~ Experience with ML frameworks like PyTorch or JAX We prefer... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    23 hours ago
  •  ...documents codebases autonomously. As a Founding ML Engineer , you’ll architect the intelligence...  ...workflows. You’ll work across product, infra, and full-stack teams to embed real-time...  ...Develop and refine agent behavior models for planning, implementation, and review... 
    Remote work
    Flexible hours

    Kodezi Inc.

    San Francisco, CA
    4 days ago
  • $298k - $368k

     ...of miles of driving data from a diverse set of sensors, enabling engineers like you to (1) develop methods for efficiently and continuously learning from large scale real-world data, to (2) develop models and model training at scale, to (3) analyze real-world behavior and... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    23 hours ago
  • $250k

     ...without traditional infrastructure limitations. As a Senior ML Infrastructure Engineer, the successful candidate will help build and scale...  ...Improve and maintain inference infrastructure, including model packaging, deployment, and serving optimisation Collaborate... 
    Full time
    San Francisco, CA
    more than 2 months ago
  •  ...We’re looking for a Machine Learning Engineer to build and ship consumer-facing AI systems...  ...intelligence.” You’ll work across data, modeling, product, and engineering to translate research...  ...you’ll contribute Build and deploy ML models that improve sleep experiences... 
    Full time
    Immediate start
    Worldwide
    Night shift

    Eight Sleep

    San Francisco, CA
    13 days ago
  • $232.65k - $387.75k

    As the Sr. Director of AI/ML Engineering you will: Lead a machine learning engineering team specializing in medical imaging and multimodal...  ...years proficiency with standard deep learning algorithms and model architectures Experience developing high-performing teams through... 

    Scorpion Therapeutics

    San Francisco, CA
    2 days ago
  • A leading AI evaluation platform in San Francisco is looking for a Senior Software Engineer specializing in ML infrastructure. The successful candidate will design and develop robust real-time data and API systems, enabling insights for researchers and developers. Ideal... 

    LMArena

    San Francisco, CA
    2 days ago
  •  ...developing next-generation multimodal AI models and a proprietary, high-efficiency serving...  ...from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the...  ...applications. About the role As an ML Engineer at Sciforium, you will operate... 
    Full time
    Flexible hours

    Sciforium

    San Francisco, CA
    23 hours ago
  •  ...are at an inflection point where advances in speech, language models, and clinical AI can fundamentally change how clinicians interact...  ...: their patients. Position Overview We are hiring two ML Engineers / Researchers to help build the next generation of Knowtex's... 
    Full time

    Knowtex

    San Francisco, CA
    23 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Infra Engineer, Modeling. Be the first to apply!