ML Infra Engineer, Modeling
physicalintelligence
Physical Intelligence is bringing general-purpose AI into the physical world. We are a group of engineers, scientists, roboticists, and company builders developing foundation models and learning algorithms to power the robots of today and the physically-actuated devices of the future. In this role you will help scale and optimize our training systems and core model code. You'll own critical infrastructure for large-scale training, from managing GPU/TPU compute and job orchestration to building reusable and efficient JAX training pipelines. You'll work closely with researchers and model engineers to translate ideas into experiments-and those experiments into production training runs. This is a hands-on, high-leverage role at the intersection of ML, software engineering, and scalable infrastructure. The Team The ML Infrastructure team supports and accelerates PI's core modeling efforts by building the systems that make large-scale training reliable, reproducible, and fast. The team works closely with research, data, and platform engineers to ensure models can scale from prototype to production-grade training runs. In This Role You Will Own training/inference infrastructure: Design, implement, and maintain systems for large-scale model training, including scheduling, job management, checkpointing, and metrics/logging. Scale distributed training: Work with researchers to scale JAX-based training across TPU and GPU clusters with minimal friction. Optimize performance: Profile and improve memory usage, device utilization, throughput, and distributed synchronization. Enable rapid iteration: Build abstractions for launching, monitoring, debugging, and reproducing experiments. Partner with researchers: Translate research needs into infra capabilities and guide best practices for training at scale. Contribute to core training code: Evolve JAX model and training code to support new architectures, modalities, and evaluation metrics. What We Hope You'll Bring Strong software engineering fundamentals and experience building ML training infrastructure or internal platforms. Hands-on large-scale training experience in JAX (preferred), PyTorch. Familiarity with distributed training, multi-host setups, data loaders, and evaluation pipelines. Experience managing training workloads on cloud platforms (e.g., SLURM, Kubernetes, GCP TPU/GKE, AWS). Ability to debug and optimize performance bottlenecks across the training stack. Strong cross-functional communication and ownership mindset. Bonus Points If You Have Deep ML systems background (e.g., training compilers, runtime optimization, custom kernels). Experience operating close to hardware (GPU/TPU performance tuning). Background in robotics, multimodal models, or large-scale foundation models. Experience designing abstractions that balance researcher flexibility with system reliability. Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. #J-18808-Ljbffr physicalintelligence
- ...thinkers. We don't believe culture can be engineered - but when it falls into place, it's a... ...Position Overview We're looking for an ML infrastructure engineer to help design, build... ...collection to dataset curation to large-scale model training and deployment, help us build...SuggestedLocal area
$190k - $205k
...rates and noise characteristics. Design models that are robust, interpretable, and... ...electrical, and visual signals Production Engineering Write clean, scalable, well-tested Python... ...shared codebase. Build end-to-end ML pipelines including data processing, feature...SuggestedFull timeLive in- ...Platform, an agentic operating system for sales. You will own models end-to-end—from training and fine-tuning to production deployment... ...as needed, and collaborate closely with the founders and engineering teams to shape what we build next. #J-18808-Ljbffr HyperboundSuggested
- ...and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate will have hands-on experience with modern inference frameworks and a solid understanding...Suggested
- Harrison Clarke is seeking a technically talented engineer to join a well-funded AI startup in San Francisco, building cutting-edge ML systems at the intersection of data pipelines, model training, and production inference. You will own the full ML lifecycle—from data generation...Suggested
- Responsibilities Design, deploy, and maintain large distributed ML training and inference clusters Develop efficient, scalable end-to-end pipelines to manage petabyte-scale datasets and model training throughout the entire ML lifecycle Research and test various training...
- ...is seeking a passionate Machine Learning Engineer to design, implement, and validate algorithmic... ...systems. Join a growing New Verticals ML team and help build ML-powered products... ...business. You will work on production ML models using Python (PyTorch, TensorFlow or similar...
- Physical Intelligence is seeking an hands-on ML Infrastructure Engineer to scale and optimize our training systems and core model code. You will own critical infrastructure for large-scale training, managing GPU/TPU compute, job orchestration, and building efficient JAX...
- ...thinkers. We don’t believe culture can be engineered - but when it falls into place, it’s a once... .... Position Overview We’re looking for an ML engineer to design, train, and ship the vision-language-action (VLA) foundation model at the core of Humble’s autonomous driving...Local area
- David Joseph & Company seeks a talented ML infrastructure engineer to own the distributed training and inference backbone for a large foundation model. You will stand up clusters, build data pipelines for petabyte-scale datasets, and squeeze GPU performance across model...Relocation package
$161.26k - $332.01k
Pinterest is seeking a skilled Research Engineer to join our visual modeling team focusing on generative models, including text-to-image development. Candidates should have significant experience in computer vision and strong background in diffusion models. The role promotes...- ...evals that keep our speech and multimodal models honest in production. Harness the... ...slide decks — partner with research and infra to prototype, train, and deploy state-... ...Expert-level PyTorch. Proven software engineer who loves ML; comfortable writing production code...Full timeContract workFlexible hoursShift work
$250k - $295k
...protection. We use advanced physics and AI to model catastrophic risk at the asset level, then... .... The real product is a scalable risk engine, our Stand World Model . We stay when traditional... ...outcomes Build on and extend scalable ML infrastructure Partner with Stand’s...Full timeTemporary workH1bWork at officeRemote workVisa sponsorshipWork visaFlexible hours- ...Perplexity is seeking experienced ML engineers to design, build, and optimize the recommendation systems that power core experiences on... ...drive impact on core user and business metrics. Build user modeling that captures intent, preference, and propensity, and powers more...Full time
- A leading tech company in San Francisco seeks a Machine Learning Engineer to build and maintain infrastructure for large-scale model training. In this hands-on role, you will design systems, work closely with researchers, and optimize training processes. Candidates should...
- ...Francisco is looking for a Senior Software Engineer to build scalable infrastructure for... ...scale training and fine-tuning of foundation models. You will design distributed training... ...candidates have over 5 years of experience in ML infrastructure and a strong background in...
- OpenAI is seeking an Applied Machine Learning Engineer in San Francisco, CA to bridge cutting-... ...engineers to optimize, deploy, and scale models like GPT-4 and custom architectures. The... ...engineering with at least 2 years in ML systems, plus deep knowledge of distributed...
- Parallel Bio in San Francisco is looking for a candidate to own the training pipeline behind models essential for both their search stack and agents. You will be responsible for building pathways from real product usage to high-quality training data while rigorously fine...
- SF Tensor in San Francisco is building the fastest GPU compiler and enterprise model-foundry to accelerate AI research. This role owns the post-training pipeline end-to-end, from data curation through RL, distillation and deployment, working with customers and domain experts...Relocation package
- Parallel is seeking a professional who will own the training pipeline behind models that support both the search stack and agents. Responsibilities include building pathways from product usage to high-quality training data, rigorously fine-tuning models, and shipping them...
- Jaide Health is seeking an engineer for their Model Efficiency team in San Francisco. The role focuses on building reliable ML systems while enhancing core performance metrics across model execution. You'll work with advanced performance techniques such as GPU/CUDA optimizations...Remote job
$170k - $216k
...from a diverse set of sensors, enabling engineers like you to (1) develop methods for efficiently... ...scale real-world data, to (2) develop models and model training at scale, to (3)... ...Experience with Python ~ Experience with ML frameworks like PyTorch or JAX We prefer...Full timeRemote work- ...documents codebases autonomously. As a Founding ML Engineer , you’ll architect the intelligence... ...workflows. You’ll work across product, infra, and full-stack teams to embed real-time... ...Develop and refine agent behavior models for planning, implementation, and review...Remote workFlexible hours
$298k - $368k
...of miles of driving data from a diverse set of sensors, enabling engineers like you to (1) develop methods for efficiently and continuously learning from large scale real-world data, to (2) develop models and model training at scale, to (3) analyze real-world behavior and...Full timeRemote work$250k
...without traditional infrastructure limitations. As a Senior ML Infrastructure Engineer, the successful candidate will help build and scale... ...Improve and maintain inference infrastructure, including model packaging, deployment, and serving optimisation Collaborate...Full time- ...We’re looking for a Machine Learning Engineer to build and ship consumer-facing AI systems... ...intelligence.” You’ll work across data, modeling, product, and engineering to translate research... ...you’ll contribute Build and deploy ML models that improve sleep experiences...Full timeImmediate startWorldwideNight shift
$232.65k - $387.75k
As the Sr. Director of AI/ML Engineering you will: Lead a machine learning engineering team specializing in medical imaging and multimodal... ...years proficiency with standard deep learning algorithms and model architectures Experience developing high-performing teams through...- A leading AI evaluation platform in San Francisco is looking for a Senior Software Engineer specializing in ML infrastructure. The successful candidate will design and develop robust real-time data and API systems, enabling insights for researchers and developers. Ideal...
- ...developing next-generation multimodal AI models and a proprietary, high-efficiency serving... ...from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the... ...applications. About the role As an ML Engineer at Sciforium, you will operate...Full timeFlexible hours
- ...are at an inflection point where advances in speech, language models, and clinical AI can fundamentally change how clinicians interact... ...: their patients. Position Overview We are hiring two ML Engineers / Researchers to help build the next generation of Knowtex's...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Infra Engineer, Modeling. Be the first to apply!
- ai ml engineer San Francisco, CA
- graduate machine learning engineer San Francisco, CA
- junior machine learning research engineer San Francisco, CA
- senior ml engineer San Francisco, CA
- machine learning ai engineer San Francisco, CA
- computer vision machine learning engineer San Francisco, CA
- data scientist machine learning engineer San Francisco, CA
- machine learning engineer San Francisco, CA
- machine learning software engineer San Francisco, CA
- artificial intelligence - machine learning intern San Francisco, CA


