Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Systems Engineer

$300k - $400k

Periodic Labs

About Periodic Labs The most important scientific discoveries of our time won't happen in a traditional lab. We're an AI and physical sciences company building state-of-the-art models to accelerate breakthroughs across materials, energy, and beyond. Backed by world-class investors and growing rapidly, we operate at the pace the frontier requires. Our team brings deep expertise, genuine ownership, and an insatiable drive to push the boundaries of what's scientifically possible. About the Role You will own the systems layer that makes our frontier model training and inference fast, efficient, and tightly coupled to the RL feedback loop that drives scientific discovery. This is not a pure infrastructure role and it is not a pure research role — it sits exactly at their intersection. You will go deep into the stack: scheduling, kernels, RDMA, weight synchronization, and communication primitives, while working shoulder-to-shoulder with researchers to co-design the algorithms and infrastructure together. The RL loop is central to how Periodic Labs works. Models propose experiments, experiments generate data, data feeds back into training. The speed and reliability of that loop is a direct multiplier on the pace of scientific discovery. You will own the infrastructure that makes it fast. What You'll Do Build rack and topology-aware scheduling for GB series GPUs across Ray, Slurm, and Kubernetes, minimizing latency and maximizing utilization across heterogeneous cluster configurations Build online and offline profilers that surface bottlenecks across the training and inference stack and translate findings into actionable optimizations Implement direct S3 checkpoint streaming to eliminate I/O bottlenecks in large-scale training runs Run methodical benchmarking to identify optimal RL training configurations across model sizes, batch strategies, and hardware topologies Write and optimize communication and GPU kernels to extract maximum throughput from the hardware Design and implement zero-copy RDMA weight synchronization between training and inference to keep the RL loop tight and low-latency Build fast sandbox execution environments that allow rapid rollout of model-generated actions and return of rewards without blocking the training pipeline Engage directly with the SGLang, Megatron, and Ray communities — contributing upstream, influencing roadmaps, and pulling in improvements that benefit Periodic Labs’ workloads Work in close collaboration with RL and pretraining researchers to co-design algorithms and infrastructure together — you will shape what is possible at the research level by knowing what is achievable at the systems level, and vice versa The net result: high-throughput, fault-tolerant training and inference systems tightly coupled with a low-latency RL feedback loop that accelerates scientific discovery at every turn. You Might Thrive in This Role if You Have Experience With Large-scale inference infrastructure: load balancing, traffic shifting, scheduling, and serving architecture at production scale Low-level systems programming: RDMA, NVLink, kernel-level work, and network stack optimization GPU cluster scheduling and orchestration across Ray, Slurm, or Kubernetes, with awareness of rack topology and hardware locality Writing and optimizing CUDA kernels, communication primitives, or distributed training collective operations Profiling and benchmarking distributed ML systems to identify and eliminate bottlenecks across compute, memory, and network Checkpoint management and streaming at scale, including direct cloud storage integration Building or contributing to open source ML infrastructure projects (e.g., SGLang, Megatron-LM, vLLM, Ray) Working directly with ML researchers on algorithm-infrastructure co-design — you understand the research well enough to make systems decisions that serve it Mechanics Minimum education: Bachelor’s degree or an equivalent combination of education and training or experience Location: Our lab is located in Menlo Park and we prefer folks to be located in Menlo Park or San Francisco but can be flexible based on role Compensation: The annual compensation range for this role - $300,00-$400,000 Visa sponsorship: Yes, we sponsor visas and will do everything we can to assist in this process with our legal support. We’re building a team of the world’s best — the scientists, engineers, and problem-solvers who don’t just follow the frontier, they define it. If you’re driven to bring AI to life in the physical world and make discoveries that have never been made before, you belong here. #J-18808-Ljbffr Periodic Labs

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the ML Systems Engineer in Menlo Park, CA vacancy
  • $250k - $350k

     ...the boundaries of what's scientifically possible. About the Role You'll work alongside some of the world's leading ML systems engineers, including leaders behind Megatron-LM, SGLang, Liger Kernel, TorchRec, CleanRL, TorchRL, and JAX-MD . We're looking for... 
    Suggested
    Visa sponsorship

    Periodic Labs

    Menlo Park, CA
    2 days ago
  •  ...power frontier robotics and world model teams. 3+ yrs distributed systems / ML infra. About Orbifold AI Orbifold AI is building the...  ...are building. Role Overview We are hiring a Machine Learning Engineer to scale and optimize the ML infrastructure behind our pipelines... 
    Suggested

    Orbifold AI

    Palo Alto, CA
    4 days ago
  •  ...work sits at the intersection of distributed systems, GPU performance, model training frameworks, RL pipelines, and production engineering. Your responsibilities Build and maintain...  ...with distributed model training, large-scale ML systems, or GPU cluster workloads.... 
    Suggested

    Nebius B.V.

    Palo Alto, CA
    4 days ago
  •  ...requiring tight integration of hardware and software performance engineering. We seek a senior engineer who will shape core infrastructure...  ...the team as they implement high-throughput inference systems. You will own the inference engine, optimize runtimes, and push... 
    Suggested

    Sanas

    Palo Alto, CA
    5 days ago
  •  ...and production-grade workflows. You will work at the intersection of distributed systems, GPU performance, and ML framework integration. The role requires strong Python and PyTorch engineering skills, hands-on experience with distributed model training, and the ability to... 
    Suggested

    Nebius B.V.

    Palo Alto, CA
    4 days ago
  • $90.1k - $191.8k

     ...and understand the world!The Data Labeling Engineering team designs, builds, and operates high‑...  ...engineering, data engineering, and ML, defining labeling strategies, tooling, and...  ...technical leadership, and work directly on systems that unblock the next generation of AV models... 
    Full time
    Work experience placement
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    2 days ago
  • $174.9k - $261.3k

     ...and understand the world!The Data Labeling Engineering team designs, builds, and operates hybrid...  ...engineering, data engineering, and AI/ML, defining the strategies, tooling, and quality...  ...leadership, and direct impact on systems that unblock the next generation of AV capabilities... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    2 days ago
  • $200k - $300k

     ...Machine Learning Systems EngineerLocation - Palo Alto, CA (On-site) - Five days per week...  ...highly technical, and research-driven, with engineers working directly alongside world-class...  ...operate infrastructure supporting large-scale ML training and inference systemsDevelop... 
    H1b
    Work at office

    Recruiting from Scratch

    Palo Alto, CA
    1 day ago
  •  ...ML Systems Engineer Role OverviewWe’re seeking an experienced engineer to build our ML data infrastructure platform. You’ll create the systems and tools that enable efficient data preparation, feature engineering, and dataset management for machine learning. This role... 

    My3Tech Inc

    Sunnyvale, CA
    3 days ago
  • Rhoda AI is hiring a Senior/Staff-level Research Engineer to ensure our robot-learning pipeline is reliable from data collection through...  ..., and real-robot evaluation. You will build validation systems, observability, and robust operating practices to distinguish model... 

    Socket.dev

    Mountain View, CA
    3 days ago
  • HP, Inc. in Palo Alto is seeking a Senior Software Engineer to build software platforms spanning client apps, backend services, cloud integrations, device interfaces, and enterprise-ready workflows. This role participates in a start-up-like incubation within HP's PC Design... 

    HP Inc.

    Palo Alto, CA
    1 day ago
  • JPMorgan Chase & Co. is seeking a Senior Machine Learning Engineer-Digital Intelligence in the Digital Intelligence team. You will specialize...  ..., interpretability, and related algorithms, driving end-to-end ML solutions in a fast-paced environment. Ideal candidates bring a... 

    JPMorgan Chase & Co.

    Palo Alto, CA
    1 day ago
  • General Motors’ Data Labeling Engineering team is building cutting‑edge labeling tools and pipelines that power autonomous vehicle ML models. The role sits at the intersection of software...  ...and ML, focusing on scalable labeling systems and foundations for foundation‑model... 

    General Motors

    Mountain View, CA
    3 days ago
  • Rhoda AI in Mountain View is seeking a Staff / Principal ML Training Systems Engineer to lead the performance of large-scale multimodal training systems. This role involves improving training efficiency and collaborating closely with research teams to accelerate model... 

    Rhoda AI

    Mountain View, CA
    5 days ago
  • A leading technology company is seeking a Principal Software Engineer for the Economy ML team. You will lead data engineering efforts, setting standards for high-scale data systems and pipelines. Collaborate with Product and Data Science teams to prioritize business growth... 

    Jobleads-US

    San Mateo, CA
    3 days ago
  •  ...About the Role We’re looking for an Applied ML Engineer to design, evaluate, and scale recommendation and ranking systems that power how content, ads, and interactive experiences are selected and surfaced in real time. This role focuses on decision-making systems, with... 
    Full time

    Darwin

    Palo Alto, CA
    1 day ago
  • $169k - $338k

     ...Segment: Home OfficePosition Summary...As a Distinguished AI/ML Engineer within Walmart Global Tech's Site Reliability Engineering organization...  ...lead the technical development of next-generation agentic AI systems and intelligent automation solutions that ensure mission-... 
    Full time
    Temporary work
    Part time

    Walmart

    Sunnyvale, CA
    4 days ago
  •  ...leading AI solutions provider in Palo Alto is looking for a Senior ML Engineer to take ownership of the entire machine learning lifecycle....  ...in Python, and significant expertise with production systems using PyTorch. You will work closely with product and data teams... 

    MetAntz

    Palo Alto, CA
    3 days ago
  • $250k - $350k

    About the RoleWe are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this...  ...improvements to model speed and efficiency, ensuring our creative AI systems deliver industry-leading user experiences at scale.You will... 
    Work at office
    3 days per week

    Pika

    Palo Alto, CA
    3 days ago
  • $148.5k - $223.9k

     ...the right place! Agentforce is the future of AI, and you are the future of Salesforce.This role is for a Senior Machine Learning Engineer within the Trust Intelligence Platform team who will architect data-driven strategies for threat detection across the security organization... 
    Full time

    Salesforce

    Palo Alto, CA
    2 days ago
  • Rhoda AI is seeking a Staff / Principal ML Training Systems Engineer in Mountain View to enhance the training systems performance. This role focuses on large-scale multimodal training, driving efficiency, and scalability in compute use across thousands of GPUs. The ideal... 

    Rhoda AI

    Mountain View, CA
    2 days ago
  • $230k - $260k

     ...enterprise marketing. What You’ll Do As a Principal Machine Learning Engineer, you will operate at the company level—defining technical...  ...AI at Typeface. You will lead the design of large-scale ML systems and shared platforms that power all generative capabilities across... 
    Work at office
    Immediate start
    3 days per week

    Typeface

    Palo Alto, CA
    2 days ago
  • $124k - $250k

     ...support of others.A Day in the LifeAs a member of our software engineering infra team, you'll solve technical challenges, including...  ...turn provide the foundation for rapid development of novel new systems that integrate into that ecosystem and improve it.Our infra team... 

    AppLovin

    Palo Alto, CA
    2 days ago
  • $190k - $234k

     ...What You’ll Do As a StaffMachine Learning Engineer/Applied Scientist, you will be responsible for building machine learning models/systems and innovative web applications that deliver...  ...with building and evolving ML Training and Inferencing systems at significant... 
    Work at office
    Local area
    3 days per week

    Typeface

    Palo Alto, CA
    2 days ago
  •  ...Job Title: ML Engineer What You Will Own End‑to‑End ML Lifecycle across real products: data ingestion, feature design, model selection,...  ...deployment, monitoring and iteration. No handoffs. Production‑grade ML systems built with PyTorch or TensorFlow, focusing on latency,... 

    MetAntz

    Palo Alto, CA
    4 days ago
  •  ...The Mission: As a Senior Machine Learning Engineer, you will be responsible for building machine learning models/systems and innovative web applications that deliver the power...  ...field Experience with building and evolving ML Training and Inferencing systems at significant... 
    Local area

    Typeface

    Palo Alto, CA
    4 days ago
  • $120k - $140k

     ...Deep Learning / NLP Engineer Join Avoma and work on some of the most challenging NLP problems in our mission to make every meeting...  ...looking for a Deep Learning / NLP Engineer to improve and build our systems to extract key insights and topics from conversations. For our... 

    Avoma Inc

    Palo Alto, CA
    1 day ago
  •  ...ML Engineer Palo Alto, California, United States About the Job Our client is a rapidly growing Tier 1 VC backed startup based...  ...long-term growth trajectory in the evolving world of intelligent systems. Location New York, NY Work Type Full Time... 
    Full time

    Catalyst Labs, LLC

    Palo Alto, CA
    3 days ago
  • $148.5k - $223.9k

     ...consider applying for a maximum of 3 roles within 12 months to ensure you are not duplicating efforts. Job Category Software Engineering Job Details About Salesforce Salesforce is the #1 AI CRM, where humans with agents drive customer success together. Here,... 

    Salesforce.Com Inc

    Palo Alto, CA
    2 days ago
  • $138.5k - $225.5k

     ...value across the platform. About the role We're hiring a ML Engineer as one of the founding engineers on Intelligence Org. You'll...  ...Responsibilities Design, train, evaluate, and ship ML systems that power governance and security capabilities, starting with... 
    Full time
    Remote work
    Work from home
    Home office
    Visa sponsorship
    Shift work

    Docker

    Palo Alto, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Systems Engineer. Be the first to apply!