Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Systems Engineer

$300k - $400k

Periodic Labs

About Periodic Labs

The most important scientific discoveries of our time won't happen in a traditional lab. We're an AI and physical sciences company building state-of-the-art models to accelerate breakthroughs across materials, energy, and beyond. Backed by world-class investors and growing rapidly, we operate at the pace the frontier requires. Our team brings deep expertise, genuine ownership, and an insatiable drive to push the boundaries of what's scientifically possible.

About the Role

You will own the systems layer that makes our frontier model training and inference fast, efficient, and tightly coupled to the RL feedback loop that drives scientific discovery.

This is not a pure infrastructure role and it is not a pure research role — it sits exactly at their intersection. You will go deep into the stack: scheduling, kernels, RDMA, weight synchronization, and communication primitives, while working shoulder-to-shoulder with researchers to co-design the algorithms and infrastructure together.

The RL loop is central to how Periodic Labs works. Models propose experiments, experiments generate data, data feeds back into training. The speed and reliability of that loop is a direct multiplier on the pace of scientific discovery. You will own the infrastructure that makes it fast.

What You'll Do
  • Build rack and topology-aware scheduling for GB series GPUs across Ray, Slurm, and Kubernetes, minimizing latency and maximizing utilization across heterogeneous cluster configurations

  • Build online and offline profilers that surface bottlenecks across the training and inference stack and translate findings into actionable optimizations

  • Implement direct S3 checkpoint streaming to eliminate I/O bottlenecks in large-scale training runs

  • Run methodical benchmarking to identify optimal RL training configurations across model sizes, batch strategies, and hardware topologies

  • Write and optimize communication and GPU kernels to extract maximum throughput from the hardware

  • Design and implement zero-copy RDMA weight synchronization between training and inference to keep the RL loop tight and low-latency

  • Build fast sandbox execution environments that allow rapid rollout of model-generated actions and return of rewards without blocking the training pipeline

  • Engage directly with the SGLang, Megatron, and Ray communities — contributing upstream, influencing roadmaps, and pulling in improvements that benefit Periodic Labs’ workloads

  • Work in close collaboration with RL and pretraining researchers to co-design algorithms and infrastructure together — you will shape what is possible at the research level by knowing what is achievable at the systems level, and vice versa

The net result: high-throughput, fault-tolerant training and inference systems tightly coupled with a low-latency RL feedback loop that accelerates scientific discovery at every turn.

You Might Thrive in This Role if You Have Experience With
  • Large-scale inference infrastructure: load balancing, traffic shifting, scheduling, and serving architecture at production scale

  • Low-level systems programming: RDMA, NVLink, kernel-level work, and network stack optimization

  • GPU cluster scheduling and orchestration across Ray, Slurm, or Kubernetes, with awareness of rack topology and hardware locality

  • Writing and optimizing CUDA kernels, communication primitives, or distributed training collective operations

  • Profiling and benchmarking distributed ML systems to identify and eliminate bottlenecks across compute, memory, and network

  • Checkpoint management and streaming at scale, including direct cloud storage integration

  • Building or contributing to open source ML infrastructure projects (e.g., SGLang, Megatron-LM, vLLM, Ray)

  • Working directly with ML researchers on algorithm-infrastructure co-design — you understand the research well enough to make systems decisions that serve it

Mechanics

Minimum education: Bachelor’s degree or an equivalent combination of education and training or experience

Location: Our lab is located in Menlo Park and we prefer folks to be located in Menlo Park or San Francisco but can be flexible based on role

Compensation: The annual compensation range for this role - $300,00-$400,000

Visa sponsorship: Yes, we sponsor visas and will do everything we can to assist in this process with our legal support.

We’re building a team of the world’s best — the scientists, engineers, and problem-solvers who don’t just follow the frontier, they define it. If you’re driven to bring AI to life in the physical world and make discoveries that have never been made before, you belong here.

#J-18808-Ljbffr
Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the ML Systems Engineer in Menlo Park, CA vacancy
  •  ...We bridge this exact gap by applying deep systems programming, software-defined networking,...  ...at UT Austin and world-renowned ML systems researcher with a pedigree spanning...  ...Seniority ~5+ years of production experience engineering ML systems, OR a PhD from a top-tier... 
    Suggested
    Shift work

    Success Matcher Recruitment

    Sunnyvale, CA
    a month ago
  •  ...behaviors, from industrial machinery and systems to wearable devices and smart environments...  ...in our agent runtime, built on Rust ML stacks (candle, Burn) with custom GPU kernels...  ...Key Qualifications ~6+ years software engineering, several of them in ML systems, inference... 
    Suggested

    Archetype AI Inc.

    San Mateo, CA
    18 hours ago
  • $179k - $223k

     ...is a foundational pillar of the Autonomy stack. In this Senior ML Engineer role, you will play a key role in delivering high-quality,...  ...evaluation and monitoring benchmarks. Identify and root-cause top-tier system anomalies, prioritizing high-impact optimizations to... 
    Suggested
    Full time
    Contract work
    Local area

    Rivian

    Palo Alto, CA
    1 day ago
  • $500 per month

     ...on Astrocade starts as an idea from a creator. An agentic AI system turns that idea into something playable. How good that experience...  ...north-star metric. Your job is to move that metric. As an ML Engineer on our AI team, you'll own how our creation system behaves,... 
    Suggested
    Work at office
    Flexible hours

    Astrocade Inc

    Palo Alto, CA
    1 hour ago
  • $208k - $244k

     ...Join to apply for the Staff ML Engineer role at Grindr Join to apply for the Staff ML Engineer role at Grindr Get AI-powered advice on...  ...our long term ML strategy. Recommendations That Reshape: Build systems that match millions to their next big moment, adapting to a range... 
    Suggested
    Full time
    Casual work
    Work at office
    Immediate start
    Flexible hours

    Grindr

    Palo Alto, CA
    1 day ago
  • $230k - $260k

     ...enterprise marketing.What You’ll Do As a Principal Machine Learning Engineer, you willoperateat the company level—defining technical...  ...generative AI at Typeface.You will lead the design of large-scale ML systems and shared platforms that power all generative capabilities... 
    Work at office
    Immediate start
    3 days per week

    Typeface

    Palo Alto, CA
    4 days ago
  • $180k

     ...SpaceXAI’s mission is to create AI systems that can accurately understand the universe and...  ...small, highly motivated, and focused on engineering excellence. This organization is for individuals...  ...teammates. ABOUT THE ROLE: As an ML Infrastructure Engineer, you will play a... 
    Temporary work
    Work experience placement

    Socket

    Palo Alto, CA
    1 day ago
  • $228k - $285k

     ...foundational pillar of the Autonomy safety architecture. In this Staff ML Software Engineer role, you will play a key part in driving and delivering...  ..., you will own the end-to-end lifecycle of the failsafe system: defining safety metrics, engineering diagnostic algorithms,... 
    Odd job
    Full time
    Contract work
    Local area

    Rivian

    Palo Alto, CA
    3 days ago
  • $228k - $285k

     ...future generations. Role Summary As a Staff Software Engineer for ML Optimization and Hardware Acceleration, you will be a lead...  ...employment, social media/website, network/device, recruiting system usage/interaction, security and preference information.... 
    Full time
    Contract work
    Local area

    Rivian

    Palo Alto, CA
    2 days ago
  • $190k - $234k

     ....What You’ll Do As aStaffMachine Learning Engineer/Applied Scientist, you will be responsible for building machine learning models/systems and innovative web applications that deliver...  ...with building and evolving ML Training and Inferencing systems at significant... 
    Work at office
    Local area
    3 days per week

    Typeface

    Palo Alto, CA
    4 days ago
  • $300k

     ...term relationship. We're doubling down on ML as the future of Grindr, and in these...  ...for our long term ML strategy. Build systems that match millions to their next big moment...  .... Collaborate cross-functionally with engineering, data science and product teams to turn bold... 
    Casual work
    Work at office
    Immediate start
    Worldwide
    Flexible hours

    Grindr

    Palo Alto, CA
    1 day ago
  • $174.9k - $261.3k

     ...and understand the world!The Data Labeling Engineering team designs, builds, and operates hybrid...  ...engineering, data engineering, and AI/ML, defining the strategies, tooling, and quality...  ...leadership, and direct impact on systems that unblock the next generation of AV capabilities... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    1 day ago
  • $90.1k - $191.8k

     ...and understand the world!The Data Labeling Engineering team designs, builds, and operates high‑...  ...engineering, data engineering, and ML, defining labeling strategies, tooling, and...  ...technical leadership, and work directly on systems that unblock the next generation of AV models... 
    Full time
    Work experience placement
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    1 day ago
  • $150k - $224k

     ...practices, tooling for data infrastructure used across ML teams Required Qualifications Strong software engineering fundamentals, with experience building high-throughput, fault-tolerant distributed systems Hands-on experience with distributed computing... 
    Full time

    AppLovin

    Palo Alto, CA
    4 days ago
  • $220k - $350k

     ...Job Description Job Description ML Infrastructure Engineer Company: Dyna Robotics Location: Redwood City, CA (in office 5 days per week...  ...Early-stage or founding infrastructure hire Multimodal systems (video, audio, multimedia models) DeepSpeed or... 
    Full time
    H1b
    Work at office
    Visa sponsorship

    Transparent Search Group

    Redwood City, CA
    5 days ago
  •  ...Data Systems Engineer - ELK/Kafka/Linux Alpharetta, GA or Menlo Park, CA - Hybrid 3 Days a Week Onsite 12 months+ Interview Process: # Screening of Technical Background - check Linux exp (Team member - 1hr) # Technical Panel (in person) Bachelor's Degree... 
    3 days per week

    Veterans Sourcing Group LLC

    Menlo Park, CA
    18 hours ago
  • $200k - $250k

     ...miles with ones on vehicles that are more affordable, more enjoyable and 10-50x more efficient. ALSO is looking for a Systems Engineering Manager to lead requirements and feature development across our electric mobility vehicle platforms. What You'll Do... 
    Local area
    Flexible hours

    ALSO

    Palo Alto, CA
    8 days ago
  •  ...About the Role We are seeking a Senior Data / AI / ML Software Engineer with 7+ years of experience building data-intensive systems. This role is ideal for someone who enjoys designing and improving core platform components at the intersection of software engineering,... 
    Full time
    Contract work
    Internship

    Next Ventures

    Palo Alto, CA
    18 hours ago
  • $150k - $250k

     ...Array Labs builds advanced radar systems to help humanity understand and respond to changes across the physical world. We’re launching...  ...About the Job We’re seeking an experienced Space Systems Engineer to join our team as we move from prototypes into proliferated systems... 
    Permanent employment
    Full time
    Remote work

    ArrayLabs, LLC

    Redwood City, CA
    1 day ago
  • $54.21 - $70.47 per hour

     ...LP_00023865-2660117 Job Description JOB SUMMARY This paragraph summarizes the general nature, level and purpose of the job. The System Engineer acts as a liaison between PCHA clinics and LPCH Clinical Technology. This includes responsibility to the entire program cycle... 
    Hourly pay
    Work experience placement

    Lucile Packard Children's Hospital Stanford

    Palo Alto, CA
    1 day ago
  • $61k - $101k

     ...graph-based formats. We need solid software engineering capabilities and the ability to build robust, production-ready systems that operate reliably at scale. We require...  ...qualifications include a publication record at top AI/ML venues. Preferred qualifications include... 
    Full time
    Work experience placement

    J.P. Morgan

    Palo Alto, CA
    21 days ago
  • $61k - $101k

     ...graph-based formats. We need strong software engineering skills and the ability to build robust, production-grade systems that operate reliably at scale. We require...  ...systems. Preferred: publication record in top AI/ML venues. Preferred: experience optimizing... 
    Full time
    Work experience placement

    J.P. Morgan

    Palo Alto, CA
    14 days ago
  • $131k - $271.6k

     ...paced and agile innovation team, consisting of machine learning engineers and software engineers. As a Senior Machine Learning Engineer...  ...Dev-Ops and AI engineers. · Contribute to all processes of the ML lifecycle: data collection, annotation, modeling, evaluation,... 
    Permanent employment
    Full time
    Worldwide
    Flexible hours

    SAP

    Palo Alto, CA
    5 days ago
  •  ...develop, and manufacture light eVTOL aircraft for personal, public service, and defense applications. Pivotal is seeking a Systems Engineer – GNC & Flight Controls to support requirements development, system integration, and Verification & Validation (V&V) of our... 
    Work at office

    Pivotal

    Palo Alto, CA
    7 days ago
  •  ...About the Role Our next Senior Machine Learning Engineer will work on architecting the next phase of our core ML platform, the developer experience of creating...  ...features. Platform Excellence Monitor production systems built on our Bun/TypeScript backend and Postgres,... 

    RiseMe

    Palo Alto, CA
    18 hours ago
  •  ...time Location Type On-site Department Engineering Role Overview As a Machine Learning...  ...functional teams to identify opportunities where ML can drive product value, architect robust model-centric systems, and ensure their seamless integration into real... 
    Full time

    NACE

    Palo Alto, CA
    18 hours ago
  • $112.7k - $169.1k

     ...Vector AI team builds the machine learning systems that decide which ads reach which players...  ...monthly users on the world's leading game engine. Recommendation and ranking systems are the...  ...working with large‑scale data and ML systems, whether through research or industry... 
    Internship
    Work at office
    Worldwide
    Relocation package
    Shift work

    Unity South APAC (SEA, ANZ, IND Subcont.)

    Mountain View, CA
    18 hours ago
  • $173k - $259k

     ...Saturn, and other digital services. Snap Engineering teams build fun and technically...  ...millions of Snapchatters Apply modern ML techniques to solve large‑scale, real‑world...  ...Adaptability in learning and applying evolving AI systems and tools to remain at the forefront of... 
    Work experience placement
    Live in
    Local area

    Snap

    Palo Alto, CA
    18 hours ago
  • $190k - $202k

     ...Wing is looking for a Machine Learning Engineer to join our Perception team . This role...  ...advancing our autonomous drone delivery system, leveraging deep learning to ensure our aircraft...  ...: Build and deploy highly reliable ML solutions both for the backend as well as... 
    Full time
    Local area

    Socket

    Palo Alto, CA
    18 hours ago
  •  ...Conduct functional and performance testing of prototypes and document results. Create and manage BOMs (Bill of Materials) and engineering change requests in SAP. Ensure compliance with design standards, safety regulations, and quality requirements. Support continuous... 

    Omni Inclusive

    Palo Alto, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Systems Engineer. Be the first to apply!