Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Systems Engineer

$250k - $350k

Periodic Labs

ML Systems EngineerYou'll work alongside some of the world's leading ML systems engineers, including leaders behind Megatron-LM, SGLang, Liger Kernel, TorchRec, CleanRL, TorchRL, and JAX-MD.We're looking for exceptional ML Systems Engineers to build the agentic infrastructure powering our large-scale training, inference, and reinforcement learning. You'll own critical pieces of the ML systems stack to maximize performance, scalability, reliability, and productivity for both engineers and AI agents.What You'll DoBuild and optimize large-scale training and reinforcement learning infrastructure while ensuring its correctnessDevelop high-performance inference and serving systemsDesign distributed runtimes and scheduling systems for complex ML workloadsBuild secure and large-scale sandboxing and execution environmentsOptimize memory, GPU kernels and communication for maximum throughput and end-to-end efficiencyImprove scalability, reliability, and efficiency across the ML systems stackWhat We're Looking ForStrong systems programming and performance engineering skillsExperience building high-performance ML infrastructure at scaleAbility to own complex technical problems end-to-endStrong coding ability and engineering judgment, including the ability to work effectively with AI agents to design, implement, test, and debug complex systemsHigh ownership, fast execution, and a passion for pushing the frontier of AI systems and accelerating scientific discoveryYou should have deep expertise in at least one of the following:Training: Strong experience building, debugging and optimizing large-scale training systems with Megatron-LM. Familiarity with TorchTitan, FSDP, veRL, Slime, or other distributed training systems is a plus.Distributed Runtime: Strong experience with Ray. Familiarity with Monarch or other distributed execution frameworks is a plus.Inference: Strong experience with SGLang. Familiarity with vLLM, TensorRT-LLM, or production LLM serving systems is a plus.Sandboxing: Strong experience with secure execution environments, containers, virtualization, or code sandboxing.GPU Kernels: Strong experience with CUDA, Triton, CUTLASS, CuTe, or custom GPU kernel development.GPU Communication: Strong experience with NCCL, NVLink, InfiniBand, RDMA, GPUDirect RDMA, or large-scale communication optimization.MechanicsMinimum education: Bachelor's degree or similar experienceLocation: Menlo Park, CA (Soon: San Francisco, too)Compensation: $250,000-$350,000 base + equityVisa sponsorship: Yes, we sponsor visas and will do everything we can to assist in this process with our legal support.We're building a team of the world's best — the scientists, engineers, and problem-solvers who don't just follow the frontier, they define it. If you're driven to bring AI to life in the physical world and make discoveries that have never been made before, you belong here.

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the ML Systems Engineer in Menlo Park, CA vacancy
  •  ...work sits at the intersection of distributed systems, GPU performance, model training frameworks, RL pipelines, and production engineering. Your responsibilities Build and maintain...  ...with distributed model training, large-scale ML systems, or GPU cluster workloads.... 
    Suggested

    Nebius B.V.

    Palo Alto, CA
    1 day ago
  •  ...power frontier robotics and world model teams. 3+ yrs distributed systems / ML infra. About Orbifold AI Orbifold AI is building the...  ...are building. Role Overview We are hiring a Machine Learning Engineer to scale and optimize the ML infrastructure behind our pipelines... 
    Suggested

    Orbifold AI

    Palo Alto, CA
    2 days ago
  •  ...and production-grade workflows. You will work at the intersection of distributed systems, GPU performance, and ML framework integration. The role requires strong Python and PyTorch engineering skills, hands-on experience with distributed model training, and the ability to... 
    Suggested

    Nebius B.V.

    Palo Alto, CA
    2 days ago
  • Recruiting From Scratch is seeking a Machine Learning Systems Engineer to design and operate large-scale ML training and inference infrastructure in Palo Alto. The role focuses on building high-performance, GPU-accelerated systems for model serving and deployment across... 
    Suggested

    Recruiting from Scratch

    Palo Alto, CA
    3 days ago
  • Arch Systems is looking for a talented individual to design and implement optimal algorithms for Wi-Fi network performance, leveraging...  ...at least three years of experience in software or systems engineering. Key skills include Wi-Fi products development, WLAN management... 
    Suggested
    Remote work

    Arch Systems

    Palo Alto, CA
    4 days ago
  • $300k - $400k

     ...possible. About the Role You will own the systems layer that makes our frontier model...  ...Profiling and benchmarking distributed ML systems to identify and eliminate bottlenecks...  ...team of the world’s best — the scientists, engineers, and problem-solvers who don’t just follow... 
    Visa sponsorship
    Flexible hours
    Shift work

    Periodic Labs

    Menlo Park, CA
    2 days ago
  •  ...Member of Technical Staff — Training to design and optimize large-scale distributed training systems for frontier AI models. Candidates should have 5+ years of experience in ML systems and be proficient in Python along with another systems language, such as C++. This role... 

    RadixArk

    Palo Alto, CA
    1 day ago
  • $174.9k - $261.3k

     ...and understand the world!The Data Labeling Engineering team designs, builds, and operates hybrid...  ...engineering, data engineering, and AI/ML, defining the strategies, tooling, and quality...  ...leadership, and direct impact on systems that unblock the next generation of AV capabilities... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    5 days ago
  • $144.7k - $261.3k

     ...challenges for autonomous vehicle development. We engineer high-performance tools that identify top-...  ...models and partner with data-intensive ML teams to drive rapid innovation. Why...  ...of next-generation autonomous systems. About the Role As a Senior Engineer... 
    Full time
    Local area
    Remote work
    Work from home
    Flexible hours

    General Motors

    Sunnyvale, CA
    3 days ago
  • $90.1k - $191.8k

     ...and understand the world!The Data Labeling Engineering team designs, builds, and operates high‑...  ...engineering, data engineering, and ML, defining labeling strategies, tooling, and...  ...technical leadership, and work directly on systems that unblock the next generation of AV models... 
    Full time
    Work experience placement
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    22 hours ago
  •  ...own the execution layer of our intelligence platform in Palo Alto. You will translate research direction into reliable, scalable ML systems deployed in production, collaborating with infra, product, and applications teams. Responsibilities include end-to-end ML pipelines... 

    A1

    Palo Alto, CA
    5 days ago
  • Rhoda AI is building next-generation generalist robots and is seeking a Staff / Principal ML Training Systems Engineer to own training systems performance end-to-end. You will optimize large-scale multimodal training, define parallelism strategies, and drive efficiency... 

    Rhoda AI

    Palo Alto, CA
    5 days ago
  • $204k - $259k

     ...autonomous driving technology company is looking for an experienced engineer to improve compute performance in machine learning systems. This hybrid role involves collaboration with a world-class ML team and requires strong expertise in ML software or systems. The ideal... 

    Waymo

    Mountain View, CA
    4 days ago
  • Rhoda is building the next generation of generalist robotic systems in Mountain View, CA. We are seeking a senior or staff-level Research Engineer or ML Systems Engineer to make the robot-learning pipeline reliable and measurable from end to end. You will own the supported... 

    RHODA

    Mountain View, CA
    2 days ago
  • JPMorgan Chase & Co. is seeking a Senior Machine Learning Engineer-Digital Intelligence in the Digital Intelligence team. You will specialize...  ..., interpretability, and related algorithms, driving end-to-end ML solutions in a fast-paced environment. Ideal candidates bring a... 

    JPMorgan Chase & Co.

    Palo Alto, CA
    4 days ago
  • $200k - $300k

    Machine Learning Systems Engineer Location - Palo Alto, CA (On-site) - Five days per week in-office in the Bay Area. Compensation - $200,000...  ...Design, build, and operate infrastructure supporting large-scale ML training and inference systems Develop high-performance... 
    H1b
    Work at office

    Recruiting from Scratch

    Palo Alto, CA
    3 days ago
  • SpaceXAI is seeking exceptional Applied engineers to join a high-priority project used by hundreds of millions of users monthly. You will...  ...and real-world impact, applying your skills to recommendation systems, ranking algorithms, search technologies, and more. You will design... 

    Pantera Capital

    Palo Alto, CA
    3 days ago
  •  ...deployed on-premise inside sovereign data centers, delivering low-latency, high-throughput AI for multi-node GPU workloads. As a Senior Engineer, you will shape core infrastructure and architecture decisions, lead performance optimizations, and own the inference engine to... 

    Sanas

    Palo Alto, CA
    2 days ago
  • $227.87k

     ...developing and executing solutions that enhance the Pinterest marketplace, primarily focusing on backend systems and statistical models. Candidates need strong software engineering skills and a graduate degree in a related field. The position is unique, allowing collaboration... 
    Work at office
    Flexible hours

    Pinterest

    Palo Alto, CA
    5 days ago
  • General Motors’ Data Labeling Engineering team is building cutting‑edge labeling tools and pipelines that power autonomous vehicle ML models. The role sits at the intersection of software...  ...and ML, focusing on scalable labeling systems and foundations for foundation‑model... 

    General Motors

    Mountain View, CA
    1 day ago
  • Rhoda AI is hiring a Senior/Staff-level Research Engineer to ensure our robot-learning pipeline is reliable from data collection through...  ..., and real-robot evaluation. You will build validation systems, observability, and robust operating practices to distinguish model... 

    Socket.dev

    Mountain View, CA
    1 day ago
  • Rhoda AI in Mountain View is seeking a Staff / Principal ML Training Systems Engineer to lead the performance of large-scale multimodal training systems. This role involves improving training efficiency and collaborating closely with research teams to accelerate model... 

    Rhoda AI

    Mountain View, CA
    3 days ago
  • A leading technology company is seeking a Principal Software Engineer for the Economy ML team. You will lead data engineering efforts, setting standards for high-scale data systems and pipelines. Collaborate with Product and Data Science teams to prioritize business growth... 

    Jobleads-US

    San Mateo, CA
    1 day ago
  • $169k - $338k

     ...Segment: Home OfficePosition Summary...As a Distinguished AI/ML Engineer within Walmart Global Tech's Site Reliability Engineering organization...  ...lead the technical development of next-generation agentic AI systems and intelligent automation solutions that ensure mission-... 
    Full time
    Temporary work
    Part time

    Walmart

    Sunnyvale, CA
    2 days ago
  •  ...leading AI solutions provider in Palo Alto is looking for a Senior ML Engineer to take ownership of the entire machine learning lifecycle....  ...in Python, and significant expertise with production systems using PyTorch. You will work closely with product and data teams... 

    MetAntz

    Palo Alto, CA
    1 day ago
  • $250k - $350k

    About the RoleWe are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this...  ...improvements to model speed and efficiency, ensuring our creative AI systems deliver industry-leading user experiences at scale.You will... 
    Work at office
    3 days per week

    Pika

    Palo Alto, CA
    1 day ago
  • $148.5k - $223.9k

     ...the right place! Agentforce is the future of AI, and you are the future of Salesforce.This role is for a Senior Machine Learning Engineer within the Trust Intelligence Platform team who will architect data-driven strategies for threat detection across the security organization... 
    Full time

    Salesforce

    Palo Alto, CA
    5 days ago
  • $189.72k - $332.01k

    A leading social media platform based in Palo Alto is seeking a Machine Learning Engineer. The role involves building innovative systems using deep learning and machine learning, improving their models across various product areas, and utilizing data-driven methods for... 

    Pinterest

    Palo Alto, CA
    1 day ago
  • Rhoda AI is seeking a Staff / Principal ML Training Systems Engineer in Mountain View to enhance the training systems performance. This role focuses on large-scale multimodal training, driving efficiency, and scalability in compute use across thousands of GPUs. The ideal... 

    Rhoda AI

    Mountain View, CA
    5 days ago
  •  ...We bridge this exact gap by applying deep systems programming, software-defined networking,...  ...at UT Austin and world-renowned ML systems researcher with a pedigree spanning...  ...Seniority ~5+ years of production experience engineering ML systems, OR a PhD from a top-tier... 
    Shift work

    Success Matcher Recruitment

    Sunnyvale, CA
    23 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Systems Engineer. Be the first to apply!