Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Machine Learning Engineer, GPU Kernel and Runtime

$213k - $263k
Full-time

Waymo

Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver—The World's Most Experienced Driver™—to improve access to mobility while saving thousands of lives now lost to traffic crashes. The Waymo Driver powers Waymo’s fully autonomous ride-hail service and can also be applied to a range of vehicle platforms and product use cases. The Waymo Driver has provided over ten million rider-only trips, enabled by its experience autonomously driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states.

The Waymo ML Infrastructure team accelerates Waymo’s mission, by building the best ecosystem for sustainably innovating and shipping ML powered intelligence.

Research, Production, and the Hardware teams are our primary stakeholders and our work powers the development of the state of the art models in the areas of Perception and Trajectory planning that are core to our autonomous driving software. We enable our partners by offering the best in class solutions for the entire model development lifecycle. These solutions include understanding the model business goals and platform hardware characteristics, and codesign the models for the hardwares. These solutions are developed in close collaboration with teams at different modeling teams. Scale and efficiency are core tenets our infra follows.

We are looking for engineers with ML software & systems expertise to help build the next generation Waymo onboard ML inference engine for Waymo fundamental model. You'll work across the entire ML stack from the system perspective, from efficient deep learning models, model compression, ML software (e.g. JAX, XLA, Triton, and CUDA), to . You will be pleasantly challenged with deploying Waymo ML models on limited computation resources. In this hybrid role, you will report to the Senior Manager of Runtime and Optimization.

You will:

  • Collaborate with ML practitioners on models for perception, behavior prediction, and planning, to understand their models and accelerate them onboard through custom NVIDIA GPU kernel development.

  • Deep dive into the NVIDIA ML software and runtime stack, from custom CUDA ops to the XLA:GPU compiler and low-level libraries. Analyze numeric behaviors, debug complex compiler issues, and ensure inference results are stable and consistent. Develop tools/system software for optimal resource usage, hardware efficiency, and platform reliability in an ML serving system.

  • Analyze ML workload performance at the hardware level; apply manual and AI agent-assisted techniques and develop highly optimized, custom CUDA/Triton operator libraries tailored to Waymo's specific architectures.

  • Build tools to benchmark, profile GPU execution, and productize deep learning models for a streamlined and robust onboard and offboard deployment.

You Have:

  • B.S. or M.S. in CS, EE, Deep Learning or a related field

  • 5+ years of industry experience on system performance, hardware-level GPU optimization, or ML compilers

  • Strong C++ and CUDA programming skills

  • Extensive experience in NVIDIA GPU Kernel development to accelerate deep learning models

  • Proven debugging and optimization experience on the XLA:GPU compiler, as well as the NVIDIA runtime stack

  • Passion for developing and optimizing ML software stacks for modern ML accelerator architectures (framework, runtime library, ML compiler, efficient deep learning etc.)

We Prefer:

  • M.Sc or PhD in Computer Science, Mathematics or a related field.

  • Strong Python programming skills.

  • Experience with advanced NVIDIA profiling (e.g., Nsight Compute) and debugging (e.g. cuda-gdb) tools.

  • Solid experience with designing, training and debugging deep learning models to achieve the highest scores/accuracies.

  • In-depth knowledge of ML frameworks, ML compilers, and IRs (Triton, HLO, MLIR, CuTe DSL, cuTile) or modern ML system architectures.

  • Role-Related Knowledge

    • C++ Coding

    • CUDA profiling & debugging

    • ML runtime optimization

    • Custom GPU kernel development

The expected base salary range for this full-time position across US locations is listed below. Actual starting pay will be based on job-related factors, including exact work location, experience, relevant training and education, and skill level. Your recruiter can share more about the specific salary range for the role location or, if the role can be performed remote, the specific salary range for your preferred location, during the hiring process.

Waymo employees are also eligible to participate in Waymo’s discretionary annual bonus program, equity incentive plan, and generous Company benefits program, subject to eligibility requirements.

Salary Range

$213,000—$263,000 USD

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Machine Learning Engineer, GPU Kernel and Runtime in Mountain View, CA vacancy
  • $159.05k - $199.3k

     ...intelligence to every moving machine on the planet. Applied...  ...; Seoul; and Tokyo. Learn more at applied.co. We...  ...looking for a software engineer with deep experience in...  ...‑grade embedded runtime environments. You’ll work...  ...with ML accelerators, GPU, CPU, SoC architecture... 
    Suggested
    Full time
    For contractors
    For subcontractor
    Casual work
    Work at office
    Remote work
    Day shift

    Applied Intuition

    Sunnyvale, CA
    2 days ago
  • $119.25k - $150.85k

     ...Solutions team deploys machine learning models from training...  ...IQ, and we’re hiring engineers to help deliver the next...  ...with our sister teams (kernels, compiler, reduced...  ...for compiler, kernel, runtime, and parity bugs. Collaborate...  ..., ML compilers, GPU programming (CUDA, OpenAI... 
    Suggested
    Full time
    Internship
    Local area
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    1 day ago
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance...  ...across kernel execution, compiler decisions, and runtime scheduling.Direct experience with LLM inference frameworks... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $204k - $259k

     ...automate the lifecycle of the machine learning workflow, including feature...  .... We are looking for engineers with ML software or ML systems...  ...to the Senior Manager of Runtime and Optimization. You will...  ...accelerators (e.g., GPU/TPU). Deep knowledge of... 
    Suggested
    Full time
    Remote work

    Waymo

    Mountain View, CA
    2 days ago
  •  ...intelligence to every moving machine on the planet....  ...Seoul; and Tokyo. Learn more at applied.co....  ...for a performance engineer who specializes in...  ...wastes 30% of its GPU‑hours on stalled...  ...preprocessing, augmentation, kernel execution, gradient...  ..., TensorRT, ONNX Runtime, Ray, or similar)... 
    Suggested
    Full time
    For contractors
    For subcontractor
    Casual work
    Work at office
    Remote work
    Day shift

    Applied Intuition

    Sunnyvale, CA
    9 hours ago
  • $150k

     ..., data scientists, and engineers, tackling the most fundamental...  ...computing in deep learning, driving impactful...  ...optimizing performance for the machine learning software...  ...propose solutions to enhance GPU utilization Support...  ...to develop appropriate kernels and systems for new... 
    Full time
    Work experience placement
    Visa sponsorship

    Institute Of Foundation Models

    Sunnyvale, CA
    1 day ago
  • $213k - $263k

     ...automate the lifecycle of the machine learning workflow, including feature...  .... We are looking for engineers with ML software & systems expertise...  ...to the Senior Manager of Runtime and Optimization. You...  ...Experience with custom kernel development (e.g., CUDA/CUDA... 
    Full time
    Remote work

    Waymo

    Mountain View, CA
    2 days ago
  •  ...Senior ML Infrastructure Engineer (PyTorch, Kubernetes, GPU Training) Short Job Description...  ...infrastructure powering large-scale machine learning training workloads. In this...  ...data pipelines, Triton, TensorRT, custom ML kernels, or ML compiler/runtime optimization.... 
    Full time

    Finoit Inc

    Redwood City, CA
    3 days ago
  • Tilde Research is seeking a Kernel Engineer in Palo Alto to design, implement, and optimize high-performance GPU kernels for training and inference workloads. You will collaborate with ML researchers and engineers to push hardware limits, improve throughput, and reduce... 

    Tilde Research

    Palo Alto, CA
    5 days ago
  •  ...Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design and optimize kernels, build new abstractions for LLM serving engines, and contribute to accelerators and runtimes that power large language models and AI workloads... 

    NVIDIA

    Santa Clara, CA
    9 hours ago
  •  ...California is looking for a Member of Technical Staff for Kernel/Compiler/Communication. This critical role requires strong expertise in CUDA and GPU optimization, along with 5+ years of experience in performance engineering. The ideal candidate will design high-performance... 

    RadixArk

    Palo Alto, CA
    3 days ago
  • CoreWeave is seeking a Senior Software Engineer for the Systems Engineering team to own kernel tracing and patching across Kubernetes and container runtimes. You will debug complex failures, trace root causes in the Linux kernel, and upstream fixes where appropriate. This... 

    CoreWeave

    Sunnyvale, CA
    6 days ago
  • $165.2k - $223.6k

     ...used to accelerate deep learning and GenAI workloads on Amazon’s custom machine learning accelerators,...  ...Trainium.The Acceleration Kernel Library team is at the...  ...software boundary, our engineers craft high-performance...  ...an ML compiler, runtime, and application framework... 
    Internship
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $300k

     ..., data scientists, and engineers, tackling the most fundamental...  ...computing in deep learning, driving impactful...  ...across multi-node, multi-GPU clusters  • Own experiment...  ...with large-scale machine learning workloads (strong...  ...performance profiling, kernel fusion, or memory optimization... 
    Full time
    Flexible hours

    Institute Of Foundation Models

    Sunnyvale, CA
    1 day ago
  •  ...patients worldwide.We’re a team of engineers, clinicians, and innovators...  ...This Opportunity:As a Staff Machine Learning Engineer, you will be responsible...  ...for real-time inference, GPU/throughput optimization (e.g., TensorRT, ONNX Runtime, mixed precision), and building... 
    Work at office
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    Sunnyvale, CA
    4 days ago
  • $165.45k - $259.75k

     ...Senior Software Engineer Description - HP is seeking a Senior Software...  ..., evaluate, and deploy machine learning models — including LLMs, vision...  ...TensorFlow / TensorFlow Lite, ONNX Runtime , or Hugging Face...  ...Cloud , containerized inference, GPU/NPU acceleration). Strong Python... 
    Full time
    Temporary work
    Local area
    Relocation
    Flexible hours
    Shift work

    HP

    Palo Alto, CA
    1 day ago
  • $150.4k - $277.6k

     ...Compute Frameworks team in GPU, Graphics and Displays org provides...  ...algebra, image processing, machine learning, along with other projects...  ...and GPU programming engineers who are passionate about providing...  ...optimized GPU compute kernels across Machine Learning, Image... 
    Relocation

    Apple Inc.

    Cupertino, CA
    6 days ago
  • $200k - $300k

     ...Machine Learning Systems EngineerLocation - Palo Alto, CA (On-site) - Five...  ..., and research-driven, with engineers working directly alongside...  ...learning, distributed systems, GPU infrastructure, and high-...  ...such as vLLM, TensorRT, ONNX Runtime, and SGLangDevelop and optimize... 
    H1b
    Work at office

    Recruiting from Scratch

    Palo Alto, CA
    11 hours ago
  • Junior Distributed Machine Learning Engineer Join to apply for the Junior Distributed Machine Learning Engineer role...  ..., network and propose solutions to enhance GPU utilization • Support the team to develop appropriate kernels and systems for new model architectures and... 
    Full time
    Work experience placement
    H1b

    Jobright.ai

    Sunnyvale, CA
    3 days ago
  • Apple Inc. in Cupertino, CA, is seeking extraordinary machine learning and GPU programming engineers to deliver robust compute solutions on Apple Silicon. You will contributes to optimized kernels, ML workflows, and high-performance GPU architectures across iOS, macOS... 

    Apple Inc.

    Cupertino, CA
    6 days ago
  • $2,000 per month

     ...Machine Learning Research EngineerCupertino, CAEtched is building...  ...tools, and runtime abstractions by implementing...  ...models on Sohu to achieve GPU-impossible latencies...  ...with GPU kernels, the CUDA compilation...  ...Cupertino, and greatly value engineering skills. We do not have... 

    ETCHED LLC

    Cupertino, CA
    2 days ago
  •  ...CoreWeave is seeking a Senior Engineer for its Benchmarking & Performance team to own kernel-level optimization for LLM inference and end-to-end model serving, focusing on CUDA kernels and throughput/latency improvements. You will lead kernel design reviews, mentor... 

    Neura Market

    Sunnyvale, CA
    9 hours ago
  • $250k - $350k

     ...Senior/Staff level Inference Engineers to accelerate the performance...  ...edge inference acceleration, GPU parallelism, advanced model deployment...  ...high-performance computing kernels and distributed workloads...  ...acceleration, and deep learning compiler stacks.GPU & Parallelism... 
    Work at office
    3 days per week

    Pika

    Palo Alto, CA
    3 days ago
  • $195.2k - $262.2k

     ...infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to...  ...Partner with GPU kernel engineers and platform engineers...  ...across model code, kernels, runtime, scheduler, gateway, and...  ...Career growth and learning opportunities Flexibility... 
    Full time
    Temporary work
    Immediate start
    Remote work

    Nebius

    Palo Alto, CA
    6 hours ago
  •  ...Technical Staff to push the limits of performance at the kernel, compiler, and communication layers. You will optimize runtimes and libraries to unlock maximum efficiency on modern accelerators across large GPU clusters. You will design high-performance kernels, improve... 

    RadixArk

    Palo Alto, CA
    6 days ago
  • $184k - $287.5k

     ...computing. An era in which our GPU acts as the brains of...  ...for a Senior Software Engineer to join the...  ...ships the production GPU kernels and software interfaces...  ...power equivariant deep learning throughout the scientific...  ...concepts in geometric machine learning - equivariance... 

    Nvidia Corporation

    Santa Clara, CA
    2 days ago
  •  ...production-grade workflows. You will work at the intersection of distributed systems, GPU performance, and ML framework integration. The role requires strong Python and PyTorch engineering skills, hands-on experience with distributed model training, and the ability to... 

    Nebius B.V.

    Palo Alto, CA
    4 days ago
  • $184k - $287.5k

    We are now looking for a Senior Formal Verification Engineer for GPU Kernels! Modern AI performance relies on highly optimized GPU kernels — performance...  ...can be hard to catch and expensive to miss. NVIDIA's Deep Learning Safety Team is hiring engineers to build the verification... 
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    4 days ago
  • $124k - $195.5k

     ...Deep Learning Software Engineer, TensorRT PerformanceNVIDIA is seeking an experienced...  ...specialize in developing GPU-accelerated deep learning...  ...modeling, performance analysis, kernel development and inference...  ...the foundation for machines to learn, perceive, reason... 

    NVIDIA

    Santa Clara, CA
    2 days ago
  • $184.7k - $324.8k

     ...Cupertino, California, United States Machine Learning and AI Video is at the core of nearly all Apple products, and as a research engineer on our team, you will build the infrastructure...  ...custom ops in CUDA or low-level GPU kernel optimization Research experience in ML... 
    Relocation

    Apple

    Cupertino, CA
    11 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Machine Learning Engineer, GPU Kernel and Runtime. Be the first to apply!