Machine Learning Engineer, GPU Kernel and Runtime
$213k - $263kWaymo
Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver—The World's Most Experienced Driver™—to improve access to mobility while saving thousands of lives now lost to traffic crashes. The Waymo Driver powers Waymo’s fully autonomous ride-hail service and can also be applied to a range of vehicle platforms and product use cases. The Waymo Driver has provided over ten million rider-only trips, enabled by its experience autonomously driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states.
The Waymo ML Infrastructure team accelerates Waymo’s mission, by building the best ecosystem for sustainably innovating and shipping ML powered intelligence.
Research, Production, and the Hardware teams are our primary stakeholders and our work powers the development of the state of the art models in the areas of Perception and Trajectory planning that are core to our autonomous driving software. We enable our partners by offering the best in class solutions for the entire model development lifecycle. These solutions include understanding the model business goals and platform hardware characteristics, and codesign the models for the hardwares. These solutions are developed in close collaboration with teams at different modeling teams. Scale and efficiency are core tenets our infra follows.
We are looking for engineers with ML software & systems expertise to help build the next generation Waymo onboard ML inference engine for Waymo fundamental model. You'll work across the entire ML stack from the system perspective, from efficient deep learning models, model compression, ML software (e.g. JAX, XLA, Triton, and CUDA), to . You will be pleasantly challenged with deploying Waymo ML models on limited computation resources. In this hybrid role, you will report to the Senior Manager of Runtime and Optimization.
You will:
-
Collaborate with ML practitioners on models for perception, behavior prediction, and planning, to understand their models and accelerate them onboard through custom NVIDIA GPU kernel development.
-
Deep dive into the NVIDIA ML software and runtime stack, from custom CUDA ops to the XLA:GPU compiler and low-level libraries. Analyze numeric behaviors, debug complex compiler issues, and ensure inference results are stable and consistent. Develop tools/system software for optimal resource usage, hardware efficiency, and platform reliability in an ML serving system.
-
Analyze ML workload performance at the hardware level; apply manual and AI agent-assisted techniques and develop highly optimized, custom CUDA/Triton operator libraries tailored to Waymo's specific architectures.
-
Build tools to benchmark, profile GPU execution, and productize deep learning models for a streamlined and robust onboard and offboard deployment.
You Have:
-
B.S. or M.S. in CS, EE, Deep Learning or a related field
-
5+ years of industry experience on system performance, hardware-level GPU optimization, or ML compilers
-
Strong C++ and CUDA programming skills
-
Extensive experience in NVIDIA GPU Kernel development to accelerate deep learning models
-
Proven debugging and optimization experience on the XLA:GPU compiler, as well as the NVIDIA runtime stack
-
Passion for developing and optimizing ML software stacks for modern ML accelerator architectures (framework, runtime library, ML compiler, efficient deep learning etc.)
We Prefer:
-
M.Sc or PhD in Computer Science, Mathematics or a related field.
-
Strong Python programming skills.
-
Experience with advanced NVIDIA profiling (e.g., Nsight Compute) and debugging (e.g. cuda-gdb) tools.
-
Solid experience with designing, training and debugging deep learning models to achieve the highest scores/accuracies.
-
In-depth knowledge of ML frameworks, ML compilers, and IRs (Triton, HLO, MLIR, CuTe DSL, cuTile) or modern ML system architectures.
-
Role-Related Knowledge
-
-
C++ Coding
-
CUDA profiling & debugging
-
ML runtime optimization
-
Custom GPU kernel development
-
The expected base salary range for this full-time position across US locations is listed below. Actual starting pay will be based on job-related factors, including exact work location, experience, relevant training and education, and skill level. Your recruiter can share more about the specific salary range for the role location or, if the role can be performed remote, the specific salary range for your preferred location, during the hiring process.
Waymo employees are also eligible to participate in Waymo’s discretionary annual bonus program, equity incentive plan, and generous Company benefits program, subject to eligibility requirements.
Salary Range
$213,000—$263,000 USD
$159.05k - $199.3k
...intelligence to every moving machine on the planet. Applied... ...; Seoul; and Tokyo. Learn more at applied.co. We... ...looking for a software engineer with deep experience in... ...‑grade embedded runtime environments. You’ll work... ...with ML accelerators, GPU, CPU, SoC architecture...SuggestedFull timeFor contractorsFor subcontractorCasual workWork at officeRemote workDay shift$119.25k - $150.85k
...Solutions team deploys machine learning models from training... ...IQ, and we’re hiring engineers to help deliver the next... ...with our sister teams (kernels, compiler, reduced... ...for compiler, kernel, runtime, and parity bugs. Collaborate... ..., ML compilers, GPU programming (CUDA, OpenAI...SuggestedFull timeInternshipLocal areaWork from homeRelocation packageFlexible hours$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance... ...across kernel execution, compiler decisions, and runtime scheduling.Direct experience with LLM inference frameworks...SuggestedFull time$204k - $259k
...automate the lifecycle of the machine learning workflow, including feature... .... We are looking for engineers with ML software or ML systems... ...to the Senior Manager of Runtime and Optimization. You will... ...accelerators (e.g., GPU/TPU). Deep knowledge of...SuggestedFull timeRemote work- ...intelligence to every moving machine on the planet.... ...Seoul; and Tokyo. Learn more at applied.co.... ...for a performance engineer who specializes in... ...wastes 30% of its GPU‑hours on stalled... ...preprocessing, augmentation, kernel execution, gradient... ..., TensorRT, ONNX Runtime, Ray, or similar)...SuggestedFull timeFor contractorsFor subcontractorCasual workWork at officeRemote workDay shift
$150k
..., data scientists, and engineers, tackling the most fundamental... ...computing in deep learning, driving impactful... ...optimizing performance for the machine learning software... ...propose solutions to enhance GPU utilization Support... ...to develop appropriate kernels and systems for new...Full timeWork experience placementVisa sponsorship$213k - $263k
...automate the lifecycle of the machine learning workflow, including feature... .... We are looking for engineers with ML software & systems expertise... ...to the Senior Manager of Runtime and Optimization. You... ...Experience with custom kernel development (e.g., CUDA/CUDA...Full timeRemote work- ...Senior ML Infrastructure Engineer (PyTorch, Kubernetes, GPU Training) Short Job Description... ...infrastructure powering large-scale machine learning training workloads. In this... ...data pipelines, Triton, TensorRT, custom ML kernels, or ML compiler/runtime optimization....Full time
- Tilde Research is seeking a Kernel Engineer in Palo Alto to design, implement, and optimize high-performance GPU kernels for training and inference workloads. You will collaborate with ML researchers and engineers to push hardware limits, improve throughput, and reduce...
- ...Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design and optimize kernels, build new abstractions for LLM serving engines, and contribute to accelerators and runtimes that power large language models and AI workloads...
- ...California is looking for a Member of Technical Staff for Kernel/Compiler/Communication. This critical role requires strong expertise in CUDA and GPU optimization, along with 5+ years of experience in performance engineering. The ideal candidate will design high-performance...
- CoreWeave is seeking a Senior Software Engineer for the Systems Engineering team to own kernel tracing and patching across Kubernetes and container runtimes. You will debug complex failures, trace root causes in the Linux kernel, and upstream fixes where appropriate. This...
$165.2k - $223.6k
...used to accelerate deep learning and GenAI workloads on Amazon’s custom machine learning accelerators,... ...Trainium.The Acceleration Kernel Library team is at the... ...software boundary, our engineers craft high-performance... ...an ML compiler, runtime, and application framework...InternshipLocal areaWork from homeFlexible hours$300k
..., data scientists, and engineers, tackling the most fundamental... ...computing in deep learning, driving impactful... ...across multi-node, multi-GPU clusters • Own experiment... ...with large-scale machine learning workloads (strong... ...performance profiling, kernel fusion, or memory optimization...Full timeFlexible hours- ...patients worldwide.We’re a team of engineers, clinicians, and innovators... ...This Opportunity:As a Staff Machine Learning Engineer, you will be responsible... ...for real-time inference, GPU/throughput optimization (e.g., TensorRT, ONNX Runtime, mixed precision), and building...Work at officeLocal areaWorldwideFlexible hours
$165.45k - $259.75k
...Senior Software Engineer Description - HP is seeking a Senior Software... ..., evaluate, and deploy machine learning models — including LLMs, vision... ...TensorFlow / TensorFlow Lite, ONNX Runtime , or Hugging Face... ...Cloud , containerized inference, GPU/NPU acceleration). Strong Python...Full timeTemporary workLocal areaRelocationFlexible hoursShift work$150.4k - $277.6k
...Compute Frameworks team in GPU, Graphics and Displays org provides... ...algebra, image processing, machine learning, along with other projects... ...and GPU programming engineers who are passionate about providing... ...optimized GPU compute kernels across Machine Learning, Image...Relocation$200k - $300k
...Machine Learning Systems EngineerLocation - Palo Alto, CA (On-site) - Five... ..., and research-driven, with engineers working directly alongside... ...learning, distributed systems, GPU infrastructure, and high-... ...such as vLLM, TensorRT, ONNX Runtime, and SGLangDevelop and optimize...H1bWork at office- Junior Distributed Machine Learning Engineer Join to apply for the Junior Distributed Machine Learning Engineer role... ..., network and propose solutions to enhance GPU utilization • Support the team to develop appropriate kernels and systems for new model architectures and...Full timeWork experience placementH1b
- Apple Inc. in Cupertino, CA, is seeking extraordinary machine learning and GPU programming engineers to deliver robust compute solutions on Apple Silicon. You will contributes to optimized kernels, ML workflows, and high-performance GPU architectures across iOS, macOS...
$2,000 per month
...Machine Learning Research EngineerCupertino, CAEtched is building... ...tools, and runtime abstractions by implementing... ...models on Sohu to achieve GPU-impossible latencies... ...with GPU kernels, the CUDA compilation... ...Cupertino, and greatly value engineering skills. We do not have...- ...CoreWeave is seeking a Senior Engineer for its Benchmarking & Performance team to own kernel-level optimization for LLM inference and end-to-end model serving, focusing on CUDA kernels and throughput/latency improvements. You will lead kernel design reviews, mentor...
$250k - $350k
...Senior/Staff level Inference Engineers to accelerate the performance... ...edge inference acceleration, GPU parallelism, advanced model deployment... ...high-performance computing kernels and distributed workloads... ...acceleration, and deep learning compiler stacks.GPU & Parallelism...Work at office3 days per week$195.2k - $262.2k
...infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to... ...Partner with GPU kernel engineers and platform engineers... ...across model code, kernels, runtime, scheduler, gateway, and... ...Career growth and learning opportunities Flexibility...Full timeTemporary workImmediate startRemote work- ...Technical Staff to push the limits of performance at the kernel, compiler, and communication layers. You will optimize runtimes and libraries to unlock maximum efficiency on modern accelerators across large GPU clusters. You will design high-performance kernels, improve...
$184k - $287.5k
...computing. An era in which our GPU acts as the brains of... ...for a Senior Software Engineer to join the... ...ships the production GPU kernels and software interfaces... ...power equivariant deep learning throughout the scientific... ...concepts in geometric machine learning - equivariance...- ...production-grade workflows. You will work at the intersection of distributed systems, GPU performance, and ML framework integration. The role requires strong Python and PyTorch engineering skills, hands-on experience with distributed model training, and the ability to...
$184k - $287.5k
We are now looking for a Senior Formal Verification Engineer for GPU Kernels! Modern AI performance relies on highly optimized GPU kernels — performance... ...can be hard to catch and expensive to miss. NVIDIA's Deep Learning Safety Team is hiring engineers to build the verification...Full timeWork experience placement$124k - $195.5k
...Deep Learning Software Engineer, TensorRT PerformanceNVIDIA is seeking an experienced... ...specialize in developing GPU-accelerated deep learning... ...modeling, performance analysis, kernel development and inference... ...the foundation for machines to learn, perceive, reason...$184.7k - $324.8k
...Cupertino, California, United States Machine Learning and AI Video is at the core of nearly all Apple products, and as a research engineer on our team, you will build the infrastructure... ...custom ops in CUDA or low-level GPU kernel optimization Research experience in ML...Relocation
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Machine Learning Engineer, GPU Kernel and Runtime. Be the first to apply!
- ai ml engineer Mountain View, CA
- senior ml engineer Mountain View, CA
- computer vision machine learning engineer Mountain View, CA
- machine learning engineer Mountain View, CA
- machine learning ai engineer Mountain View, CA
- machine learning software engineer Mountain View, CA
- machine learning scientist Mountain View, CA
- machine learning intern Mountain View, CA
- machine learning remote Mountain View, CA
- machine learning researcher Mountain View, CA




