Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

GPU Kernel Engineer — Fast ML Training & Inference

Tilde Research

Tilde Research is seeking a Kernel Engineer in Palo Alto to design, implement, and optimize high-performance GPU kernels for training and inference workloads. You will collaborate with ML researchers and engineers to push hardware limits, improve throughput, and reduce latency across models and infrastructure. This role emphasizes performance-aware software, scalable kernels, and hands-on experimentation with CUDA, PyTorch, and Triton. #J-18808-Ljbffr Tilde Research

Vacancy posted 22 hours ago
Similar jobs that could be interesting for youBased on the GPU Kernel Engineer — Fast ML Training & Inference in Palo Alto, CA vacancy
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...with patients. We have trained our own LLMs as part...  ...together. To support fast collaboration and a strong...  ...an experienced LLM Inference Engineer to optimize our large...  ...deployment scenarios and GPU types What You...  ...Experience with custom CUDA kernels Track record of... 
    Training
    Work at office

    Hippocratic AI

    Menlo Park, CA
    1 day ago
  •  ...deliver industry‑leading training and inference speeds and empowers...  ...effortlessly run large‑scale ML applications, without...  ...10 times faster than GPU‑based hyperscale cloud...  ..., TensorRT‑LLM), GPU kernel‑level optimization...  ...Collaborate with Product and Engineering to identify where... 
    Training
    Contract work
    Shift work

    Cerebras

    Sunnyvale, CA
    22 hours ago
  • $152k - $241.5k

    NVIDIA's invention of the GPU 1999 sparked the...  ...top-tier AI Compiler Engineers to drive innovation within...  ...focusing on kernel generation and computational...  ...for AI workloads (both inference and training) and successfully transition...  ...in a dynamic, fast-paced, and product-oriented... 
    Training
    Full time

    Nvidia

    Santa Clara, CA
    7 hours ago
  • $300k - $400k

     ...makes our frontier model training and inference fast, efficient, and...  ...the stack: scheduling, kernels, RDMA, weight synchronization...  ...communication and GPU kernels to extract...  ...distributed ML systems to identify and...  ...best — the scientists, engineers, and problem-solvers... 
    Training
    Visa sponsorship
    Flexible hours
    Shift work

    Periodic Labs

    Menlo Park, CA
    4 days ago
  •  ...push the limits of performance at the kernel, compiler, and communication layers. You...  ...efficiency on modern accelerators across large GPU clusters. You will design high-...  ...and runtime stacks, and collaborate with training and inference teams to reduce latency and increase... 
    Training

    RadixArk

    Palo Alto, CA
    22 hours ago
  •  ...intelligence. About The Role As a Kernel Engineer at Tilde, you'll design,...  ...and optimize high-performance GPU kernels that are critical to scaling our training and inference workloads. Your work will enable...  .... You'll work closely with ML researchers and engineers to co... 
    Training
    Full time
    Internship

    Tilde Research

    Palo Alto, CA
    22 hours ago
  • $195.2k - $361.2k

     ...Role Summary Make models fast on the hardware people...  ...own. You optimize inference engines (llama.cpp, vLLM) for...  ...and edge environments — GPU/iGPUs, Vulkan backends...  ...impact with the Post-Training team Cut CPU overhead...  ...Metal) or SIMD / CPU kernels Familiarity with quantization... 
    Training
    Internship
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    1 day ago
  • CoreWeave is seeking a Senior Engineer for its Benchmarking & Performance team to own kernel-level optimization for LLM inference and end-to-end model serving, focusing on CUDA kernels and throughput/latency improvements. You will lead kernel design reviews, mentor engineers... 

    Neura Market

    Sunnyvale, CA
    2 days ago
  • $166k - $244k

     ...looking for a Software Engineer, Edge Systems &...  ...pipelines and GPU runtime engines to...  ...inputs into model inference loops.Collaborate...  ...smooth handoffs from training pipelines to production...  ....On-Device ML Deployment: 3+ years...  ...writing custom CUDA kernels, custom TensorRT... 
    Training
    Full time

    X Company

    Mountain View, CA
    3 days ago
  •  ...Corporation in Santa Clara, CA is seeking outstanding AI systems engineers to develop groundbreaking inference technologies for the hardware-accelerated stack. You will create libraries, code generators, and GPU kernel innovations for LLM workloads. Join a team that designs... 

    NVIDIA Corporation

    Santa Clara, CA
    4 days ago
  • NVIDIA is seeking outstanding AI systems engineers in Santa Clara to advance the inference software stack. You will build libraries, code generators, and GPU kernels for NVIDIA hardware, designing abstractions for LLM serving engines and JIT compilers to accelerate large... 

    Segment (Twilio)

    Santa Clara, CA
    4 days ago
  • $180k

     ...the world’s largest AI supercomputers. You will design and optimize massive GPU clusters, ensuring fast and reliable AI training. Ideal candidates will possess deep programming skills, GPU kernel optimization experience, and a strong grasp of large-scale distributed... 
    Training

    xAI

    Palo Alto, CA
    2 days ago
  • $198k - $326k

     ...layer between model training and production...  ...Senior Staff Software Engineer with deep expertise...  ...machine learning, GPU infrastructure, and large-scale inference. This is a highly technical...  ...runtime, compiler, kernel, and hardware...  ...improvementsPartner closely with ML, infrastructure,... 
    Training
    For contractors
    Work at office
    Flexible hours

    Linkedin

    Mountain View, CA
    2 days ago
  • $184k - $287.5k

     ...skilled and motivated software engineers to join us and build AI inference systems that serve large-...  ...stacks, optimize GPU kernels and compilers, drive industry...  ...for the field of ML Systems; survey recent publications...  ...; ability to excel in a fast-paced, multi-functional... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...a Sr. HPC Performance engineer to join our team of scientists...  ...machine learning (ML) frameworks. Starting...  ...scale, CUDA-backed ML training frameworks, using low...  ...strategies such as kernel design, GPU porting, data structure...  ...engineering teams are growing fast in some of the hottest... 
    Training
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • Nebius B.V. is building an AI training and model post-training capability focused on frontier...  ...intersection of distributed systems, GPU performance, and ML framework integration. The role requires strong Python and PyTorch engineering skills, hands-on experience with... 
    Training

    Nebius B.V.

    Palo Alto, CA
    4 days ago
  •  ...the multimodal video, training, and RL pipelines...  ...distributed systems / ML infra. About...  ...a Machine Learning Engineer to scale and optimize...  ...distributed training jobs and inference deployments to maximize GPU/CPU utilization and...  ...end to end in fast-paced applied research... 
    Training

    Orbifold AI

    Palo Alto, CA
    4 days ago
  •  ...advance your career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of GPU-accelerated...  ...bottlenecks spanning HBM bandwidth, compute utilization, kernel scheduling, memory allocation, and PCIe/Infinity Fabric... 

    AMD

    Santa Clara, CA
    2 days ago
  • $250k - $350k

     ...seeking Senior/Staff level Inference Engineers to accelerate the...  ...inference acceleration, GPU parallelism, advanced...  ...performance computing kernels and distributed...  ...into production.Improve Training Efficiency: (Bonus) Contribute...  ...ambiguity in a fast-paced startup environment... 
    Training
    Work at office
    3 days per week

    Pika

    Palo Alto, CA
    3 days ago
  • $250k - $300k

    Performance/ Benchmark Engineer - NVIDIA GPU SystemsNVIDIA GPU Systems / AI Inference / Performance EngineeringOverview Seeking...  ...NVIDIA GPU compute platforms and AI/ML performance benchmarking.Strong...  ...workloads, distributed training, or large-scale GPU clusters.Familiarity... 
    Training

    Yoh

    Santa Clara, CA
    2 days ago
  • $120k - $275k

     ...including hardware and software to train and run the largest ML workloads for AGI. MatX is...  ...-architects and design engineers to join our team as we...  ...experience in SoC, AI accelerator, GPU, networking ASIC, or high-...  ...to work independently in a fast-paced environment.Excellent... 
    Training
    Full time
    Work experience placement
    Local area
    Remote work
    Monday to Friday
    Flexible hours

    MatX

    Mountain View, CA
    4 days ago
  •  ...Machine Learning Software Engineers to build compute-...  ...across teams to deploy CV/ML models into production...  ...for dataset management, training, and deployment Participate...  ...(both training and inference) ~ Knowledge of model...  ...responsibility in a fast-paced startup environment... 
    Training
    Full time
    Visa sponsorship

    Corvus Robotics

    Mountain View, CA
    1 day ago
  •  ...Systems Performance Engineer Palo Alto,...  ...talented and driven ML performance engineer...  ...for large‑scale AI inference. Responsibilities...  ...Compiler, runtime, or kernel‑level optimization...  ...or multimodal model training and inference. Background...  ...TensorRT. Strong GPU programming skills... 
    Training
    Full time
    Temporary work
    Local area
    Flexible hours

    SambaNova

    Palo Alto, CA
    22 hours ago
  •  ...Moveworks was also named one of Fast Company’s 2025 Most...  ...automation with Moveworks’ Reasoning Engine and natural language...  ...Engineer to help build cutting edge ML infrastructure for building and...  ...including distributed training and inference pipeline for large language models... 
    Training
    Work at office
    Remote work
    Flexible hours

    Moveworks

    Mountain View, CA
    3 days ago
  • $120k - $275k

     ...including hardware and software to train and run the largest ML workloads for AGI. MatX is...  ...-architects and design engineers to join our team as we...  ...directed and effective in a fast-moving, low-overhead startup...  ...HaveExperience with AI accelerator, GPU, or custom ASIC... 
    Training
    Full time
    Work experience placement
    Local area
    Remote work
    Monday to Friday
    Flexible hours

    MatX

    Mountain View, CA
    1 day ago
  •  ...expertise will be critical in enhancing GPU kernels, deep learning models, and training/inference performance across multi-GPU and...  ...technologies and advanced engineering principles to drive continuous...  ...integrate GPU-accelerated compute into ML frameworks (e.g., PyTorch,... 
    Training

    AMD

    Santa Clara, CA
    2 days ago
  • $150k - $230k

     ...Driven Fabrics to increase GPU cluster utilization...  ...researchers and veteran systems engineers who share a vision for...  ...distributed GPU training. You'll work at the...  ...operated them. Examples: Kernel subsystems, device...  ...networking (RDMA, InfiniBand) ML framework or runtime... 
    Training

    Clockwork.io

    Palo Alto, CA
    more than 2 months ago
  • $124k - $195.5k

     ...optimization, custom kernel development, and cluster...  ...systems from a single GPU to supercomputer...  ...performance bottlenecks in both training and inference pipelines.Collaborate...  ...Science, Computer Engineering, Electrical Engineering...  ...compilers and ML systems, including graph... 
    Training
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • RadixArk in California is seeking a Member of Technical Staff - CI Engineer to own the infrastructure powering SGLang. You’ll keep 300+ GPU tests across diverse hardware green and fast. You’ll build regression‑based CI, harden runners, reduce CI time, and improve developer... 

    RadixArk

    Palo Alto, CA
    22 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to GPU Kernel Engineer — Fast ML Training & Inference. Be the first to apply!