GPU Kernel Engineer — Fast ML Training & Inference
Tilde Research
Tilde Research is seeking a Kernel Engineer in Palo Alto to design, implement, and optimize high-performance GPU kernels for training and inference workloads. You will collaborate with ML researchers and engineers to push hardware limits, improve throughput, and reduce latency across models and infrastructure. This role emphasizes performance-aware software, scalable kernels, and hands-on experimentation with CUDA, PyTorch, and Triton. #J-18808-Ljbffr Tilde Research
$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop...SuggestedFull time- ...with patients. We have trained our own LLMs as part... ...together. To support fast collaboration and a strong... ...an experienced LLM Inference Engineer to optimize our large... ...deployment scenarios and GPU types What You... ...Experience with custom CUDA kernels Track record of...TrainingWork at office
- ...deliver industry‑leading training and inference speeds and empowers... ...effortlessly run large‑scale ML applications, without... ...10 times faster than GPU‑based hyperscale cloud... ..., TensorRT‑LLM), GPU kernel‑level optimization... ...Collaborate with Product and Engineering to identify where...TrainingContract workShift work
$152k - $241.5k
NVIDIA's invention of the GPU 1999 sparked the... ...top-tier AI Compiler Engineers to drive innovation within... ...focusing on kernel generation and computational... ...for AI workloads (both inference and training) and successfully transition... ...in a dynamic, fast-paced, and product-oriented...TrainingFull time$300k - $400k
...makes our frontier model training and inference fast, efficient, and... ...the stack: scheduling, kernels, RDMA, weight synchronization... ...communication and GPU kernels to extract... ...distributed ML systems to identify and... ...best — the scientists, engineers, and problem-solvers...TrainingVisa sponsorshipFlexible hoursShift work- ...push the limits of performance at the kernel, compiler, and communication layers. You... ...efficiency on modern accelerators across large GPU clusters. You will design high-... ...and runtime stacks, and collaborate with training and inference teams to reduce latency and increase...Training
- ...intelligence. About The Role As a Kernel Engineer at Tilde, you'll design,... ...and optimize high-performance GPU kernels that are critical to scaling our training and inference workloads. Your work will enable... .... You'll work closely with ML researchers and engineers to co...TrainingFull timeInternship
$195.2k - $361.2k
...Role Summary Make models fast on the hardware people... ...own. You optimize inference engines (llama.cpp, vLLM) for... ...and edge environments — GPU/iGPUs, Vulkan backends... ...impact with the Post-Training team Cut CPU overhead... ...Metal) or SIMD / CPU kernels Familiarity with quantization...TrainingInternshipLocal areaImmediate startShift work- CoreWeave is seeking a Senior Engineer for its Benchmarking & Performance team to own kernel-level optimization for LLM inference and end-to-end model serving, focusing on CUDA kernels and throughput/latency improvements. You will lead kernel design reviews, mentor engineers...
$166k - $244k
...looking for a Software Engineer, Edge Systems &... ...pipelines and GPU runtime engines to... ...inputs into model inference loops.Collaborate... ...smooth handoffs from training pipelines to production... ....On-Device ML Deployment: 3+ years... ...writing custom CUDA kernels, custom TensorRT...TrainingFull time- ...Corporation in Santa Clara, CA is seeking outstanding AI systems engineers to develop groundbreaking inference technologies for the hardware-accelerated stack. You will create libraries, code generators, and GPU kernel innovations for LLM workloads. Join a team that designs...
- NVIDIA is seeking outstanding AI systems engineers in Santa Clara to advance the inference software stack. You will build libraries, code generators, and GPU kernels for NVIDIA hardware, designing abstractions for LLM serving engines and JIT compilers to accelerate large...
$180k
...the world’s largest AI supercomputers. You will design and optimize massive GPU clusters, ensuring fast and reliable AI training. Ideal candidates will possess deep programming skills, GPU kernel optimization experience, and a strong grasp of large-scale distributed...Training$198k - $326k
...layer between model training and production... ...Senior Staff Software Engineer with deep expertise... ...machine learning, GPU infrastructure, and large-scale inference. This is a highly technical... ...runtime, compiler, kernel, and hardware... ...improvementsPartner closely with ML, infrastructure,...TrainingFor contractorsWork at officeFlexible hours$184k - $287.5k
...skilled and motivated software engineers to join us and build AI inference systems that serve large-... ...stacks, optimize GPU kernels and compilers, drive industry... ...for the field of ML Systems; survey recent publications... ...; ability to excel in a fast-paced, multi-functional...Full time$184k - $287.5k
...a Sr. HPC Performance engineer to join our team of scientists... ...machine learning (ML) frameworks. Starting... ...scale, CUDA-backed ML training frameworks, using low... ...strategies such as kernel design, GPU porting, data structure... ...engineering teams are growing fast in some of the hottest...TrainingFull time- Nebius B.V. is building an AI training and model post-training capability focused on frontier... ...intersection of distributed systems, GPU performance, and ML framework integration. The role requires strong Python and PyTorch engineering skills, hands-on experience with...Training
- ...the multimodal video, training, and RL pipelines... ...distributed systems / ML infra. About... ...a Machine Learning Engineer to scale and optimize... ...distributed training jobs and inference deployments to maximize GPU/CPU utilization and... ...end to end in fast-paced applied research...Training
- ...advance your career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of GPU-accelerated... ...bottlenecks spanning HBM bandwidth, compute utilization, kernel scheduling, memory allocation, and PCIe/Infinity Fabric...
$250k - $350k
...seeking Senior/Staff level Inference Engineers to accelerate the... ...inference acceleration, GPU parallelism, advanced... ...performance computing kernels and distributed... ...into production.Improve Training Efficiency: (Bonus) Contribute... ...ambiguity in a fast-paced startup environment...TrainingWork at office3 days per week$250k - $300k
Performance/ Benchmark Engineer - NVIDIA GPU SystemsNVIDIA GPU Systems / AI Inference / Performance EngineeringOverview Seeking... ...NVIDIA GPU compute platforms and AI/ML performance benchmarking.Strong... ...workloads, distributed training, or large-scale GPU clusters.Familiarity...Training$120k - $275k
...including hardware and software to train and run the largest ML workloads for AGI. MatX is... ...-architects and design engineers to join our team as we... ...experience in SoC, AI accelerator, GPU, networking ASIC, or high-... ...to work independently in a fast-paced environment.Excellent...TrainingFull timeWork experience placementLocal areaRemote workMonday to FridayFlexible hours- ...Machine Learning Software Engineers to build compute-... ...across teams to deploy CV/ML models into production... ...for dataset management, training, and deployment Participate... ...(both training and inference) ~ Knowledge of model... ...responsibility in a fast-paced startup environment...TrainingFull timeVisa sponsorship
- ...Systems Performance Engineer Palo Alto,... ...talented and driven ML performance engineer... ...for large‑scale AI inference. Responsibilities... ...Compiler, runtime, or kernel‑level optimization... ...or multimodal model training and inference. Background... ...TensorRT. Strong GPU programming skills...TrainingFull timeTemporary workLocal areaFlexible hours
- ...Moveworks was also named one of Fast Company’s 2025 Most... ...automation with Moveworks’ Reasoning Engine and natural language... ...Engineer to help build cutting edge ML infrastructure for building and... ...including distributed training and inference pipeline for large language models...TrainingWork at officeRemote workFlexible hours
$120k - $275k
...including hardware and software to train and run the largest ML workloads for AGI. MatX is... ...-architects and design engineers to join our team as we... ...directed and effective in a fast-moving, low-overhead startup... ...HaveExperience with AI accelerator, GPU, or custom ASIC...TrainingFull timeWork experience placementLocal areaRemote workMonday to FridayFlexible hours- ...expertise will be critical in enhancing GPU kernels, deep learning models, and training/inference performance across multi-GPU and... ...technologies and advanced engineering principles to drive continuous... ...integrate GPU-accelerated compute into ML frameworks (e.g., PyTorch,...Training
$150k - $230k
...Driven Fabrics to increase GPU cluster utilization... ...researchers and veteran systems engineers who share a vision for... ...distributed GPU training. You'll work at the... ...operated them. Examples: Kernel subsystems, device... ...networking (RDMA, InfiniBand) ML framework or runtime...Training$124k - $195.5k
...optimization, custom kernel development, and cluster... ...systems from a single GPU to supercomputer... ...performance bottlenecks in both training and inference pipelines.Collaborate... ...Science, Computer Engineering, Electrical Engineering... ...compilers and ML systems, including graph...TrainingFull time- RadixArk in California is seeking a Member of Technical Staff - CI Engineer to own the infrastructure powering SGLang. You’ll keep 300+ GPU tests across diverse hardware green and fast. You’ll build regression‑based CI, harden runners, reduce CI time, and improve developer...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to GPU Kernel Engineer — Fast ML Training & Inference. Be the first to apply!
- machine learning researcher Palo Alto, CA
- machine learning research scientist Palo Alto, CA
- machine learning Palo Alto, CA
- artificial intelligence - machine learning intern Palo Alto, CA
- machine learning intern Palo Alto, CA
- machine learning remote Palo Alto, CA
- data engineer machine learning Palo Alto, CA
- machine learning scientist Palo Alto, CA
- internship machine learning Palo Alto, CA
- machine learning drug discovery



