Machine Learning Engineer — GPU Kernel
$150kJobleads-US
About the Institute of Foundation Models
We are a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.
The Role
The GPU Kernel Engineer will play a role at the forefront of optimizing performance for the machine learning software stacks, especially at training and inference, and support the team to develop new and cutting-edge systems. The ideal candidate will have a strong background in parallel computing, and hands-on experience in system level coding, debug methodologies, and large-scale machine learning experience.
This role focuses on CUDA kernel development and optimization. Distributed training experience is a plus.
Key Responsibilities
- Understand, analyze, profile, optimize, and provide guidance to the team on deep learning workloads on state-of-the-art hardware and software platforms to improve their efficiency with different levels of optimization
- Design and implement performance benchmarks and testing methodologies to evaluate application performance
- Build tools to automate workload analysis, workload optimization, and other critical workflows
- Triage system issues and identify bottleneck and inefficiencies by analyzing the sources of issues and the impact on hardware, network and propose solutions to enhance GPU utilization
- Support the team to develop appropriate kernels and systems for new model architectures and algorithms
- Participate in, or lead design reviews with peers and stakeholders to decide amongst available technologies.
- Review code developed by other developers and provide feedback to ensure best practices (e.g., style guidelines, checking code in, accuracy, testability, and efficiency).
- Contribute to existing documentation or educational content and adapt content based on product/program updates and user feedback.
- Represent MBZUAI at industry conferences and events, showcasing the institution’s cutting-edge HPC anddeep learning capabilities and establishing MBZUAI as a global leader in AI research and innovation.
- Perform all other duties as reasonably directed by the line manager that are commensurate with thesefunctional objectives.
- Validate CUDA kernel outputs and gradients against reference implementations, and benchmark representative shapes, dtypes, and model workloads.
Technical Qualifications
Must-Haves:
- Strong C++ skills and hands-on CUDA kernel development and optimization for deep-learning workloads.
- Understanding of GPU memory hierarchy, warp/block execution, and compute-memory trade-offs, with demonstrated profiling-driven optimization.
- Strong Python skills and experience integrating kernels with PyTorch or an equivalent framework, including numerical and gradient validation where needed.
Nice-to-Haves:
- Experience with Triton, CUTLASS, or PTX/SASS analysis.
- Experience with multi-node distributed training or inference systems.
- Experience validating mixed-precision computations, such as BF16 or FP8.
$150,000 - $450,000 a year
The posted salary range represents the company’s good faith estimate of the compensation for this position upon hire. The actual compensation offered may vary within this range depending on individual qualifications, including but not limited to relevant skills, experience, education, certifications, geographic location, and specific business needs.
Benefits Include
- Comprehensive medical, dental, and vision benefits
- Bonus
- 401K Plan
- Generous paid time off, sick leave and holidays
- Paid Parental Leave
- Employee Assistance Program
- Life insurance and disability
$193.3k - $261.5k
...used to accelerate deep learning and GenAI workloads on Amazon’s custom machine learning accelerators, Inferentia... ....The Acceleration Kernel Library team is at the... ...-software boundary, our engineers craft high-performance... ...architectures- Experience with GPU kernel optimization and...SuggestedInternshipLocal areaWork from homeFlexible hours$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop...SuggestedFull time- ...NVIDIA Corporation in Santa Clara, CA, seeks a Sr. Inference Engineer to accelerate LLM inference through GPU kernel optimization. You will lead kernel benchmarking, model-level performance analysis, and AI-driven optimization workflows across silicon and software stacks...Suggested
- ...speeds; over 10 times faster than GPU-based hyperscale cloud... ...Core ML team develops novel machine learning algorithms that take advantage... ...of the Cerebras Wafer-Scale Engine. Our work spans efficient LLM... ...compilers, runtimes, and low-level kernels to implement new algorithmic...Suggested
$150k
..., data scientists, and engineers, tackling the most fundamental... ...computing in deep learning, driving impactful... ...across multi-node, multi-GPU clusters Own experiment... ...with large-scale machine learning workloads (strong... ...performance profiling, kernel fusion, or memory optimization...SuggestedVisa sponsorshipFlexible hours- ...the foundations to democratize AI and Machine Learning for Atlassian’s teams, customers, and... ...About This RoleAs a Senior ML System Engineer on the AI & ML Platform’s Inference team... ...scaling) to deep low-level optimizations (GPU kernels, quantization, speculative decoding)....Work at officeLocal area
$197.5k - $272k
...Edge. We are looking for a great Staff Machine Learning Engineer to join our seasoned AI team and lead... ...process unstructured application logs, kernel traces, and multi-modalities.Integrate... ...-grade models for execution on CPU/GPU-bound targets or embedded NPUs.Apply quantization...Work at officeWorldwideFlexible hoursShift work3 days per week$278.1k - $347.6k
...entirely within that runtime. As our Principal Engineer for On-Device AI Inference & Systems, you... ..., through export, optimization, and kernel-level tuning, to a shipped feature... ...specific kernel tuning across NPU, mobile GPU, and desktop/laptop GPU. Make authoritative...Work at officeWorldwideRelocation package$150k
..., data scientists, and engineers, tackling the most fundamental... ...computing in deep learning, driving impactful... ...across multi-node, multi-GPU clusters Own experiment... ...with large-scale machine learning workloads (strong... ...performance profiling, kernel fusion, or memory optimization...Flexible hours$184k - $287.5k
...computing. An era where our GPU serves as the intelligence behind... ...with Product, Program, Engineering, and Data Procurement teams... ...proven experience in applied machine learning or AI research. ~ Deep expertise... ..., Triton, or low-level GPU kernel development for inference...$165.6k
...degree in computer science or an equivalent background Responsibilities: Design and build high-performance compute kernels for machine learning operations using the Neuron architecture and programming models Evaluate and improve kernel-level performance across...Full timeInternship$119.25k - $150.85k
...Inference Solutions team deploys machine learning models from training... ...Escalade IQ, and we’re hiring engineers to help deliver the next generation... ...with our sister teams (kernels, compiler, reduced precision... ...ML systems, ML compilers, GPU programming (CUDA, OpenAI Triton...Full timeInternshipLocal areaWork from homeRelocation packageFlexible hours$2,000 per month
...product could never be achieved on a typical GPU Implement diffusion models on Sohu to... ...with Rust Familiarity with GPU kernels, the CUDA compilation stack and related tools... ...person team in Cupertino, and greatly value engineering skills. We do not have boundaries between...Work at officeRelocation package- ...Meta's Santa Clara-based MTIA Software Team is hiring a Software Engineer, Systems ML specializing in compilers and kernels. You will help develop the AI compiler stack, contribute to PyTorch core components, and optimize high-performance kernels for next-generation hardware...
- ...across image, video, and world-model workloads. You will work with a founding team on kernels, runtimes, and distributed engines that power production-scale ML stacks. You’ll optimize GPU performance, profile bottlenecks with Nsight, and implement low-level CUDA and...
$152k - $241.5k
NVIDIA is looking for a talented Machine Learning Engineer to drive the development, evaluation, deployment and end-to-end lifecycle management of... ...efficiently across distributed infrastructure. You will manage GPU orchestration, prompt-tune models, and build advanced AI...Full timeFlexible hours- ...THE ROLE: We are looking for a talented engineer to join our team: developing heterogeneous... ...: Design, develop, and optimize GPU/CPU software for computer vision, image... ...Experience in video codecs, image processing and machine learning frameworksFamiliarity with computer...
$224k - $356.5k
NVIDIA is looking for a Machine Learning Engineer to join the GPU accelerated Apache Spark team.Apache Spark is the most popular data processing engine in data centers for running massive scale workloads for ETL, SQL, and ML/DL model training and inference pipelines, spanning...Full time$171k - $231.5k
...the Fintech Risk AI/ML Smart Money Services (SMS) as a Senior Machine Learning Engineer. The SMS team is responsible for detecting and preventing... ...scalable software supporting millions or more usersExperience with GPU acceleration (i.e CUDA and cuDNN)Experience with integrating...WorldwideShift work$165.2k - $223.6k
...revolution? At Amazon our vision is to make deep learning pervasive for everyday developers and to... ...workloads.This role is for a software engineer in the Compiler team for AWS Neuron. As... ....- Experience in compiler design for CPU/GPU/Vector engines/ML-accelerators.-...InternshipLocal areaFlexible hours- ...patients worldwide.We’re a team of engineers, clinicians, and innovators... ..., CA for a SeniorMachine Learning Engineer with focus on Computer... ...computer vision and machine learning applications; or PhD... ...Strong hands-on experience with GPU accelerated algorithms and implementations...Local areaImmediate startWorldwideFlexible hours
$143k - $286k
...WalmartBusiness Segment: Home OfficeRole summary: The Senior Machine Learning Engineer will lead the design, development, and deployment of... ...pipeline development.Familiarity with large language modeling and GPU optimization techniques. At Walmart, we offer competitive pay...Full timeTemporary workPart timeWork experience placement$150k - $230k
...information, visit About the RoleWe are looking for a hands-on Machine Learning Engineer to drive the post-training of our large language models,... ...on tight timelines.Run large-scale training on mid-to-large GPU clusters, applying distributed-training techniques (data parallelism...Full timeLocal areaWork from home$184k - $287.5k
Intelligent machines powered by Artificial Intelligence computers that can learn, reason and interact with people are no longer science fiction. GPU Deep Learning has provided the foundation for machines... ...Senior Perception Engineer to develop and productize NVIDIA...Odd jobFull timeWork experience placementRemote work$207k - $300k
...implement the real-time on-device machine learning foundation that turns raw... ...of outlook models for NPU, GPU, and DSP execution,... ...architecture design, and custom kernel development, in partnership... ...qualifications:Master’s degree or PhD in Engineering, Computer Science, or a...Immediate start$184k - $287.5k
We are seeking a self‑motivated senior engineer for the Aerial Omniverse Digital Twin team.... ...design and implementation of a real‑time, GPU‑accelerated propagation engine that... ...and compute‑vs‑bandwidth trade‑offs at the kernel level.Working knowledge of electromagnetic...Full time$152k - $241.5k
...Automotive, VR, Gaming, Deep Learning, and High Performance Computing... ...helped craft as a member of the GPU Foundations Developer Tools... ...Collaborate with developer tools engineers, software library developers,... ...and both user-mode and kernel-mode drivers.Proven knowledge...Full time- ...advance your career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of GPU-accelerated... ...bottlenecks spanning HBM bandwidth, compute utilization, kernel scheduling, memory allocation, PCIe/Infinity Fabric data...
$152k - $241.5k
...NVIDIA is hiring a Senior Compiler Engineer to join our team driving the next generation of GPU systems programming. We are... ...memory-safe, high-performance GPU kernels in idiomatic Rust.What you’ll be... ...to high-performance CUDA PTX and machine code.Build compiler IRs and...Full timeRemote work$160k - $200k
Santa Clara, CAData Engineering - ML Infrastructure /Full-time /HybridPlusAI... ...for managing large-scale GPU clusters. This role offers... ...integrated with state-of-the-art deep learning frameworks like PyTorch or... ...of what's possible in machine learning infrastructure and contribute...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Machine Learning Engineer — GPU Kernel. Be the first to apply!
- computer vision machine learning engineer Sunnyvale, CA
- machine learning software engineer Sunnyvale, CA
- ai ml engineer Sunnyvale, CA
- senior ml engineer Sunnyvale, CA
- machine learning ai engineer Sunnyvale, CA
- machine learning engineer Sunnyvale, CA
- data engineer machine learning Sunnyvale, CA
- artificial intelligence - machine learning intern Sunnyvale, CA
- machine learning intern Sunnyvale, CA
- internship machine learning Sunnyvale, CA



