Machine Learning Engineer -- GPU Kernel
Institute of Foundation Models
Job Description
Job Description
About the Institute of Foundation Models
We are a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.
As part of our team, you’ll have the opportunity to work on the core of cutting-edge foundation model training, alongside world-class researchers, data scientists, and engineers, tackling the most fundamental and impactful challenges in AI development. You will participate in the development of groundbreaking AI solutions that have the potential to reshape entire industries. Strategic and innovative problem-solving skills will be instrumental in establishing MBZUAI as a global hub for high-performance computing in deep learning, driving impactful discoveries that inspire the next generation of AI pioneers.
The Role
The GPU Kernel Engineer will play a role at the forefront of optimizing performance for the machine learning software stacks, especially at training and inference, and support the team to develop new and cutting-edge systems. The ideal candidate will have a strong background in parallel computing, and hands-on experience in system level coding, debug methodologies, and large-scale machine learning experience.
This role focuses on CUDA kernel development and optimization. Distributed training experience is a plus.
Key Responsibilities- Understand, analyze, profile, optimize, and provide guidance to the team on deep learning workloads on state-of-the-art hardware and software platforms to improve their efficiency with different levels of optimization
- Design and implement performance benchmarks and testing methodologies to evaluate application performance
- Build tools to automate workload analysis, workload optimization, and other critical workflows
- Triage system issues and identify bottleneck and inefficiencies by analyzing the sources of issues and the impact on hardware, network and propose solutions to enhance GPU utilization
- Support the team to develop appropriate kernels and systems for new model architectures and algorithms
- Participate in, or lead design reviews with peers and stakeholders to decide amongst available technologies.
- Review code developed by other developers and provide feedback to ensure best practices (e.g., style guidelines, checking code in, accuracy, testability, and efficiency).
- Contribute to existing documentation or educational content and adapt content based on product/program updates and user feedback.
- Represent MBZUAI at industry conferences and events, showcasing the institution’s cutting-edge HPC and deep learning capabilities and establishing MBZUAI as a global leader in AI research and innovation.
- Perform all other duties as reasonably directed by the line manager that are commensurate with these functional objectives.
- Validate CUDA kernel outputs and gradients against reference implementations, and benchmark representative shapes, dtypes, and model workloads.
Must-Haves:
- Strong C++ skills and hands-on CUDA kernel development and optimization for deep-learning workloads.
- Understanding of GPU memory hierarchy, warp/block execution, and compute-memory trade-offs, with demonstrated profiling-driven optimization.
- Strong Python skills and experience integrating kernels with PyTorch or an equivalent framework, including numerical and gradient validation where needed.
- Experience with Triton, CUTLASS, or PTX/SASS analysis.
- Experience with multi-node distributed training or inference systems.
- Experience validating mixed-precision computations, such as BF16 or FP8.
Salary Range
The posted salary range represents the company’s good faith estimate of the compensation for this position upon hire. The actual compensation offered may vary within this range depending on individual qualifications, including but not limited to relevant skills, experience, education, certifications, geographic location, and specific business needs.
The posted salary range represents the company’s good faith estimate of the compensation for this position upon hire. The actual compensation offered may vary within this range depending on individual qualifications, including but not limited to relevant skills, experience, education, certifications, geographic location, and specific business needs.
Benefits Include
*Comprehensive medical, dental, and vision benefits
*Bonus
*401K Plan
*Generous paid time off, sick leave and holidays
*Paid Parental Leave
*Employee Assistance Program
*Life insurance and disability
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
$193.3k - $261.5k
...used to accelerate deep learning and GenAI workloads on Amazon’s custom machine learning accelerators, Inferentia... ....The Acceleration Kernel Library team is at the... ...-software boundary, our engineers craft high-performance... ...architectures- Experience with GPU kernel optimization and...SuggestedInternshipLocal areaWork from homeFlexible hours$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop...SuggestedFull time- ...speeds; over 10 times faster than GPU-based hyperscale cloud... ...Core ML team develops novel machine learning algorithms that take advantage... ...of the Cerebras Wafer-Scale Engine. Our work spans efficient LLM... ...compilers, runtimes, and low-level kernels to implement new algorithmic...Suggested
- ...the foundations to democratize AI and Machine Learning for Atlassian’s teams, customers, and... ...About This RoleAs a Senior ML System Engineer on the AI & ML Platform’s Inference team... ...scaling) to deep low-level optimizations (GPU kernels, quantization, speculative decoding)....SuggestedWork at officeLocal area
$197.5k - $272k
...Edge. We are looking for a great Staff Machine Learning Engineer to join our seasoned AI team and lead... ...process unstructured application logs, kernel traces, and multi-modalities.Integrate... ...-grade models for execution on CPU/GPU-bound targets or embedded NPUs.Apply quantization...SuggestedWork at officeWorldwideFlexible hoursShift work3 days per week$278.1k - $417.1k
...entirely within that runtime. As our Principal Engineer for On-Device AI Inference & Systems, you... ..., through export, optimization, and kernel-level tuning, to a shipped feature... ...specific kernel tuning across NPU, mobile GPU, and desktop/laptop GPU. Make authoritative...Work at officeWorldwideRelocation package- ..., data scientists, and engineers, tackling the most fundamental... ...computing in deep learning, driving impactful... ...across multi-node, multi-GPU clusters Own experiment... ...with large-scale machine learning workloads (strong... ...performance profiling, kernel fusion, or memory optimization...Flexible hours
$119.25k - $150.85k
...Inference Solutions team deploys machine learning models from training... ...Escalade IQ, and we’re hiring engineers to help deliver the next generation... ...with our sister teams (kernels, compiler, reduced precision... ...ML systems, ML compilers, GPU programming (CUDA, OpenAI Triton...Full timeInternshipLocal areaWork from homeRelocation packageFlexible hours$2,000 per month
...product could never be achieved on a typical GPU Implement diffusion models on Sohu to... ...with Rust Familiarity with GPU kernels, the CUDA compilation stack and related tools... ...person team in Cupertino, and greatly value engineering skills. We do not have boundaries between...Work at officeRelocation package- ...Meta's Santa Clara-based MTIA Software Team is hiring a Software Engineer, Systems ML specializing in compilers and kernels. You will help develop the AI compiler stack, contribute to PyTorch core components, and optimize high-performance kernels for next-generation hardware...
$152k - $241.5k
NVIDIA is looking for a talented Machine Learning Engineer to drive the development, evaluation, deployment and end-to-end lifecycle management of... ...efficiently across distributed infrastructure. You will manage GPU orchestration, prompt-tune models, and build advanced AI...Full timeFlexible hours$184k - $287.5k
...computing. An era where our GPU serves as the intelligence behind... ...scale. We seek a Senior ML Engineer to compose and deliver next-... ...proven experience in applied machine learning or AI research.Deep expertise... ...CUDA, Triton, or low-level GPU kernel development for inference...Full time$165.2k - $223.6k
...revolution? At Amazon our vision is to make deep learning pervasive for everyday developers and to... ...workloads.This role is for a software engineer in the Compiler team for AWS Neuron. As... ....- Experience in compiler design for CPU/GPU/Vector engines/ML-accelerators.-...InternshipLocal areaFlexible hours$171k - $231.5k
...the Fintech Risk AI/ML Smart Money Services (SMS) as a Senior Machine Learning Engineer. The SMS team is responsible for detecting and preventing... ...scalable software supporting millions or more usersExperience with GPU acceleration (i.e CUDA and cuDNN)Experience with integrating...WorldwideShift work$224k - $356.5k
NVIDIA is looking for a Machine Learning Engineer to join the GPU accelerated Apache Spark team.Apache Spark is the most popular data processing engine in data centers for running massive scale workloads for ETL, SQL, and ML/DL model training and inference pipelines, spanning...Full time$150k - $230k
...information, visit About the RoleWe are looking for a hands-on Machine Learning Engineer to drive the post-training of our large language models,... ...on tight timelines.Run large-scale training on mid-to-large GPU clusters, applying distributed-training techniques (data parallelism...Full timeLocal areaWork from home$184k - $287.5k
Intelligent machines powered by Artificial Intelligence computers that can learn, reason and interact with people are no longer science fiction. GPU Deep Learning has provided the foundation for machines... ...Senior Perception Engineer to develop and productize NVIDIA...Odd jobFull timeWork experience placementRemote work$152k - $241.5k
...NVIDIA is hiring a Senior Compiler Engineer to join our team driving the next generation of GPU systems programming. We are... ...memory-safe, high-performance GPU kernels in idiomatic Rust.What you’ll be... ...to high-performance CUDA PTX and machine code.Build compiler IRs and...Full timeRemote work- ...advance your career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of GPU-accelerated... ...bottlenecks spanning HBM bandwidth, compute utilization, kernel scheduling, memory allocation, PCIe/Infinity Fabric data...
$184k - $287.5k
We are seeking a self‑motivated senior engineer for the Aerial Omniverse Digital Twin team.... ...design and implementation of a real‑time, GPU‑accelerated propagation engine that... ...and compute‑vs‑bandwidth trade‑offs at the kernel level.Working knowledge of electromagnetic...Full time$152k - $241.5k
...Automotive, VR, Gaming, Deep Learning, and High Performance Computing... ...helped craft as a member of the GPU Foundations Developer Tools... ...Collaborate with developer tools engineers, software library developers,... ...and both user-mode and kernel-mode drivers.Proven knowledge...Full time$207k - $300k
...implement the real-time on-device machine learning foundation that turns raw... ...of outlook models for NPU, GPU, and DSP execution,... ...architecture design, and custom kernel development, in partnership... ...qualifications:Master’s degree or PhD in Engineering, Computer Science, or a...Immediate start$160k - $200k
Santa Clara, CAData Engineering - ML Infrastructure /Full-time /HybridPlusAI... ...for managing large-scale GPU clusters. This role offers... ...integrated with state-of-the-art deep learning frameworks like PyTorch or... ...of what's possible in machine learning infrastructure and contribute...Full time$170k - $240.8k
...and model development initiatives. As a Senior ML Engineer, you will collaborate closely with machine learning engineers, research scientists, and other partners... ..., or similarExperience with distributed computing, GPU computing, and cloud environments (AWS, GCP, Azure)...Full timeLocal areaWork from homeRelocation packageFlexible hours$174.72k - $295.68k
...transportation through cutting-edge R&D in AI, machine learning, and smart connectivity.We are seeking Machine Learning Engineers with strong expertise in generative modeling... ..., such as mixed precision, torch.compile, kernel-level optimization, and multi-node scaling.What...Full time- ...Senior Machine Learning Engineer The AI Engineering group within Intuitive Surgical has an immediate opening in Sunnyvale, CA for a Senior Machine... ...-threaded programming. Strong hands-on experience with GPU accelerated algorithms and implementations. Strong hands-...Immediate start
- ...teams work, discover, and create.We’re seeking a Principal Machine Learning Systems Engineer (P60) to lead technical directions of GenAI Products &... ...HaveBackground in distributed systems, high-performance computing, or GPU optimization.Familiarity with search/GenAI evaluation...Work at officeLocal area
- ...General Motors is seeking a Senior Performance Engineer to join the AV Capacity and Performance Engineering team. You will help develop and optimize large-scale ML infrastructure and GPU platforms for autonomous-vehicle research and deployment. Responsibilities include...
$224k - $356.5k
We are seeking exceptional Senior Machine Learning and Simulation Engineers to join NVIDIA's Autonomous Vehicles (AV) Simulation team! This role requires... ...reliability and performance of training workflows on large GPU clusters through the development of robust monitoring...Full time$272k - $431.25k
NVIDIA is seeking a Senior MLOps Engineering Manager to join our Autonomous Driving organization in Santa Clara, CA. This role offers... ...: Autonomous Vehicles, Robotics, Computer Vision, Deep Learning, or GPU‑accelerated computing.Excellent communication and leadership...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Machine Learning Engineer -- GPU Kernel. Be the first to apply!
- computer vision machine learning engineer Sunnyvale, CA
- machine learning software engineer Sunnyvale, CA
- ai ml engineer Sunnyvale, CA
- senior ml engineer Sunnyvale, CA
- machine learning ai engineer Sunnyvale, CA
- machine learning engineer Sunnyvale, CA
- data engineer machine learning Sunnyvale, CA
- artificial intelligence - machine learning intern Sunnyvale, CA
- machine learning intern Sunnyvale, CA
- internship machine learning Sunnyvale, CA



