Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Machine Learning Engineer — GPU Kernel

$150k

Jobleads-US

About the Institute of Foundation Models

We are a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.

The Role

The GPU Kernel Engineer will play a role at the forefront of optimizing performance for the machine learning software stacks, especially at training and inference, and support the team to develop new and cutting-edge systems. The ideal candidate will have a strong background in parallel computing, and hands-on experience in system level coding, debug methodologies, and large-scale machine learning experience.

This role focuses on CUDA kernel development and optimization. Distributed training experience is a plus.

Key Responsibilities

  • Understand, analyze, profile, optimize, and provide guidance to the team on deep learning workloads on state-of-the-art hardware and software platforms to improve their efficiency with different levels of optimization
  • Design and implement performance benchmarks and testing methodologies to evaluate application performance
  • Build tools to automate workload analysis, workload optimization, and other critical workflows
  • Triage system issues and identify bottleneck and inefficiencies by analyzing the sources of issues and the impact on hardware, network and propose solutions to enhance GPU utilization
  • Support the team to develop appropriate kernels and systems for new model architectures and algorithms
  • Participate in, or lead design reviews with peers and stakeholders to decide amongst available technologies.
  • Review code developed by other developers and provide feedback to ensure best practices (e.g., style guidelines, checking code in, accuracy, testability, and efficiency).
  • Contribute to existing documentation or educational content and adapt content based on product/program updates and user feedback.
  • Represent MBZUAI at industry conferences and events, showcasing the institution’s cutting-edge HPC anddeep learning capabilities and establishing MBZUAI as a global leader in AI research and innovation.
  • Perform all other duties as reasonably directed by the line manager that are commensurate with thesefunctional objectives.
  • Validate CUDA kernel outputs and gradients against reference implementations, and benchmark representative shapes, dtypes, and model workloads.

Technical Qualifications

Must-Haves:

  • Strong C++ skills and hands-on CUDA kernel development and optimization for deep-learning workloads.
  • Understanding of GPU memory hierarchy, warp/block execution, and compute-memory trade-offs, with demonstrated profiling-driven optimization.
  • Strong Python skills and experience integrating kernels with PyTorch or an equivalent framework, including numerical and gradient validation where needed.

Nice-to-Haves:

  • Experience with Triton, CUTLASS, or PTX/SASS analysis.
  • Experience with multi-node distributed training or inference systems.
  • Experience validating mixed-precision computations, such as BF16 or FP8.

$150,000 - $450,000 a year

The posted salary range represents the company’s good faith estimate of the compensation for this position upon hire. The actual compensation offered may vary within this range depending on individual qualifications, including but not limited to relevant skills, experience, education, certifications, geographic location, and specific business needs.

Benefits Include

  • Comprehensive medical, dental, and vision benefits
  • Bonus
  • 401K Plan
  • Generous paid time off, sick leave and holidays
  • Paid Parental Leave
  • Employee Assistance Program
  • Life insurance and disability
#J-18808-Ljbffr Jobleads-US
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Machine Learning Engineer — GPU Kernel in Sunnyvale, CA vacancy
  • $193.3k - $261.5k

     ...used to accelerate deep learning and GenAI workloads on Amazon’s custom machine learning accelerators, Inferentia...  ....The Acceleration Kernel Library team is at the...  ...-software boundary, our engineers craft high-performance...  ...architectures- Experience with GPU kernel optimization and... 
    Suggested
    Internship
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...NVIDIA Corporation in Santa Clara, CA, seeks a Sr. Inference Engineer to accelerate LLM inference through GPU kernel optimization. You will lead kernel benchmarking, model-level performance analysis, and AI-driven optimization workflows across silicon and software stacks... 
    Suggested

    Jobleads-US

    Santa Clara, CA
    1 day ago
  •  ...speeds; over 10 times faster than GPU-based hyperscale cloud...  ...Core ML team develops novel machine learning algorithms that take advantage...  ...of the Cerebras Wafer-Scale Engine. Our work spans efficient LLM...  ...compilers, runtimes, and low-level kernels to implement new algorithmic... 
    Suggested

    Cerebras Systems

    Sunnyvale, CA
    20 hours ago
  • $150k

     ..., data scientists, and engineers, tackling the most fundamental...  ...computing in deep learning, driving impactful...  ...across multi-node, multi-GPU clusters Own experiment...  ...with large-scale machine learning workloads (strong...  ...performance profiling, kernel fusion, or memory optimization... 
    Suggested
    Visa sponsorship
    Flexible hours

    Institute of Foundation Models

    Sunnyvale, CA
    2 days ago
  •  ...the foundations to democratize AI and Machine Learning for Atlassian’s teams, customers, and...  ...About This RoleAs a Senior ML System Engineer on the AI & ML Platform’s Inference team...  ...scaling) to deep low-level optimizations (GPU kernels, quantization, speculative decoding).... 
    Work at office
    Local area

    Atlassian

    Mountain View, CA
    1 day ago
  • $197.5k - $272k

     ...Edge. We are looking for a great Staff Machine Learning Engineer to join our seasoned AI team and lead...  ...process unstructured application logs, kernel traces, and multi-modalities.Integrate...  ...-grade models for execution on CPU/GPU-bound targets or embedded NPUs.Apply quantization... 
    Work at office
    Worldwide
    Flexible hours
    Shift work
    3 days per week

    Sonatus

    Sunnyvale, CA
    2 days ago
  • $278.1k - $347.6k

     ...entirely within that runtime. As our Principal Engineer for On-Device AI Inference & Systems, you...  ..., through export, optimization, and kernel-level tuning, to a shipped feature...  ...specific kernel tuning across NPU, mobile GPU, and desktop/laptop GPU. Make authoritative... 
    Work at office
    Worldwide
    Relocation package

    Unity

    Mountain View, CA
    2 days ago
  • $150k

     ..., data scientists, and engineers, tackling the most fundamental...  ...computing in deep learning, driving impactful...  ...across multi-node, multi-GPU clusters Own experiment...  ...with large-scale machine learning workloads (strong...  ...performance profiling, kernel fusion, or memory optimization... 
    Flexible hours

    Jobleads-US

    Sunnyvale, CA
    10 hours ago
  • $184k - $287.5k

     ...computing. An era where our GPU serves as the intelligence behind...  ...with Product, Program, Engineering, and Data Procurement teams...  ...proven experience in applied machine learning or AI research. ~ Deep expertise...  ..., Triton, or low-level GPU kernel development for inference... 

    Jobleads-US

    Santa Clara, CA
    10 hours ago
  • $165.6k

     ...degree in computer science or an equivalent background Responsibilities: Design and build high-performance compute kernels for machine learning operations using the Neuron architecture and programming models Evaluate and improve kernel-level performance across... 
    Full time
    Internship

    Annapurna Labs (U.S.) Inc.

    Cupertino, CA
    4 days ago
  • $119.25k - $150.85k

     ...Inference Solutions team deploys machine learning models from training...  ...Escalade IQ, and we’re hiring engineers to help deliver the next generation...  ...with our sister teams (kernels, compiler, reduced precision...  ...ML systems, ML compilers, GPU programming (CUDA, OpenAI Triton... 
    Full time
    Internship
    Local area
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    4 days ago
  • $2,000 per month

     ...product could never be achieved on a typical GPU Implement diffusion models on Sohu to...  ...with Rust Familiarity with GPU kernels, the CUDA compilation stack and related tools...  ...person team in Cupertino, and greatly value engineering skills. We do not have boundaries between... 
    Work at office
    Relocation package

    ETCHED LLC

    Cupertino, CA
    2 days ago
  •  ...Meta's Santa Clara-based MTIA Software Team is hiring a Software Engineer, Systems ML specializing in compilers and kernels. You will help develop the AI compiler stack, contribute to PyTorch core components, and optimize high-performance kernels for next-generation hardware... 

    Jobleads-US

    Santa Clara, CA
    1 day ago
  •  ...across image, video, and world-model workloads. You will work with a founding team on kernels, runtimes, and distributed engines that power production-scale ML stacks. You’ll optimize GPU performance, profile bottlenecks with Nsight, and implement low-level CUDA and... 

    Jobleads-US

    Menlo Park, CA
    20 hours ago
  • $152k - $241.5k

    NVIDIA is looking for a talented Machine Learning Engineer to drive the development, evaluation, deployment and end-to-end lifecycle management of...  ...efficiently across distributed infrastructure. You will manage GPU orchestration, prompt-tune models, and build advanced AI... 
    Full time
    Flexible hours

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...THE ROLE: We are looking for a talented engineer to join our team: developing heterogeneous...  ...: Design, develop, and optimize GPU/CPU software for computer vision, image...  ...Experience in video codecs, image processing and machine learning frameworksFamiliarity with computer... 

    AMD

    Santa Clara, CA
    1 day ago
  • $224k - $356.5k

    NVIDIA is looking for a Machine Learning Engineer to join the GPU accelerated Apache Spark team.Apache Spark is the most popular data processing engine in data centers for running massive scale workloads for ETL, SQL, and ML/DL model training and inference pipelines, spanning... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $171k - $231.5k

     ...the Fintech Risk AI/ML Smart Money Services (SMS) as a Senior Machine Learning Engineer. The SMS team is responsible for detecting and preventing...  ...scalable software supporting millions or more usersExperience with GPU acceleration (i.e CUDA and cuDNN)Experience with integrating... 
    Worldwide
    Shift work

    Intuit

    Mountain View, CA
    4 days ago
  • $165.2k - $223.6k

     ...revolution? At Amazon our vision is to make deep learning pervasive for everyday developers and to...  ...workloads.This role is for a software engineer in the Compiler team for AWS Neuron. As...  ....- Experience in compiler design for CPU/GPU/Vector engines/ML-accelerators.-... 
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  •  ...patients worldwide.We’re a team of engineers, clinicians, and innovators...  ..., CA for a SeniorMachine Learning Engineer with focus on Computer...  ...computer vision and machine learning applications; or PhD...  ...Strong hands-on experience with GPU accelerated algorithms and implementations... 
    Local area
    Immediate start
    Worldwide
    Flexible hours

    Intuitive Surgical

    Sunnyvale, CA
    1 day ago
  • $143k - $286k

     ...WalmartBusiness Segment: Home OfficeRole summary: The Senior Machine Learning Engineer will lead the design, development, and deployment of...  ...pipeline development.Familiarity with large language modeling and GPU optimization techniques. At Walmart, we offer competitive pay... 
    Full time
    Temporary work
    Part time
    Work experience placement

    Walmart

    Sunnyvale, CA
    1 day ago
  • $150k - $230k

     ...information, visit About the RoleWe are looking for a hands-on Machine Learning Engineer to drive the post-training of our large language models,...  ...on tight timelines.Run large-scale training on mid-to-large GPU clusters, applying distributed-training techniques (data parallelism... 
    Full time
    Local area
    Work from home

    News Break

    Mountain View, CA
    2 days ago
  • $184k - $287.5k

    Intelligent machines powered by Artificial Intelligence computers that can learn, reason and interact with people are no longer science fiction. GPU Deep Learning has provided the foundation for machines...  ...Senior Perception Engineer to develop and productize NVIDIA... 
    Odd job
    Full time
    Work experience placement
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $207k - $300k

     ...implement the real-time on-device machine learning foundation that turns raw...  ...of outlook models for NPU, GPU, and DSP execution,...  ...architecture design, and custom kernel development, in partnership...  ...qualifications:Master’s degree or PhD in Engineering, Computer Science, or a... 
    Immediate start

    Google

    Mountain View, CA
    2 days ago
  • $184k - $287.5k

    We are seeking a self‑motivated senior engineer for the Aerial Omniverse Digital Twin team....  ...design and implementation of a real‑time, GPU‑accelerated propagation engine that...  ...and compute‑vs‑bandwidth trade‑offs at the kernel level.Working knowledge of electromagnetic... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

     ...Automotive, VR, Gaming, Deep Learning, and High Performance Computing...  ...helped craft as a member of the GPU Foundations Developer Tools...  ...Collaborate with developer tools engineers, software library developers,...  ...and both user-mode and kernel-mode drivers.Proven knowledge... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...advance your career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of GPU-accelerated...  ...bottlenecks spanning HBM bandwidth, compute utilization, kernel scheduling, memory allocation, PCIe/Infinity Fabric data... 

    AMD

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...NVIDIA is hiring a Senior Compiler Engineer to join our team driving the next generation of GPU systems programming. We are...  ...memory-safe, high-performance GPU kernels in idiomatic Rust.What you’ll be...  ...to high-performance CUDA PTX and machine code.Build compiler IRs and... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    20 hours ago
  • $160k - $200k

    Santa Clara, CAData Engineering - ML Infrastructure /Full-time /HybridPlusAI...  ...for managing large-scale GPU clusters. This role offers...  ...integrated with state-of-the-art deep learning frameworks like PyTorch or...  ...of what's possible in machine learning infrastructure and contribute... 
    Full time

    Plus.ai

    Santa Clara, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Machine Learning Engineer — GPU Kernel. Be the first to apply!