Principal Engineer, CUDA UMD - GPU Kernel Scheduling
$272k - $431.25kNVIDIA
NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. We're looking to grow our company, and form teams with the smartest people in the world. Join us at the forefront of technological advancement.Are you a motivated system software engineer with a deep understanding of device drivers who has phenomenal C/C++ skills? If so, this role might be for you. We are looking for a seasoned software professional to work on the CUDA Driver, a core component of our platform for accelerating general purpose computation on the GPU. You will be an integral part of a team that delivers features and improvements to better realize the potential of NVIDIA hardware for a growing range of computational workloads, ranging from deep learning, scientific computation, data science and self-driving cars to video games and virtual reality.What you'll be doing:As a member of our team, you will use your design abilities, coding expertise, and creativity to deliver the best compute platform in the world. You will craft elegant solutions to exciting problems and shape the future direction of CUDA as you collaborate with your peers across NVIDIA.Evangelize, architect, and implement new featuresCoordinate and drive development efforts across multiple teamsHelp define forward-looking improvements to the CUDA APIs and programming modelExtend important CUDA programming models and functionality such as CUDA GraphsExplore ways to use Graphs to improve the scheduling of AI/ML workloads on our GPUS to be more efficient and faster. Write effective, maintainable, and well-tested codeDevelop code for multiple operating systemsWhat we need to see:BS or MS degree in Computer Science, Electrical Engineering or related field (or equivalent experience)Strong C and C++ programming skillsMinimum of 15+ years of related development experience (multiple positions for varying experience levels open)Experience driving projects across multiple teamsExperience working with large codebasesBackground with operating system interfaces for threads, process control, and virtual memoryExperience writing and debugging multithreaded programsGood written communication as well as presentation skillsWays to stand out from the crowd:Prior experience with parallel computing - preferably writing CUDA Programs or Libraries that use CUDAUnderstanding of system level architecture, such as interconnects, memory hierarchy, interrupts, and memory-mapped IOKnowledge of memory coherence and consistency modelsBackground with kernel mode developmentExperience with Linux Systems Software development as well as experience maintaining and extending programming models or higher-level language support for similar environmentsYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until September 18, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time
$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push... ..., compiler decisions, and runtime scheduling.Direct experience with LLM inference frameworks... ...of GPU kernel optimization — CUDA, CUTLASS, Triton, or equivalent — and...SuggestedFull time- ...MBZUAI is seeking a GPU Kernel Engineer to optimize machine learning software stacks for training and inference on state-of-the-art hardware. You will work on CUDA kernel development, profiling, and performance tuning, with opportunities to lead design reviews and contribute...Suggested
- ...THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance... ...rocprof-sys, Omniperf) and NVIDIA (CUDA, Nsight Systems/Compute, DCGM) GPUs... ...HBM bandwidth, compute utilization, kernel scheduling, memory allocation, PCIe/Infinity Fabric...Suggested
$272k - $431.25k
...We are seeking software engineers to work on next-generation high-speed interconnect technologies... ...demanding high-speed IO applications a GPU or high-performance computing server... ...diagnostic software using debug CUDA/kernel driver features This job will require an...Suggested- ...ROLEWe are seeking a Principal GenAI Inference Optimization Engineer to join our Models... ...workloads on AMD GPU platforms. You will... ...multiple layers—from kernels and runtimes to... ...memory bandwidth, scheduling).- Contribute to cross... ...systems language (C++/CUDA/HIP).- Experience...Suggested
- ...NVIDIA Corporation is seeking a Senior Compiler Engineer to advance Rust-to-GPU tooling, with focus on safe, high-performance code. You will design... ...IR frameworks, and JIT components that translate Rust into CUDA PTX and native GPU code. You will collaborate across teams...
- ...NVIDIA Corporation in Santa Clara, CA, seeks a Sr. Inference Engineer to accelerate LLM inference through GPU kernel optimization. You will lead kernel benchmarking, model-level performance analysis, and AI-driven optimization workflows across silicon and software stacks...
$272k - $431.25k
...transforming every industry. GPU-accelerated deep learning... ...are seeking an exceptional Principal Perception Engineer to lead the design and productization... ...pipelines.Experience with CUDA development and optimizing... ...through custom CUDA kernels or other GPU-accelerated...Full time$272k - $431.25k
...era of computing. An era in which our GPU acts as the brains of computers,... ...a lasting impact on the world. As a Principal Linux Kernel Engineer, you will lead the effort to develop... ...subsystems including memory management, scheduling, virtualization, PCIe, ACPI, RAS, CXL...- ...We are seeking a Senior GPU Systems & Fabric Engineer to serve as the critical bridge... ...deep expertise in Linux kernel internals, GPU architectures... ...drivers, and runtime libraries (CUDA, NCCL) to resolve... ...workloads. Collaborate with the Scheduling and Storage engineering...Full timeLocal area
$152k - $241.5k
...NVIDIA is hiring a Senior Compiler Engineer to join our team driving the next generation of GPU systems programming. We are... ...tooling of Rust to native GPU and CUDA development. On this team, you will... ...-safe, high-performance GPU kernels in idiomatic Rust.What you’ll be...Full timeRemote work$184k - $287.5k
...seeking a self‑motivated senior engineer for the Aerial Omniverse... ...implementation of a real‑time, GPU‑accelerated propagation engine... ...experience.Hands‑on proficiency with CUDA and at least one GPU ray‑... ...vs‑bandwidth trade‑offs at the kernel level.Working knowledge of electromagnetic...Full time$152k - $241.5k
...era of computing. An era in which our GPU acts as the brains of computers, robots... ...impact on the world.We are hiring software engineers for the CUDA Tile team. NVIDIA GPUs are at the... ...optimize the performance of tile-based kernels to ensure they execute efficiently across...Full timeRemote work$152k - $241.5k
...helped craft as a member of the GPU Foundations Developer Tools... ...Collaborate with developer tools engineers, software library developers,... ...and both user-mode and kernel-mode drivers.Proven knowledge... ...monitor hardware.Experience in CUDA or GPU programming models.Prior...Full time$207k - $340k
...approval. We’re hiring a Principal Staff Software Engineer to lead LinkedIn’s GPU-Based Retrieval Platform,... ...distributed serving, GPU scheduling, memory efficiency, batching, and kernel-level optimization. You will... ...Optimize GPU performance across CUDA, Triton, memory...For contractorsWork at officeRemote workWork from homeFlexible hours- ...collaborating across compiler, libraries, and research groups. You will oversee performance tuning for large models (LLM, multimodal, generative AI), using CUDA, Triton, and multi-GPU communications. Equity and comprehensive benefits are offered. #J-18808-Ljbffr Jobleads-US
$206k - $333k
...this role We're looking for a Principal Engineer to be the technical lead of... ...tracks with NVIDIA (CUDA, cuDNN, TensorRT/TensorRT-LLM... ...6/FP8/FP4), batch sizes, and GPU types. Maintain a corpus of representative... ...s controllers/operators to schedule benchmarks at scale;...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$278.1k - $417.1k
...that runtime. As our Principal Engineer for On-Device AI Inference... ..., optimization, and kernel-level tuning, to a... ...tuning across NPU, mobile GPU, and desktop/laptop... ...SPIR-V compute, D3D12, CUDA); profile with browser... ...game engine: real-time scheduling, threading, memory pooling...Work at officeWorldwideRelocation package$220.92k - $311.89k
Job Details:Job Description: The Role and Impact As a GPU Platform Hardware Design Engineer, you will play a pivotal role in designing and developing high-quality GPU hardware platforms that drive innovation in high-performance computing, graphics, and visualization technologies...Full timeLocal areaImmediate startShift work- ...NVIDIA Corporation is seeking software engineers to design, implement, and optimize... ...performance sparse linear algebra kernels for cuSPARSE and cuDSS. You’ll work on GPU-accelerated libraries,... ...feature and performance goals across CUDA-enabled architectures. Ideal candidates...
- ...inference speeds; over 10 times faster than GPU-based hyperscale cloud inference... ...speed inference. About the Role As a Kernel Engineer at Cerebras, you will develop high-performance... ...-level programming, assembly language, CUDA, OpenCL, or a domain-specific language....Internship
$140k - $224.25k
NVIDIA's invention of the GPU 1999 sparked the growth of the PC... ...are looking for a Verification Engineer, Compute Performance to join our... ...and be productive under tight schedules, and have strong analytical... ...programming and/or testing in C/C++/CUDA as well as scripting languages...Full timeWork experience placement- ...motivated PhD AI Systems & GPU Performance Engineering Intern to join our team.... ...bandwidth, communication, and kernel execution. Your... ...optimization, operator fusion, scheduling strategies, mixed precision... ...programming technologies such as CUDA, HIP, Triton, or similar...Full timeSummer workInternshipSummer internshipWorldwide
$140k - $224.25k
...creative, and hands-on software engineer with a test to failure... ...optimize the testing workflows in GPU domain.Write maintainable, reliable... ...create a realistic delivery schedule and work on challenging... ...DLSS, Frame Generation, Reflex, CUDA, G-Sync, etc.The ability to collaborate...Full time$133.5k - $272k
...trust the Netskope One platform, its Zero Trust Engine, and the powerful NewEdge network to gain... ...computer architecture - multi-threading, CPU scheduling, memory management with practical knowledge of Linux at system/kernel levelA strong understanding of available AI tools...$133.5k - $272k
...trust the Netskope One platform, its Zero Trust Engine, and the powerful NewEdge network to gain... ...computer architecture - multi-threading, CPU scheduling, memory management with practical knowledge of Linux at system/kernel levelHands-on experience with Docker and Kubernetes...$219k - $351k
..., customers, partners, and communities.Principal Engineer, Architecture & Performance Research Engineer... ...-V, ARM, or x86 cores; experience with GPU/NPU vector or accelerator architectures... ...-end: Out-of-order execution, scheduling, RVV/SIMD, vector execution, or matrix/...Work at officeFlexible hoursShift work$248k - $396.75k
...accelerated computing, inventing the GPU and pioneering the... ...our time.NVIDIA's IT Storage Engineering team architects, designs, deploys... ...performance, including kernel-level concepts around I/O subsystems... ...with HPC job schedulers such as IBM LSF or Slurm — understanding...Full timeShift work$200k - $260k
...for talented individuals to join its fast-growing teams.As a Principal Engineer on our Mapping and Localization team, you will design and... ...experience with LOAM, lego-LOAM, ORB-SLAM, or VINS.Experience with CUDA/GPU programming or optimization for ARM-based embedded platforms...Odd jobFull timeShift work- ...pioneering initiatives at the intersection of CUDA and Deep Learning Systems. You will tackle... ...optimization across DL workloads, high-performance kernels, and distributed AI systems, delivering scalable performance from a single GPU to large clusters. Join a research-...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Engineer, CUDA UMD - GPU Kernel Scheduling. Be the first to apply!
- chief engineer Santa Clara, CA
- senior chief engineer Santa Clara, CA
- principal infrastructure engineer Santa Clara, CA
- principal developer Santa Clara, CA
- principal reliability engineer Santa Clara, CA
- general engineer Santa Clara, CA
- director software engineering Santa Clara, CA
- engineering director Santa Clara, CA
- director data engineering Santa Clara, CA
- chief design engineer Santa Clara, CA




