Senior Inference Engineer, GPU Kernel Optimization
$184k - $287.5kNVIDIA
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop silicon-measured kernel benchmarking infrastructure, model-level performance projection tooling, and agentic optimization systems that improve GPU kernels at the assembly layer. Our team works closely with compiler, kernel, hardware, and framework organizations across NVIDIA to surface bottlenecks and ship measurable gains. If driving GPU performance at the frontier of LLM inference sounds like your kind of challenge, we'd love to meet you!What you'll be doing:The role drives three interconnected systems, all aimed at accelerating NVIDIA's LLM inference stack. The first is GPU kernel microbenchmarking: measuring competing kernel implementations at real-silicon fidelity across the full configuration space that production LLM deployments demand. The second is end-to-end model performance analysis: connecting performance evidence to model-level serving economics, surfacing high-value optimization opportunities, and producing optimization policies for production inference deployments. The third is agentic kernel optimization: applying AI-driven analysis to diagnose performance gaps, explore optimization opportunities across the kernel ecosystem, and validate findings with rigorous silicon measurements. All three streams converge in close collaboration with compiler, hardware, kernel, and framework teams to deliver upstream improvements and production-grade performance gains.What we need to seeMaster's or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.6+ years of relevant industry experience.Experience building or directing agentic AI systems — code generation, automated optimization, or multi-step reasoning workflows.Strong Python and C++ skills with proven software engineering fundamentals.Hands-on GPU profiling with CUPTI, NSYS, and NCU; proven track record to attribute bottlenecks across kernel execution, compiler decisions, and runtime scheduling.Direct experience with LLM inference frameworks such as TRT-LLM, SGLang, or vLLM and clear understanding of how kernel selection drives model-level throughput and latency.Working knowledge of GPU kernel optimization — CUDA, CUTLASS, Triton, or equivalent — and the ability to read PTX or SASS output.Ways to stand out from the crowdDeep knowledge of SASS/PTX-level kernel analysis, compiler middle-end optimization, or GPU code generation pipelines (LLVM, MLIR, ptxas, or similar).Track record shipping agentic systems end-to-end — tool invent, multi-agent orchestration, and silicon-verified validation — within a performance engineering or kernel optimization context.Active contributions to open-source LLM inference or GPU kernel libraries (FlashInfer, Triton, CUTLASS, or similar).Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until July 31, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, NY, New York; US, WA, SeattleType: Full time
- ...career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance... ...HBM bandwidth, compute utilization, kernel scheduling, memory allocation, and... ...framework performance: Profile and optimize inference engines including vLLM, SGLang...Senior
$184k - $287.5k
We are now looking for a Senior Formal Verification Engineer for GPU Kernels! Modern AI performance relies on highly optimized GPU kernels — performance-critical code where bugs can be hard to catch and expensive to miss. NVIDIA's Deep Learning Safety Team is hiring engineers...SeniorFull timeWork experience placement- ...industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud... ...About The RoleWe are hiring a Senior Performance Engineer to join our Product team. You... ...SGLang, TensorRT-LLM), GPU kernel-level optimization toolchains (CUDA, Triton), and...SeniorContract workShift work
$195.2k - $361.2k
...people it servesRole SummaryMake models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments — GPU/iGPUs, Vulkan backends — not datacenter H100 environment, mostly PC/edge. KV cache, batching...SeniorFull timeInternshipLocal areaImmediate startShift work$182k - $242k
...cloud for high-performance GPU infrastructure across... ..., and real-time inference. Our stack is engineered for speed, scale, and cost... ...We're looking for a Senior Engineer for CoreWeave'... ...Performance team, focused on kernel authoring and optimization. You will write,...SeniorPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$184k - $287.5k
...are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme... ...architect and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry benchmarks, and scale...SeniorFull time- ...Instinct™ GPUs. You’ll work across kernels, distributed training, and framework... ...is passionate about software engineering and the craft of training performance... ...stability across data, model, and optimizer steps.Optimize multi‑GPU/multi‑node training and communication...Senior
$152k - $241.5k
...the world.NVIDIA is hiring a Senior Compiler Engineer to join our team driving the next generation of GPU systems programming. We are redefining... ...-safe, high-performance GPU kernels in idiomatic Rust.What you’ll... ...code.Build compiler IRs and Optimizers: Work with modern compiler...SeniorFull timeRemote work$184k - $287.5k
We are now looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior engineers... ...performance analysis and optimization to help us squeeze every... .../software stack from GPU architecture to Deep Learning... ...and multimodal model inference as part of NVIDIA Inference...SeniorFull time- ...you will play a pivotal role in optimizing and developing deep learning... ...will be critical in enhancing GPU kernels, deep learning models, and training/inference performance across multi-GPU and... ...compiler technologies and advanced engineering principles to drive continuous...Senior
$184k - $287.5k
We are seeking a self‑motivated senior engineer for the Aerial Omniverse Digital Twin team. This... ...design and implementation of a real‑time, GPU‑accelerated propagation engine that... ...compute‑vs‑bandwidth trade‑offs at the kernel level.Working knowledge of electromagnetic...SeniorFull time$184k - $287.5k
...AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars... ...impact on the worldWe are looking for an experienced Compiler Optimization Engineer for an exciting role in our Compute Compiler Team. We deliver...SeniorFull time$184k - $287.5k
...longer science fiction. GPU Deep Learning has... ...looking for an extraordinary Senior Perception Engineer to develop and... ...through KPI building and optimization. This includes careful... ...ability to implement CUDA kernels as part of training or inference pipelines.Your base salary...SeniorFull timeWork experience placement- ...infrastructure to predict, explain, and improve AMD GPU kernel performance across current and future... ...validation, and architecture-aware optimization. Hardware-software co-design experience... ...compilers, and faster kernels.Mentor engineers, write design docs, lead technical...SeniorLocal area
- ...ROLEWe are seeking a Principal GenAI Inference Optimization Engineer to join our Models and Applications team... ...AI inference workloads on AMD GPU platforms. You will contribute to optimizing... ...working across multiple layers—from kernels and runtimes to frameworks and...
$184k - $287.5k
...s invention of the GPU 1999 sparked the growth... ...'re looking for a Senior Performance Compiler Engineer to join our team... ...both training and inference. You will be immersed... ...opportunities for optimization.Designing and implementing... ...high-level kernel descriptions (written...SeniorFull timeRemote work$184k - $287.5k
...transforming every industry. GPU-accelerated deep... ...an exceptional Senior Perception Engineer to help design and productize... ...platforms, including optimization for latency, memory, and... ...training or inference pipelines through custom CUDA kernels or other GPU-accelerated...SeniorFull time$152k - $241.5k
.... An era in which our GPU acts as the brains of... ...Deep Learning Compiler Engineer. NVIDIA is hiring software... ...backbone of NVIDIA inference engine, spanning... ...and developing compiler optimization algorithms.Collaborating... ..., such as PyTorch.GPU kernel generation with high performance...SeniorFull timeRemote work$139k - $208.4k
...the world. Come build with us!Role and ResponsibilitiesAs a Senior Engineer GPU RTL Design Engineer, you will contribute to the design and... ...RTL design, front-end implementation, power and performance optimization, firmware enablement, and the infrastructure needed to...SeniorHourly payFull timeWorldwideRelocation$224k - $356.5k
...transforming every industry. GPU-accelerated deep... ...an exceptional Senior Radar Perception Engineer to help design and productize... ..., and layout optimization to support L2-L4 autonomous... ...training or inference pipelines through custom CUDA kernels or other GPU-accelerated...SeniorFull time$184k - $287.5k
NVIDIA is seeking a Compute Kernel Performance Architect who can... ...workloads with a strong focus on GPU power behavior. In this role,... ..., silicon validation engineers, and software teams to characterize... ...on experience developing and optimizing GPU kernels, including work...SeniorFull time$124k - $208.4k
...by millions of people around the world. Come build with us!Role and ResponsibilitiesAs a Senior Engineer, GPU Architect, you will work on the analysis, verification, and optimization of end-to-end system performance for Samsung’s premium mobile GPUs.In this mid-to-senior...SeniorHourly payFull timeRelocation$124k - $208.4k
...the world. Come build with us!Role and ResponsibilitiesAs a Senior Engineer, GPU Modeling Architect, you will work on the design and... ...bottlenecks, and proposing architectural and micro-architectural optimizations to improve GPU performance, power efficiency, and...SeniorHourly payFull timeRelocation$184k - $287.5k
...and low-level hardware optimization has never been more... ...optimization, custom kernel development, and cluster... ...systems from a single GPU to supercomputer clusters... ...in both training and inference pipelines.Collaborate... ...Computer Science, Computer Engineering, Electrical...SeniorFull time$182.5k - $260.5k
...Netskope One platform, its Zero Trust Engine, and the powerful NewEdge network to gain... ...Instagram.Positions are available at Senior Staff and above. Candidates are... ...Machine Learning Scientist, you own the inference and optimization layer that makes AI in agentic workflows...Senior$152k - $241.5k
...applications? Join NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of technology, enabling high-performance AI inference solutions for automotive safety and... ...requirementsContribute to performance optimization and benchmarking efforts for...SeniorFull time$160k - $253k
...accelerated computing is the engine of artificial intelligence. Our... ...at scale. We are looking for a Senior Technical Marketing Engineer... ...pivotal in showcasing NVIDIA's GPU architecture, server-level platforms... ...and efficiency for AI inference & training.What you’ll be doing...SeniorFull time$152k - $241.5k
...built. We are seeking a Senior Software Engineer - AI Inference to advance open‑source LLM... ...Contribute features, fixes, and optimizations upstream to vLLM/SGLang:... ...to C++/CUDA kernels-using data to guide optimization... ...optimization work.Improve multi‑GPU inference performance and...SeniorFull timeRemote work$152k - $241.5k
We are now looking for a Senior Software Engineer for Deep Learning Inference! Would you like to make... ...to build graph parsers, optimizers, and tools for effective... ...deep learning experts, GPU architects and DevOps engineers... ...as Background in GPU kernel programming using CUDA...SeniorFull time$152k - $241.5k
NVIDIA's invention of the GPU 1999 sparked the growth... ...top-tier AI Compiler Engineers to drive innovation... ...development focusing on kernel generation and computational graph optimizations for next-generation NVIDIA... ...for AI workloads (both inference and training) and...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Inference Engineer, GPU Kernel Optimization. Be the first to apply!
- senior manager tax Santa Clara, CA
- senior devops Santa Clara, CA
- senior director digital marketing Santa Clara, CA
- senior international accountant Santa Clara, CA
- senior vmware engineer Santa Clara, CA
- sr marketing manager Santa Clara, CA
- sr technical product manager Santa Clara, CA
- senior resident engineer Santa Clara, CA
- senior performance engineer Santa Clara, CA
- senior storage engineer Santa Clara, CA
