Senior Inference Engineer, GPU Kernel Optimization
$184k - $287.5kNVIDIA
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop silicon-measured kernel benchmarking infrastructure, model-level performance projection tooling, and agentic optimization systems that improve GPU kernels at the assembly layer. Our team works closely with compiler, kernel, hardware, and framework organizations across NVIDIA to surface bottlenecks and ship measurable gains. If driving GPU performance at the frontier of LLM inference sounds like your kind of challenge, we'd love to meet you!What you'll be doing:The role drives three interconnected systems, all aimed at accelerating NVIDIA's LLM inference stack. The first is GPU kernel microbenchmarking: measuring competing kernel implementations at real-silicon fidelity across the full configuration space that production LLM deployments demand. The second is end-to-end model performance analysis: connecting performance evidence to model-level serving economics, surfacing high-value optimization opportunities, and producing optimization policies for production inference deployments. The third is agentic kernel optimization: applying AI-driven analysis to diagnose performance gaps, explore optimization opportunities across the kernel ecosystem, and validate findings with rigorous silicon measurements. All three streams converge in close collaboration with compiler, hardware, kernel, and framework teams to deliver upstream improvements and production-grade performance gains.What we need to seeMaster's or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.6+ years of relevant industry experience.Experience building or directing agentic AI systems — code generation, automated optimization, or multi-step reasoning workflows.Strong Python and C++ skills with proven software engineering fundamentals.Hands-on GPU profiling with CUPTI, NSYS, and NCU; proven track record to attribute bottlenecks across kernel execution, compiler decisions, and runtime scheduling.Direct experience with LLM inference frameworks such as TRT-LLM, SGLang, or vLLM and clear understanding of how kernel selection drives model-level throughput and latency.Working knowledge of GPU kernel optimization — CUDA, CUTLASS, Triton, or equivalent — and the ability to read PTX or SASS output.Ways to stand out from the crowdDeep knowledge of SASS/PTX-level kernel analysis, compiler middle-end optimization, or GPU code generation pipelines (LLVM, MLIR, ptxas, or similar).Track record shipping agentic systems end-to-end — tool invent, multi-agent orchestration, and silicon-verified validation — within a performance engineering or kernel optimization context.Active contributions to open-source LLM inference or GPU kernel libraries (FlashInfer, Triton, CUTLASS, or similar).Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until July 31, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, NY, New York; US, WA, SeattleType: Full time
$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop...SeniorFull time- ...career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance... ...HBM bandwidth, compute utilization, kernel scheduling, memory allocation, and... ...framework performance: Profile and optimize inference engines including vLLM, SGLang...Senior
$193.3k - $261.5k
We are looking for a Senior Inference Engineer to own inference for real-time multimodalconversational... ...in• Implement and optimize the inference path for... ...and tune high-performance kernels for critical operations... ...load• Implement multi-GPU inference (tensor parallelism...SeniorInternshipLocal areaFlexible hours$182k - $242k
...cloud for high-performance GPU infrastructure across... ..., and real-time inference. Our stack is engineered for speed, scale, and cost... ...We're looking for a Senior Engineer for CoreWeave'... ...Performance team, focused on kernel authoring and optimization. You will write,...SeniorPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$184k - $287.5k
...are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme... ...architect and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry benchmarks, and scale...SeniorFull time- ...Instinct™ GPUs. You’ll work across kernels, distributed training, and framework... ...is passionate about software engineering and the craft of training performance... ...stability across data, model, and optimizer steps.Optimize multi‑GPU/multi‑node training and communication...Senior
$152k - $241.5k
...the world.NVIDIA is hiring a Senior Compiler Engineer to join our team driving the next generation of GPU systems programming. We are redefining... ...-safe, high-performance GPU kernels in idiomatic Rust.What you’ll... ...code.Build compiler IRs and Optimizers: Work with modern compiler...SeniorFull timeRemote work- ...industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference... ...inference.About The RoleWe’re hiring a Senior Frontend Engineer to own and scale critical parts... ...decisions and know when to optimize vs move fast.Ability to operate...SeniorShift work
$184k - $287.5k
We are now looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior engineers... ...performance analysis and optimization to help us squeeze every... .../software stack from GPU architecture to Deep Learning... ...and multimodal model inference as part of NVIDIA Inference...SeniorFull time- ...industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference... ...inference.About The RoleWe’re hiring a Senior Frontend Engineer to own and scale critical parts... ...decisions and know when to optimize vs move fast.Ability to operate...SeniorFull timeShift work
- ...you will play a pivotal role in optimizing and developing deep learning... ...will be critical in enhancing GPU kernels, deep learning models, and training/inference performance across multi-GPU and... ...compiler technologies and advanced engineering principles to drive continuous...Senior
$184k - $287.5k
We are seeking a self‑motivated senior engineer for the Aerial Omniverse Digital Twin team. This... ...design and implementation of a real‑time, GPU‑accelerated propagation engine that... ...compute‑vs‑bandwidth trade‑offs at the kernel level.Working knowledge of electromagnetic...SeniorFull time- ...ROLEWe are seeking a Principal GenAI Inference Optimization Engineer to join our Models and Applications team... ...AI inference workloads on AMD GPU platforms. You will contribute to optimizing... ...working across multiple layers—from kernels and runtimes to frameworks and...
- ...Overview We are seeking a Senior GPU Systems & Fabric Engineer to serve as the critical... ...deep expertise in Linux kernel internals, GPU... ...plane. Configure and optimize high-performance host networking... ...enable efficient multi-tenant inference workloads and maximize...SeniorRemote jobFull timeLocal area
$184k - $287.5k
...s invention of the GPU 1999 sparked the growth... ...'re looking for a Senior Performance Compiler Engineer to join our team... ...both training and inference. You will be immersed... ...opportunities for optimization.Designing and implementing... ...high-level kernel descriptions (written...SeniorFull timeRemote work$184k - $287.5k
NVIDIA is seeking a Compute Kernel Performance Architect who can... ...workloads with a strong focus on GPU power behavior. In this role,... ..., silicon validation engineers, and software teams to characterize... ...on experience developing and optimizing GPU kernels, including work...SeniorFull time$184k - $287.5k
...transforming every industry. GPU-accelerated deep... ...an exceptional Senior Perception Engineer to help design and productize... ...platforms, including optimization for latency, memory, and... ...training or inference pipelines through custom CUDA kernels or other GPU-accelerated...SeniorFull time$224k - $356.5k
...transforming every industry. GPU-accelerated deep... ...an exceptional Senior Radar Perception Engineer to help design and productize... ..., and layout optimization to support L2-L4 autonomous... ...training or inference pipelines through custom CUDA kernels or other GPU-accelerated...SeniorFull time$152k - $241.5k
.... An era in which our GPU acts as the brains of... ...Deep Learning Compiler Engineer. NVIDIA is hiring software... ...backbone of NVIDIA inference engine, spanning... ...and developing compiler optimization algorithms.Collaborating... ..., such as PyTorch.GPU kernel generation with high performance...SeniorFull timeRemote work$124k - $208.4k
...by millions of people around the world. Come build with us!Role and ResponsibilitiesAs a Senior Engineer, GPU Architect, you will work on the analysis, verification, and optimization of end-to-end system performance for Samsung’s premium mobile GPUs.In this mid-to-senior...SeniorHourly payFull timeRelocation$124k - $208.4k
...the world. Come build with us!Role and ResponsibilitiesAs a Senior Engineer, GPU Modeling Architect, you will work on the design and... ...bottlenecks, and proposing architectural and micro-architectural optimizations to improve GPU performance, power efficiency, and...SeniorHourly payFull timeRelocation$152k - $241.5k
NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined... ....Are you a motivated system software engineer with a deep understanding of device... ...coherence and consistency modelsBackground with kernel mode developmentExperience with Linux...SeniorFull time$184k - $287.5k
...and low-level hardware optimization has never been more... ...optimization, custom kernel development, and cluster... ...systems from a single GPU to supercomputer clusters... ...in both training and inference pipelines.Collaborate... ...Computer Science, Computer Engineering, Electrical...SeniorFull time$124.5k - $272k
...Netskope One platform, its Zero Trust Engine, and the powerful NewEdge network to gain... ...Instagram.Positions are available at Senior Staff and above. Candidates are... ...Machine Learning Scientist, you own the inference and optimization layer that makes AI in agentic workflows...Senior$184k - $287.5k
NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined... ...thermal management software features in Linux Kernel and user spaceCollaborate with power architects, hardware and software engineers on platform power estimation and...SeniorFull time$160k - $253k
...accelerated computing is the engine of artificial intelligence. Our... ...at scale. We are looking for a Senior Technical Marketing Engineer... ...pivotal in showcasing NVIDIA's GPU architecture, server-level platforms... ...and efficiency for AI inference & training.What you’ll be doing...SeniorFull time$184k - $287.5k
NVIDIA seeks a Senior Software Engineer specializing in Deep Learning Inference for our growing team. As a key contributor, you... ...will help design, build, and optimize the GPU-accelerated software that... ...CUTLASS, OAI Triton, NCCL, and CUDA kernels—to implement and optimize...SeniorFull time$136k - $218.5k
...computing. An era in which our GPU acts as the brains of computers... ...the world.NVIDIA is seeking a Senior ASIC Compose Verification Infrastructure and Tools Engineer to advance the systems that qualify... ...priorities, forecast risk, and optimize the systems we operate....SeniorFull time$224k - $356.5k
.... An era in which our GPU acts as the brains of... ...leading open-source LLM inference frameworks — identify... ..., fallback paths, and optimization opportunitiesCharacterize... ...Science, Computer Engineering, Electrical Engineering... ...on experience with GPU kernel development or...SeniorFull timeLocal area- ...industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud... ...Cerebras Wafer-Scale Engine.We are hiring a Software... ...to productionize and optimize our GPU serving stack, working... ...libraries, kernels, drivers, firmware, networking...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Inference Engineer, GPU Kernel Optimization. Be the first to apply!
- senior technical service engineer Santa Clara, CA
- senior director product management Santa Clara, CA
- senior automation controls engineer Santa Clara, CA
- senior grant accountant Santa Clara, CA
- senior tax Santa Clara, CA
- senior firmware engineer Santa Clara, CA
- senior data management analyst Santa Clara, CA
- senior consulting engineer Santa Clara, CA
- sr electrical engineer Santa Clara, CA
- sr marketing manager Santa Clara, CA

