Senior Inference Engineer, GPU Kernel Optimization
$184k - $287.5kNVIDIA
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop silicon-measured kernel benchmarking infrastructure, model-level performance projection tooling, and agentic optimization systems that improve GPU kernels at the assembly layer. Our team works closely with compiler, kernel, hardware, and framework organizations across NVIDIA to surface bottlenecks and ship measurable gains. If driving GPU performance at the frontier of LLM inference sounds like your kind of challenge, we'd love to meet you!What you'll be doing:The role drives three interconnected systems, all aimed at accelerating NVIDIA's LLM inference stack. The first is GPU kernel microbenchmarking: measuring competing kernel implementations at real-silicon fidelity across the full configuration space that production LLM deployments demand. The second is end-to-end model performance analysis: connecting performance evidence to model-level serving economics, surfacing high-value optimization opportunities, and producing optimization policies for production inference deployments. The third is agentic kernel optimization: applying AI-driven analysis to diagnose performance gaps, explore optimization opportunities across the kernel ecosystem, and validate findings with rigorous silicon measurements. All three streams converge in close collaboration with compiler, hardware, kernel, and framework teams to deliver upstream improvements and production-grade performance gains.What we need to seeMaster's or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.6+ years of relevant industry experience.Experience building or directing agentic AI systems — code generation, automated optimization, or multi-step reasoning workflows.Strong Python and C++ skills with proven software engineering fundamentals.Hands-on GPU profiling with CUPTI, NSYS, and NCU; proven track record to attribute bottlenecks across kernel execution, compiler decisions, and runtime scheduling.Direct experience with LLM inference frameworks such as TRT-LLM, SGLang, or vLLM and clear understanding of how kernel selection drives model-level throughput and latency.Working knowledge of GPU kernel optimization — CUDA, CUTLASS, Triton, or equivalent — and the ability to read PTX or SASS output.Ways to stand out from the crowdDeep knowledge of SASS/PTX-level kernel analysis, compiler middle-end optimization, or GPU code generation pipelines (LLVM, MLIR, ptxas, or similar).Track record shipping agentic systems end-to-end — tool invent, multi-agent orchestration, and silicon-verified validation — within a performance engineering or kernel optimization context.Active contributions to open-source LLM inference or GPU kernel libraries (FlashInfer, Triton, CUTLASS, or similar).Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until July 31, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, NY, New York; US, WA, SeattleType: Full time
$182k - $242k
...cloud for high-performance GPU infrastructure across... ..., and real-time inference. Our stack is engineered for speed, scale, and cost... ...We're looking for a Senior Engineer for CoreWeave'... ...Performance team, focused on kernel authoring and optimization. You will write,...SeniorPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$226k - $307k
...Machine Learning and System Optimization Engineer, you will orchestrate and allocate... ...allow for more efficient inference by sharing various parts of... ...models, write custom CUDA kernels, and build highly concurrent... ...distribute system resources (CPU/GPU/interconnect) to various...SeniorFull timeTemporary workRelocation package$152k - $241.5k
...the world.NVIDIA is hiring a Senior Compiler Engineer to join our team driving the next generation of GPU systems programming. We are redefining... ...-safe, high-performance GPU kernels in idiomatic Rust.What you’ll... ...code.Build compiler IRs and Optimizers: Work with modern compiler...SeniorFull timeRemote work$184k - $287.5k
...s invention of the GPU 1999 sparked the growth... ...'re looking for a Senior Performance Compiler Engineer to join our team... ...both training and inference. You will be immersed... ...opportunities for optimization.Designing and implementing... ...high-level kernel descriptions (written...SeniorFull timeRemote work- ...intelligence. As a Model Optimization & Deployment Engineer, you will focus on bringing highly... ...ML models, write custom CUDA kernels, and build highly concurrent inference code to ensure real-time,... ...runtime efficiency optimization for GPU clusters. Experience with...SeniorTemporary workRelocation package
$224k - $356.5k
.... An era in which our GPU acts as the brains of... ...leading open-source LLM inference frameworks — identify... ..., fallback paths, and optimization opportunitiesCharacterize... ...Science, Computer Engineering, Electrical Engineering... ...on experience with GPU kernel development or...SeniorFull timeLocal area$164k - $313.3k
...ART is seeking a Senior Machine Learning (... ...Systems & Efficiency Engineer to join our R&D... ...ready improvements in inference performance,... ..., and performance optimization. You will work closely... ..., improve GPU utilization, and build... ...runtime performance.Kernel Development & System...SeniorFull timeTemporary workLocal areaWorldwide$184k - $287.5k
NVIDIA is seeking an NCX Senior Engineer to join our DSX team, collaborating closely with... ...platforms, including distributed training, inference optimization, and MLOps pipelines constructed on... ...using Kubernetes, containers, and GPU scheduling systems aligned to NCP builds...SeniorFull timeRemote work$242k - $389k
...to drive our ML Performance Optimization initiatives and make our ML models... ...a team of strong software engineers and act as a force multiplier... ...cutting-edge ML Training OR Inference performance optimization techniques... ...training.Experience with GPU-accelerated inference using...SeniorFull timeRemote work$217k - $307k
...achieve maximum throughput at the most optimal power levels. The Software... ...meet the required specifications. As a GPU performance software engineer within the Software Performance team... ...experience debugging/optimizing GPU kernels using tools like Nsight.Strong knowledge...SeniorFull timeTemporary workRelocation package- ...feedback (RLHF), and continual learning, to optimize AI customer service responses with... ...compression, quantization, and efficient inference techniques to ensure the AI customer... ...with majors in computer science, computer engineering, statistics, applied mathematics, data...SeniorWork experience placement
$315k
...committed researchers, engineers, policy experts, and... ...breakthrough innovations in GPU performance and... ...developing cutting-edge optimizations that directly enable new... ...dramatically improve inference efficiency. Working... ...from custom kernel development to distributed...Work at officeVisa sponsorshipFlexible hours$280k
...committed researchers, engineers, policy experts, and business... ...developing systems that optimize the throughput and... ...large-scale ML systems * GPU/Accelerator programming... ...models * Implement GPU kernels to adapt our models to low-precision inference * Write a custom load-...Full timeWork at officeVisa sponsorshipFlexible hours$152k - $241.5k
...AI workloads, and we are looking for an engineer focused on performance validation, analysis... ...of deep learning compilers, GPU systems, and automation infrastructure,... ...identify regressions, bottlenecks, and optimization opportunitiesPartner with compiler and architecture...SeniorFull timeRemote work$184k - $287.5k
NVIDIA is seeking a Senior Software Engineer to join our CSP Engagements team,... ...in embedded firmware, Linux kernel development, and middleware... ...device drivers, and system optimizations for GB200 and next-gen platforms... ...the crowd:Experience with GPU computing (CUDA), deep...SeniorFull time$152k - $241.5k
...for a Deep Learning and Computer Vision engineer for our Autonomous Vehicles team. The... ...Building training pipelines and real-time inference run-times (PyTorch, TensorFlow,... ...LearningInvolvement with architecture optimization, pruning, curriculum & multi-task trainingExperience...SeniorFull timeNight shift$165.2k - $223.6k
...Key job responsibilities- Research, design, and implement Linux kernel changes to meet business requirements- Drive kernel... ...from kernel to application- Profile system performance and drive optimizations across the software stack- Develop tooling for performance characterization...InternshipLocal areaFlexible hours$168.1k - $227.4k
...cloud-scale machinelearning accelerators. This role is for a senior software engineer in the Machine Learning Inference Applications team. This role is responsible for development and performance optimization of core building blocks of LLM Inference - Attention, MLP, Quantization...SeniorWork experience placementFlexible hours$180k
...motivated, and focused on engineering excellence. This organization... ...you will own both the raw GPU supercomputer and the platform... ...stack — from low-level GPU kernel optimizations and Linux kernel internals... ...— to make training and inference at SpaceXAI as fast, reliable...Temporary work$135.2k - $306.4k
Oracle hardware platform development engineering is seeking a highly driven GPU/CPU Platform System Engineer at the Principal Engineer level. The GPU... ...AI platform architectures and assist with scaling & optimizations. You will help support solution operational health visibility...Temporary workWork experience placementRemote workFlexible hours- ...a diverse F5 community where each individual can thrive.Senior Programmability Engineer At F5, we strive to bring a better digital world to life.... ...using this tooling to improve the implementation Profile and optimize performance Help set expectations with software,...SeniorFull timeWork experience placementLocal area
$200k - $287.5k
...to help redefine the future of how work gets done.As a Senior Forward Deployed Engineer (FDE), you will be embedded with customers to lead the modernization... ...the customer's modern data environment.Performance Optimization & Best Practices: Analyze and optimize data models,...Senior$148.5k - $223.9k
...of AI, and you are the future of Salesforce.Overview:As a Senior Threat Detection Engineer, you will take on complete ownership of a technical area,... ...false positives, and fine-tune detection rules for optimal efficacy.Demonstrate in-depth knowledge of fundamental security...SeniorFull time$151.8k - $332.2k
...expect We are looking for an AI Inference Engineer with a solid background in... ...inference hardware, such as GPU, TPU and AI-specific chips.... ...solutions are not available.Optimizing ASR inference systems for... ...developing and tuning custom CUDA kernels, leveraging CUDA Graphs for...Full timeWork at officeRemote work$300k
...team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together... ...beneficial AI systems. About the Role The Cloud Inference team scales and optimizes Claude to serve the massive audiences of developers and...SeniorFull timeWork at officeVisa sponsorshipFlexible hours- Amazon’s Engineering & Design team is seeking a Senior Electrical Engineer to design, analyze, and optimize complex electrical systems that power large-scale IT infrastructure. In this role, you’ll architect high‑reliability power distribution, backup, and control systems...SeniorFull time
$160k - $195k
...underlying physical infrastructure. As the Senior GPU Capacity and Optimization Planner, you will own the day-to-day... ...and our Go-To-Market (GTM) engine. You will work closely with Sales... ...signals and large-scale AI training/inference architectural trends to inform hardware...SeniorTemporary work$124k - $280k
...people in data and analytics engineering focus on leveraging advanced... ...and health plans. As a Senior Manager, you will drive use... ...and genomics) and operational optimization - Foster a collaborative environment... ..., LlamaIndex, Semantic Kernel) to build healthcare AI solutions...SeniorFull timeH1b$200k - $332k
...excited to lead our ML Performance Optimization initiatives and make our Training and Inference platform that enables autonomous... ..., and Advanced Hardware Engineering group and have the opportunity to... ...model training. Experience with GPU-accelerated inference using TensorRT...SeniorTemporary workRelocation package$142k - $220.5k
Job DescriptionA Senior Engineer on the AI Enablement team designs and delivers AI platform capabilities that let Nordstrom Technology use... ...deconstruct complex systems, platforms, and processes.Develop and optimize databases and infrastructure using Kubernetes and AWS....SeniorFull timeTemporary work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Inference Engineer, GPU Kernel Optimization. Be the first to apply!
- senior business analyst Seattle, WA
- senior cost estimator Seattle, WA
- senior manager tax Seattle, WA
- senior automation engineer Seattle, WA
- senior devops Seattle, WA
- senior recruiter Seattle, WA
- senior paralegal Seattle, WA
- senior associate vice president Seattle, WA
- senior content designer Seattle, WA
- senior director digital marketing Seattle, WA


