Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Inference Engineer, GPU Kernel Optimization

$184k - $287.5k

NVIDIA

We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop silicon-measured kernel benchmarking infrastructure, model-level performance projection tooling, and agentic optimization systems that improve GPU kernels at the assembly layer. Our team works closely with compiler, kernel, hardware, and framework organizations across NVIDIA to surface bottlenecks and ship measurable gains. If driving GPU performance at the frontier of LLM inference sounds like your kind of challenge, we'd love to meet you!What you'll be doing:The role drives three interconnected systems, all aimed at accelerating NVIDIA's LLM inference stack. The first is GPU kernel microbenchmarking: measuring competing kernel implementations at real-silicon fidelity across the full configuration space that production LLM deployments demand. The second is end-to-end model performance analysis: connecting performance evidence to model-level serving economics, surfacing high-value optimization opportunities, and producing optimization policies for production inference deployments. The third is agentic kernel optimization: applying AI-driven analysis to diagnose performance gaps, explore optimization opportunities across the kernel ecosystem, and validate findings with rigorous silicon measurements. All three streams converge in close collaboration with compiler, hardware, kernel, and framework teams to deliver upstream improvements and production-grade performance gains.What we need to seeMaster's or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.6+ years of relevant industry experience.Experience building or directing agentic AI systems — code generation, automated optimization, or multi-step reasoning workflows.Strong Python and C++ skills with proven software engineering fundamentals.Hands-on GPU profiling with CUPTI, NSYS, and NCU; proven track record to attribute bottlenecks across kernel execution, compiler decisions, and runtime scheduling.Direct experience with LLM inference frameworks such as TRT-LLM, SGLang, or vLLM and clear understanding of how kernel selection drives model-level throughput and latency.Working knowledge of GPU kernel optimization — CUDA, CUTLASS, Triton, or equivalent — and the ability to read PTX or SASS output.Ways to stand out from the crowdDeep knowledge of SASS/PTX-level kernel analysis, compiler middle-end optimization, or GPU code generation pipelines (LLVM, MLIR, ptxas, or similar).Track record shipping agentic systems end-to-end — tool invent, multi-agent orchestration, and silicon-verified validation — within a performance engineering or kernel optimization context.Active contributions to open-source LLM inference or GPU kernel libraries (FlashInfer, Triton, CUTLASS, or similar).Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until July 31, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, NY, New York; US, WA, SeattleType: Full time

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Senior Inference Engineer, GPU Kernel Optimization in Seattle, WA vacancy
  • $182k - $242k

     ...cloud for high-performance GPU infrastructure across...  ..., and real-time inference. Our stack is engineered for speed, scale, and cost...  ...We're looking for a Senior Engineer for CoreWeave'...  ...Performance team, focused on kernel authoring and optimization. You will write,... 
    Senior
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Bellevue, WA
    23 days ago
  • $226k - $307k

     ...Machine Learning and System Optimization Engineer, you will orchestrate and allocate...  ...allow for more efficient inference by sharing various parts of...  ...models, write custom CUDA kernels, and build highly concurrent...  ...distribute system resources (CPU/GPU/interconnect) to various... 
    Senior
    Full time
    Temporary work
    Relocation package

    Zoox

    Seattle, WA
    9 hours ago
  • $152k - $241.5k

     ...the world.NVIDIA is hiring a Senior Compiler Engineer to join our team driving the next generation of GPU systems programming. We are redefining...  ...-safe, high-performance GPU kernels in idiomatic Rust.What you’ll...  ...code.Build compiler IRs and Optimizers: Work with modern compiler... 
    Senior
    Full time
    Remote work

    Nvidia

    Seattle, WA
    9 hours ago
  • $184k - $287.5k

     ...s invention of the GPU 1999 sparked the growth...  ...'re looking for a Senior Performance Compiler Engineer to join our team...  ...both training and inference. You will be immersed...  ...opportunities for optimization.Designing and implementing...  ...high-level kernel descriptions (written... 
    Senior
    Full time
    Remote work

    Nvidia

    Seattle, WA
    4 days ago
  •  ...intelligence. As a Model Optimization & Deployment Engineer, you will focus on bringing highly...  ...ML models, write custom CUDA kernels, and build highly concurrent inference code to ensure real-time,...  ...runtime efficiency optimization for GPU clusters. Experience with... 
    Senior
    Temporary work
    Relocation package

    Zoox

    Seattle, WA
    28 days ago
  • $224k - $356.5k

     .... An era in which our GPU acts as the brains of...  ...leading open-source LLM inference frameworks — identify...  ..., fallback paths, and optimization opportunitiesCharacterize...  ...Science, Computer Engineering, Electrical Engineering...  ...on experience with GPU kernel development or... 
    Senior
    Full time
    Local area

    Nvidia

    Seattle, WA
    1 day ago
  • $164k - $313.3k

     ...ART is seeking a Senior Machine Learning (...  ...Systems & Efficiency Engineer to join our R&D...  ...ready improvements in inference performance,...  ..., and performance optimization. You will work closely...  ..., improve GPU utilization, and build...  ...runtime performance.Kernel Development & System... 
    Senior
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    Seattle, WA
    1 day ago
  • $184k - $287.5k

    NVIDIA is seeking an NCX Senior Engineer to join our DSX team, collaborating closely with...  ...platforms, including distributed training, inference optimization, and MLOps pipelines constructed on...  ...using Kubernetes, containers, and GPU scheduling systems aligned to NCP builds... 
    Senior
    Full time
    Remote work

    Nvidia

    Seattle, WA
    3 days ago
  • $242k - $389k

     ...to drive our ML Performance Optimization initiatives and make our ML models...  ...a team of strong software engineers and act as a force multiplier...  ...cutting-edge ML Training OR Inference performance optimization techniques...  ...training.Experience with GPU-accelerated inference using... 
    Senior
    Full time
    Remote work

    Zoox

    Seattle, WA
    2 days ago
  • $217k - $307k

     ...achieve maximum throughput at the most optimal power levels. The Software...  ...meet the required specifications. As a GPU performance software engineer within the Software Performance team...  ...experience debugging/optimizing GPU kernels using tools like Nsight.Strong knowledge... 
    Senior
    Full time
    Temporary work
    Relocation package

    Zoox

    Seattle, WA
    1 day ago
  •  ...feedback (RLHF), and continual learning, to optimize AI customer service responses with...  ...compression, quantization, and efficient inference techniques to ensure the AI customer...  ...with majors in computer science, computer engineering, statistics, applied mathematics, data... 
    Senior
    Work experience placement

    TikTok

    Seattle, WA
    3 days ago
  • $315k

     ...committed researchers, engineers, policy experts, and...  ...breakthrough innovations in GPU performance and...  ...developing cutting-edge optimizations that directly enable new...  ...dramatically improve inference efficiency. Working...  ...from custom kernel development to distributed... 
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    Seattle, WA
    more than 2 months ago
  • $280k

     ...committed researchers, engineers, policy experts, and business...  ...developing systems that optimize the throughput and...  ...large-scale ML systems * GPU/Accelerator programming...  ...models * Implement GPU kernels to adapt our models to low-precision inference * Write a custom load-... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    Seattle, WA
    1 day ago
  • $152k - $241.5k

     ...AI workloads, and we are looking for an engineer focused on performance validation, analysis...  ...of deep learning compilers, GPU systems, and automation infrastructure,...  ...identify regressions, bottlenecks, and optimization opportunitiesPartner with compiler and architecture... 
    Senior
    Full time
    Remote work

    Nvidia

    Seattle, WA
    1 day ago
  • $184k - $287.5k

    NVIDIA is seeking a Senior Software Engineer to join our CSP Engagements team,...  ...in embedded firmware, Linux kernel development, and middleware...  ...device drivers, and system optimizations for GB200 and next-gen platforms...  ...the crowd:Experience with GPU computing (CUDA), deep... 
    Senior
    Full time

    Nvidia

    Seattle, WA
    4 days ago
  • $152k - $241.5k

     ...for a Deep Learning and Computer Vision engineer for our Autonomous Vehicles team. The...  ...Building training pipelines and real-time inference run-times (PyTorch, TensorFlow,...  ...LearningInvolvement with architecture optimization, pruning, curriculum & multi-task trainingExperience... 
    Senior
    Full time
    Night shift

    Nvidia

    Seattle, WA
    2 days ago
  • $165.2k - $223.6k

     ...Key job responsibilities- Research, design, and implement Linux kernel changes to meet business requirements- Drive kernel...  ...from kernel to application- Profile system performance and drive optimizations across the software stack- Develop tooling for performance characterization... 
    Internship
    Local area
    Flexible hours

    Amazon

    Seattle, WA
    3 days ago
  • $168.1k - $227.4k

     ...cloud-scale machinelearning accelerators. This role is for a senior software engineer in the Machine Learning Inference Applications team. This role is responsible for development and performance optimization of core building blocks of LLM Inference - Attention, MLP, Quantization... 
    Senior
    Work experience placement
    Flexible hours

    Amazon

    Seattle, WA
    3 days ago
  • $180k

     ...motivated, and focused on engineering excellence. This organization...  ...you will own both the raw GPU supercomputer and the platform...  ...stack — from low-level GPU kernel optimizations and Linux kernel internals...  ...— to make training and inference at SpaceXAI as fast, reliable... 
    Temporary work

    X

    Seattle, WA
    3 days ago
  • $135.2k - $306.4k

    Oracle hardware platform development engineering is seeking a highly driven GPU/CPU Platform System Engineer at the Principal Engineer level. The GPU...  ...AI platform architectures and assist with scaling & optimizations. You will help support solution operational health visibility... 
    Temporary work
    Work experience placement
    Remote work
    Flexible hours

    Oracle Corporation

    Seattle, WA
    3 days ago
  •  ...a diverse F5 community where each individual can thrive.Senior Programmability Engineer At F5, we strive to bring a better digital world to life....  ...using this tooling to improve the implementation Profile and optimize performance Help set expectations with software,... 
    Senior
    Full time
    Work experience placement
    Local area

    F5 Networks

    Seattle, WA
    1 day ago
  • $200k - $287.5k

     ...to help redefine the future of how work gets done.As a Senior Forward Deployed Engineer (FDE), you will be embedded with customers to lead the modernization...  ...the customer's modern data environment.Performance Optimization & Best Practices: Analyze and optimize data models,... 
    Senior

    Snowflake

    Bellevue, WA
    3 days ago
  • $148.5k - $223.9k

     ...of AI, and you are the future of Salesforce.Overview:As a Senior Threat Detection Engineer, you will take on complete ownership of a technical area,...  ...false positives, and fine-tune detection rules for optimal efficacy.Demonstrate in-depth knowledge of fundamental security... 
    Senior
    Full time

    Salesforce

    Bellevue, WA
    9 hours ago
  • $151.8k - $332.2k

     ...expect We are looking for an AI Inference Engineer with a solid background in...  ...inference hardware, such as GPU, TPU and AI-specific chips....  ...solutions are not available.Optimizing ASR inference systems for...  ...developing and tuning custom CUDA kernels, leveraging CUDA Graphs for... 
    Full time
    Work at office
    Remote work

    Zoom

    Seattle, WA
    2 days ago
  • $300k

     ...team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together...  ...beneficial AI systems. About the Role The Cloud Inference team scales and optimizes Claude to serve the massive audiences of developers and... 
    Senior
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    Seattle, WA
    13 hours ago
  • Amazon’s Engineering & Design team is seeking a Senior Electrical Engineer to design, analyze, and optimize complex electrical systems that power large-scale IT infrastructure. In this role, you’ll architect high‑reliability power distribution, backup, and control systems... 
    Senior
    Full time

    Amazon

    Bellevue, WA
    3 days ago
  • $160k - $195k

     ...underlying physical infrastructure. As the Senior GPU Capacity and Optimization Planner, you will own the day-to-day...  ...and our Go-To-Market (GTM) engine. You will work closely with Sales...  ...signals and large-scale AI training/inference architectural trends to inform hardware... 
    Senior
    Temporary work

    Crusoe

    Seattle, WA
    12 days ago
  • $124k - $280k

     ...people in data and analytics engineering focus on leveraging advanced...  ...and health plans. As a Senior Manager, you will drive use...  ...and genomics) and operational optimization - Foster a collaborative environment...  ..., LlamaIndex, Semantic Kernel) to build healthcare AI solutions... 
    Senior
    Full time
    H1b

    PwC

    Seattle, WA
    3 days ago
  • $200k - $332k

     ...excited to lead our ML Performance Optimization initiatives and make our Training and Inference platform that enables autonomous...  ..., and Advanced Hardware Engineering group and have the opportunity to...  ...model training. Experience with GPU-accelerated inference using TensorRT... 
    Senior
    Temporary work
    Relocation package

    Zoox

    Seattle, WA
    more than 2 months ago
  • $142k - $220.5k

    Job DescriptionA Senior Engineer on the AI Enablement team designs and delivers AI platform capabilities that let Nordstrom Technology use...  ...deconstruct complex systems, platforms, and processes.Develop and optimize databases and infrastructure using Kubernetes and AWS.... 
    Senior
    Full time
    Temporary work

    Nordstrom

    Seattle, WA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Inference Engineer, GPU Kernel Optimization. Be the first to apply!