Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Inference Engineer, GPU Kernel Optimization

$184k - $287.5k

NVIDIA

We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop silicon-measured kernel benchmarking infrastructure, model-level performance projection tooling, and agentic optimization systems that improve GPU kernels at the assembly layer. Our team works closely with compiler, kernel, hardware, and framework organizations across NVIDIA to surface bottlenecks and ship measurable gains. If driving GPU performance at the frontier of LLM inference sounds like your kind of challenge, we'd love to meet you!What you'll be doing:The role drives three interconnected systems, all aimed at accelerating NVIDIA's LLM inference stack. The first is GPU kernel microbenchmarking: measuring competing kernel implementations at real-silicon fidelity across the full configuration space that production LLM deployments demand. The second is end-to-end model performance analysis: connecting performance evidence to model-level serving economics, surfacing high-value optimization opportunities, and producing optimization policies for production inference deployments. The third is agentic kernel optimization: applying AI-driven analysis to diagnose performance gaps, explore optimization opportunities across the kernel ecosystem, and validate findings with rigorous silicon measurements. All three streams converge in close collaboration with compiler, hardware, kernel, and framework teams to deliver upstream improvements and production-grade performance gains.What we need to seeMaster's or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.6+ years of relevant industry experience.Experience building or directing agentic AI systems — code generation, automated optimization, or multi-step reasoning workflows.Strong Python and C++ skills with proven software engineering fundamentals.Hands-on GPU profiling with CUPTI, NSYS, and NCU; proven track record to attribute bottlenecks across kernel execution, compiler decisions, and runtime scheduling.Direct experience with LLM inference frameworks such as TRT-LLM, SGLang, or vLLM and clear understanding of how kernel selection drives model-level throughput and latency.Working knowledge of GPU kernel optimization — CUDA, CUTLASS, Triton, or equivalent — and the ability to read PTX or SASS output.Ways to stand out from the crowdDeep knowledge of SASS/PTX-level kernel analysis, compiler middle-end optimization, or GPU code generation pipelines (LLVM, MLIR, ptxas, or similar).Track record shipping agentic systems end-to-end — tool invent, multi-agent orchestration, and silicon-verified validation — within a performance engineering or kernel optimization context.Active contributions to open-source LLM inference or GPU kernel libraries (FlashInfer, Triton, CUTLASS, or similar).Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until July 31, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, NY, New York; US, WA, SeattleType: Full time

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior Inference Engineer, GPU Kernel Optimization in Santa Clara, CA vacancy
  • CoreWeave is hiring a Senior Engineer for its Benchmarking & Performance team to write, profile, and optimize GPU kernels on the LLM inference path. You will improve latency and throughput and collaborate with product, orchestration, and hardware teams to achieve strict... 
    Senior

    CoreWeave

    Sunnyvale, CA
    3 days ago
  • NVIDIA in Santa Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design and optimize kernels, build new abstractions for LLM serving engines, and contribute to accelerators and runtimes that power large language models... 
    Senior

    NVIDIA

    Santa Clara, CA
    4 days ago
  •  ...career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance...  ...HBM bandwidth, compute utilization, kernel scheduling, memory allocation, and...  ...framework performance: Profile and optimize inference engines including vLLM, SGLang... 
    Senior

    AMD

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

    We are now looking for a Senior Formal Verification Engineer for GPU Kernels! Modern AI performance relies on highly optimized GPU kernels — performance-critical code where bugs can be hard to catch and expensive to miss. NVIDIA's Deep Learning Safety Team is hiring engineers... 
    Senior
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    1 day ago
  • $195.2k - $361.2k

    ## Sr. Inference Optimization Engineer (local / edge runtime)Applylocations: US, California, Santa Clara:...  ...constrained local and edge environments — GPU/iGPUs, Vulkan backends — not...  ...Vulkan, CUDA, SYCL, Metal) or SIMD / CPU kernels* Familiarity with quantization formats... 
    Senior
    Internship
    Local area
    Shift work

    Intel Corporation

    Santa Clara, CA
    3 days ago
  • $193.3k - $261.5k

    We are looking for a Senior Inference Engineer to own inference for real-time multimodalconversational...  ...in• Implement and optimize the inference path for...  ...and tune high-performance kernels for critical operations...  ...load• Implement multi-GPU inference (tensor parallelism... 
    Senior
    Internship
    Local area
    Flexible hours

    Amazon

    Sunnyvale, CA
    7 hours ago
  • CoreWeave is seeking a Senior Engineer for the Benchmarking & Performance team to own CUDA kernel authoring and optimization for AI inference. You will write, profile, and tune kernels on the critical path of large‑scale model serving to maximize throughput and minimize... 
    Senior

    CoreWeave

    Sunnyvale, CA
    6 days ago
  • Intel is seeking an Sr. Inference Optimization Engineer to accelerate models on edge devices. You will optimize inference engines for constrained hardware, including profiling, KV cache tuning, batching, and quantization. This role focuses on fast, efficient local AI with... 
    Senior
    Local area

    Intel Corporation

    Santa Clara, CA
    3 days ago
  • $195.2k - $361.2k

     ...people it servesRole SummaryMake models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments — GPU/iGPUs, Vulkan backends — not datacenter H100 environment, mostly PC/edge. KV cache, batching... 
    Senior
    Full time
    Internship
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    4 days ago
  •  ...California is looking for a Member of Technical Staff for Kernel/Compiler/Communication. This critical role requires strong expertise in CUDA and GPU optimization, along with 5+ years of experience in performance engineering. The ideal candidate will design high-performance... 
    Senior

    RadixArk

    Palo Alto, CA
    5 days ago
  • $184k - $287.5k

     ...are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme...  ...architect and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry benchmarks, and scale... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...Algorithmic Model Optimization Team specifically...  ...models for maximal inference efficiency using techniques...  ...by research and engineering teams alike...  ...now looking for a Senior Deep Learning Software...  ...high-performance kernel implementations in...  ...and profile GPU kernel-level performance... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...Instinct™ GPUs. You’ll work across kernels, distributed training, and framework...  ...is passionate about software engineering and the craft of training performance...  ...stability across data, model, and optimizer steps.Optimize multi‑GPU/multi‑node training and communication... 
    Senior

    AMD

    San Jose, CA
    1 day ago
  • $152k - $241.5k

     ...the world.NVIDIA is hiring a Senior Compiler Engineer to join our team driving the next generation of GPU systems programming. We are redefining...  ...-safe, high-performance GPU kernels in idiomatic Rust.What you’ll...  ...code.Build compiler IRs and Optimizers: Work with modern compiler... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...Overview We are seeking a Senior GPU Systems & Fabric Engineer to serve as the critical...  ...deep expertise in Linux kernel internals, GPU...  ...plane. Configure and optimize high-performance host networking...  ...enable efficient multi-tenant inference workloads and maximize... 
    Senior
    Full time
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    7 days ago
  • $184k - $287.5k

    We are now looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior engineers...  ...performance analysis and optimization to help us squeeze every...  .../software stack from GPU architecture to Deep Learning...  ...and multimodal model inference as part of NVIDIA Inference... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...you will play a pivotal role in optimizing and developing deep learning...  ...will be critical in enhancing GPU kernels, deep learning models, and training/inference performance across multi-GPU and...  ...compiler technologies and advanced engineering principles to drive continuous... 
    Senior

    AMD

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

    We are seeking a self‑motivated senior engineer for the Aerial Omniverse Digital Twin team. This...  ...design and implementation of a real‑time, GPU‑accelerated propagation engine that...  ...compute‑vs‑bandwidth trade‑offs at the kernel level.Working knowledge of electromagnetic... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...ROLEWe are seeking a Principal GenAI Inference Optimization Engineer to join our Models and Applications team...  ...AI inference workloads on AMD GPU platforms. You will contribute to optimizing...  ...working across multiple layers—from kernels and runtimes to frameworks and... 

    AMD

    San Jose, CA
    3 days ago
  • $184k - $287.5k

     ...s invention of the GPU 1999 sparked the growth...  ...'re looking for a Senior Performance Compiler Engineer to join our team...  ...both training and inference. You will be immersed...  ...opportunities for optimization.Designing and implementing...  ...high-level kernel descriptions (written... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...transforming every industry. GPU-accelerated deep...  ...an exceptional Senior Perception Engineer to help design and productize...  ...platforms, including optimization for latency, memory, and...  ...training or inference pipelines through custom CUDA kernels or other GPU-accelerated... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

     .... An era in which our GPU acts as the brains of...  ...Deep Learning Compiler Engineer. NVIDIA is hiring software...  ...backbone of NVIDIA inference engine, spanning...  ...and developing compiler optimization algorithms.Collaborating...  ..., such as PyTorch.GPU kernel generation with high performance... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $139k - $208.4k

     ...the world. Come build with us!Role and ResponsibilitiesAs a Senior Engineer GPU RTL Design Engineer, you will contribute to the design and...  ...RTL design, front-end implementation, power and performance optimization, firmware enablement, and the infrastructure needed to... 
    Senior
    Hourly pay
    Full time
    Worldwide
    Relocation

    Samsung Semiconductor

    San Jose, CA
    12 hours ago
  • $224k - $356.5k

     ...transforming every industry. GPU-accelerated deep...  ...an exceptional Senior Radar Perception Engineer to help design and productize...  ..., and layout optimization to support L2-L4 autonomous...  ...training or inference pipelines through custom CUDA kernels or other GPU-accelerated... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $166k - $225k

    Databricks is looking for a Sr. Research Engineer to join their Scaling Team in Mountain View, California. The role involves driving performance improvements and designing high-performance GPU kernels for training workloads. Candidates should have strong experience in CUDA... 
    Senior

    Databricks

    Mountain View, CA
    2 days ago
  • $184k - $287.5k

    NVIDIA is seeking a Compute Kernel Performance Architect who can...  ...workloads with a strong focus on GPU power behavior. In this role,...  ..., silicon validation engineers, and software teams to characterize...  ...on experience developing and optimizing GPU kernels, including work... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...TEAM: AMD’s Data Center GPU organization builds...  ...The Triton compiler and kernel teams are central to...  ...ROLE:AMD is seeking a Senior Product Manager to own...  ...features, performance optimization, bug fixes, stability,...  ...in Computer Science, Engineering, or related field, or... 
    Senior

    AMD

    Santa Clara, CA
    1 day ago
  • $124k - $208.4k

     ...by millions of people around the world. Come build with us!Role and ResponsibilitiesAs a Senior Engineer, GPU Architect, you will work on the analysis, verification, and optimization of end-to-end system performance for Samsung’s premium mobile GPUs.In this mid-to-senior... 
    Senior
    Hourly pay
    Full time
    Relocation

    Samsung Semiconductor

    San Jose, CA
    4 days ago
  • $124k - $208.4k

     ...the world. Come build with us!Role and ResponsibilitiesAs a Senior Engineer, GPU Modeling Architect, you will work on the design and...  ...bottlenecks, and proposing architectural and micro-architectural optimizations to improve GPU performance, power efficiency, and... 
    Senior
    Hourly pay
    Full time
    Relocation

    Samsung Semiconductor

    San Jose, CA
    4 days ago
  • $124.5k - $272k

     ...Netskope One platform, its Zero Trust Engine, and the powerful NewEdge network to gain...  ...Instagram.Positions are available at Senior Staff and above. Candidates are...  ...Machine Learning Scientist, you own the inference and optimization layer that makes AI in agentic workflows... 
    Senior

    Netskope

    Santa Clara, CA
    12 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Inference Engineer, GPU Kernel Optimization. Be the first to apply!