Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Inference Performance Engineer, AI Inference Configuration Optimization

$124k - $195.5k

NVIDIA

NVIDIA is recruiting a Senior Inference Performance Engineer to push NVIDIA's performance limits on large-scale AI inference benchmarks. This position provides an outstanding opportunity to employ your optimization knowledge in an autonomous optimization framework. AI agents use this framework to repeatedly run benchmark, profile, and tune processes, amplifying the impact of every technique you design. If you enjoy extracting maximum performance from GPUs and scaling your skills beyond your individual efforts, this role is a great fit!What you'll be doing:Distill your performance instincts into reusable skills, workflows, and evidence-backed methodologies that AI agents can complete autonomously. Review agent-generated experiments, validate findings, and curate best-known configurations.Performance improvement of AI inference workloads that methodically increase throughput-per-GPU and user interactivity by exploring configuration options, parallelism techniques, batching, KV cache handling, quantization, and speculative decoding settings.Measure and optimize both aggregated and disaggregated serving architectures across TensorRT-LLM, SGLang, vLLM, and Dynamo on NVIDIA's latest GPU platforms.Profile workloads using Nsight Systems, kernel traces, and internal analysis tools. Use roofline and speed-of-light analysis to find credible headroom and drive fixes from hypothesis to measured wins.Land improvements upstream: serving framework patches, optimized kernels, and deployment recipes that advance the public Pareto frontier while maintaining strict model correctness.Collaborate with TensorRT-LLM, SGLang, vLLM, kernel, benchmarking, and GPU architecture teams to convert profiling insights into delivered performance improvements.What we need to see:BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Applied Math, or a related field, or equivalent experience.3+ years of relevant engineering experience.Must have: Extensive knowledge of the efficiency and optimization involved in AI model execution, covering continuous batching, throughput-latency tradeoffs, KV cache and memory limitations, parallel processing techniques, MoE serving, quantization, and meeting serving SLAs.Must have: Hands-on experience benchmarking and profiling GPU workloads using tools such as Nsight Systems, Nsight Compute, CUPTI, or PyTorch profiler, and interpreting kernel-level performance data.Strong Python engineering skills and the ability to navigate and modify large C++/CUDA serving codebases.Rigorous experimental methodology with controlled single-variable comparisons, reproducible benchmarks, and evidence-backed optimization decisions.Strong written and verbal communication skills to explain performance tradeoffs clearly to both humans and documentation for autonomous systems.Ways to stand out from the crowd:Direct contributions to TensorRT-LLM, vLLM, SGLang, FlashInfer, Dynamo, or comparable inference frameworks.Experience with disaggregated serving, wide expert-parallel MoE inference, KV cache transfer, or NCCL/NIXL/NVSHMEM communication at multi-node scale.CUDA kernel authorship or optimization experience on Hopper/Blackwell architectures, focusing on Tensor Cores, TMA, and warp specialization.Proven results on public inference benchmarks such as MLPerf Inference or SemiAnalysis InferenceX.Experience building or operating agentic AI workflows to automate engineering tasks.Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family #LI-HybridYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 124,000 USD - 195,500 USD for Level 2, and 152,000 USD - 241,500 USD for Level 3.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until August 10, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, CA, RemoteType: Full time

Vacancy posted 8 hours ago
Similar jobs that could be interesting for youBased on the Inference Performance Engineer, AI Inference Configuration Optimization in Santa Clara, CA vacancy
  • $184k - $287.5k

     ...re now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to...  ...inference operation to its performance ceiling? Our LLM...  ...fidelity across the full configuration space that production LLM...  ...optimization: applying AI-driven analysis to diagnose... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    8 hours ago
  •  ...computing experiences—from AI and data centers, to PCs,...  ...looking for a Senior GPU Inference Performance Engineer to own end-to-end performance...  ...performance: Profile and optimize inference engines...  ...fusion, graph compilation), or configuration differences.Multi-server inference... 
    Performance

    AMD

    Santa Clara, CA
    4 days ago
  • $195.2k - $361.2k

     ...MissionAt Intel, our journey is to transform AI into something safer, more trustworthy...  ...the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained...  ...hardware tiers and publish honest performance comparisonsUpstream fixes and patches... 
    Performance
    Full time
    Internship
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

     ...learning ignited modern AI — the next era of...  ...top-tier AI Compiler Engineers to drive innovation within...  ...what is possible in AI performance and help build the...  ...and computational graph optimizations for next-generation NVIDIA...  ...AI workloads (both inference and training) and successfully... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme...  .... You’ll architect and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry benchmarks... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

     ...eager to work on cutting-edge AI technology for safety-...  ...team as a Senior Software Engineer, and be at the forefront of technology, enabling high-performance AI inference solutions for automotive safety...  ...to performance optimization and benchmarking efforts for... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...for a Senior DL Algorithms Engineer! NVIDIA is seeking senior...  ...engineers who are mindful of performance analysis and optimization to help us squeeze every...  ...company that leads the AI revolution.What you will be...  ...language and multimodal model inference as part of NVIDIA... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...builds the world's largest AI chip, 56 times larger than...  ...industry-leading training and inference speeds; over 10 times...  ...RoleWe are hiring a Senior Performance Engineer to join our Product team. You...  ...TensorRT-LLM), GPU kernel-level optimization toolchains (CUDA, Triton),... 
    Performance
    Contract work
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    4 days ago
  •  ...generation computing experiences—from AI and data centers, to PCs, gaming and...  ...ROLEWe are seeking a Principal GenAI Inference Optimization Engineer to join our Models and Applications team. This role focuses on improving performance, efficiency, and scalability of generative... 
    Performance

    AMD

    San Jose, CA
    3 days ago
  • $148k - $235.75k

     ...and pivotal in our inference marketing. You...  ...on working with engineering to understand the...  ...techniques (parallelisms, configurations, etc.). You will...  ...position in AI inference.Want to...  ...and high performance computing. Come grow...  ...specific frameworks & optimizations (Dynamo, Triton... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $224k - $356.5k

     ...which every new AI-powered application...  ...production AI inference for NVIDIA Inference...  ...customers deploy optimized, enterprise-...  ...optimized inference engines, model profiles/recipes...  ...runtime configurations, and security hardening...  ...integration, performance profiling/optimization... 
    Performance

    Socket.dev

    Santa Clara, CA
    12 hours ago
  • $152k - $241.5k

     ...upon which every new AI‑powered application...  ...seeking a Senior Software Engineer - AI Inference to advance open‑...  ...enjoys digging into performance bottlenecks,...  ...features, fixes, and optimizations upstream to vLLM/SGLang...  ...model and hardware configurations.Collaborate with model... 
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $182.5k - $260.5k

     ...networking for the cloud and AI era. We secure and...  ...One platform, its Zero Trust Engine, and the powerful NewEdge network...  ...and control without performance trade-offs.At Netskope, our...  ...Learning Scientist, you own the inference and optimization layer that makes AI in agentic... 
    Performance

    Netskope

    Santa Clara, CA
    8 hours ago
  •  ...computing experiences—from AI and data centers, to...  ...or Principal level engineer who is passionate about...  ...scaling training and inference for the latest Generative...  ...and inference optimizations across a variety of applications...  ...to E2E co-optimize performance on current and future... 
    Performance

    AMD

    San Jose, CA
    4 days ago
  • $250k - $350k

    About the RoleWe are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical role, you...  ...user experiences at scale.You will design and optimize inference pipelines, implement state-of-the-art acceleration... 
    Performance
    Work at office
    3 days per week

    Pika

    Palo Alto, CA
    8 hours ago
  •  ...the world's largest AI chip, 56 times...  ...leading training and inference speeds; over 10 times...  ...Manufacturing Linux / Network Engineer to design,...  ..., security, and performance that modern manufacturing...  ...fabric design.Configure, troubleshoot, and optimize Layer 2/3 networking... 
    Performance

    Cerebras Systems

    Sunnyvale, CA
    4 days ago
  •  ...About Us Hippocratic AI is the leading generative AI company in healthcare....  ...Role We're seeking an experienced LLM Inference Engineer to optimize our large language model (LLM) serving...  ...Continuously benchmark and improve system performance across various deployment scenarios... 
    Performance

    Hippocratic AI

    Palo Alto, CA
    1 day ago
  • $170.6k - $261.3k

     ...Senior Machine Learning Engineer on the State Estimation...  ...to improve model performance against those metrics....  ...efficient training and inference pipelines, including model optimization techniques (e.g., pruning...  ...including interfaces, configuration, deployment, monitoring... 
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    3 days ago
  •  ...Founder Vice President, AI Inference Software About the Company...  ...shaping the future of high-performance AI inference. The successful...  ...building and mentoring an elite engineering team, and partnering with...  ...background in building or optimizing production-scale AI... 
    Performance

    Confidential

    San Jose, CA
    5 days ago
  • $184k - $287.5k

     ...Systems Software test (lead) Engineer to join our Cloud Service...  ...with next-generation high-performance training and inference platforms. You will work...  ...modules and known-good partner configurations.Partner with NVIDIA...  ...existing vacancy. NVIDIA uses AI tools in its recruiting... 
    Performance
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    3 days ago
  • $168k - $258.75k

    Inference is the fastest growing and most competitive area in Generative AI today. It is where AI models impact our...  ...of accuracy and performance matters for quality...  ...roadmaps for model optimization softwareWork with leadership...  ...Science, Computer Engineering, or similar... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...seeking an NCX Senior Engineer to join our DSX team...  ...groundbreaking AI workloads! You will...  ...customers realize efficient performance from NVIDIA's AI...  ...distributed training, inference optimization, and MLOps pipelines...  ...deployment and configuration of GPU‑accelerated clusters... 
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $229.9k - $262.4k

     ...Overview Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform) Overview: At Capital One, we are creating responsible...  ...with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring... 
    Performance
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    more than 2 months ago
  • $152k - $241.5k

     ...driving advancements in AI and machine learning...  ...talented and motivated engineers to join our TensorRT...  ...leading deep learning inference software for NVIDIA AI...  ...implementing inference software optimizations to power AI...  ...Knowledge of close-to-metal performance analysis, optimization... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...Intel is seeking a seasoned software engineer to accelerate AI inference on edge hardware. You will optimize llama.cpp/vLLM, tune KV cache, batching and scheduling...  ...You will work across hardware tiers, benchmark performance, and contribute upstream fixes to open-source... 
    Performance

    PVH (Tommy Hilfiger/Calvin Klein)

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

     ...unlimited potential of AI to define the next era...  ...Developer Technology Engineer, you will be at the forefront...  ...in suboptimal runtime performance.Conduct hands-on...  ...deployment targeting optimal runtime performance.Improve...  ...GPU-accelerated AI inference driven by NVIDIA APIs... 
    Performance
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    1 day ago
  • $137k - $156k

     ...passionate, and committed engineers, technologists,...  ..., Big Data, HPC, AI and Storage, etc....  ..., compatibility, performance, stress, and...  ..., providing optimized benchmarks for HPC...  ...optimizing OS/network configurations, and...  ...MLPerf Training/Inference benchmark, LLM, HPL... 
    Performance
    Worldwide

    Super Micro Computer

    San Jose, CA
    4 days ago
  • $184k - $287.5k

     ...unlimited potential of AI to define the next...  ...CPUs, and a fully optimized NVIDIA AI and HPC software...  ...a highly motivated engineer to lead performance benchmarking and...  ...-world AI training, inference, and HPC workloads at...  ...system tuning, configuration optimization, and architectural... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...globally. We seek a Senior Engineer to lead technical...  ...in deploying advanced AI agent frameworks and local...  ...powerful local inference (Nemotron models) with...  ...engineering efforts to optimize the agent runtimes for...  ...particularly C++ (for performance-critical systems/OS integration... 
    Performance
    Full time
    Local area
    Shift work

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...the potential of generative AI to power the transformation...  ...Principal System Software Engineer, AI Inference ExecutionWhat you will do:The...  ...of what it takes to optimize and trade-off various aspects...  ...toolsExperience with distributed, high-performance software design and... 
    Performance
    3 days per week

    d-Matrix

    Santa Clara, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Inference Performance Engineer, AI Inference Configuration Optimization. Be the first to apply!