Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Software Engineer - AI Inference Performance

$184k - $356.5k
Full-time

NVIDIA

NVIDIA is the platform upon which every new AI-powered application is built. We are seeking a Senior Software Engineer – AI Inference Performance to advance innovative LLM and VLM inference. You will push workloads toward practical performance limits on NVIDIA GPU-accelerated systems. Your work will span models, serving software, distributed runtimes, communication, CUDA kernels, and GPU architecture. Deliver measurable gains in latency, throughput, efficiency, and scale.

This is a hands-on role for an engineer who turns performance models and profiler data into working code. You will collaborate with model, framework, kernel, networking, and GPU architecture teams. You will contribute improvements to open-source inference engines and develop methods that others can reproduce. Your work will improve production deployments and help build future NVIDIA platforms.

What you'll be doing:

  • Lead end-to-end analysis of LLM/VLM inference processes. Define representative prefill and decode workloads. Optimize time to first token, inter-token latency, P99 end-to-end latency, processing efficiency, and key-value (KV) cache capacity. For multimodal models, isolate preprocessing, encoder, and decoder costs.

  • Build speed-of-light and roofline models to quantify performance headroom. Connect arithmetic intensity, bandwidth, occupancy, memory hierarchy, and communication costs to clear optimization hypotheses.

  • Profile workloads using NVIDIA Nsight Systems, Nsight Compute, PyTorch Profiler, and custom instrumentation. Eliminate bottlenecks in host code, CUDA kernels, memory, communication, and scheduling.

  • Tune serving hyperparameters and techniques such as batching, KV-cache management, quantization, speculative decoding, CUDA Graphs, and model parallelism. Choose them based on workload, hardware, model quality, and service-level objectives.

  • Build and optimize performance-critical kernels, including attention, matrix multiplication, mixture-of-experts routing, quantization, and data movement. Use CUDA, CUTLASS, Triton, or related technologies.

  • Establish repeatable benchmarks, canonical run records, and performance regression gates. Manage aspects such as model, precision, hardware, topology, software, features, and workload; Balance between performance and accuracy. Collaborate across with various teams and contribute high-quality upgrades to TensorRT-LLM, vLLM, SGLang, or associated projects.

What we need to see:

  • More than 6 years of experience in full-stack LLM/VLM inference performance involving models, serving, distributed runtimes, kernels, and hardware. Your efforts result in measurable gains in production or production-representative environments.

  • Strong programming skills in Python, Rust and/or C++, plus hands-on experience with CUDA or another GPU programming environment.

  • Demonstrated expertise in speed-of-light analysis, roofline models, microbenchmarks, and tools including NVIDIA Nsight Systems and Nsight Compute. You convert profiles into testable hypotheses and validated progress.

  • Deep understanding of GPU architecture, including Tensor Cores, memory hierarchy, caches, occupancy, synchronization, and numerical formats across hardware generations.

  • Practical experience optimizing inference servers and model execution. You can choose techniques for the workload, including batching, scheduling, KV-cache management, quantization, speculative decoding, and various parallelism strategies

  • Understanding of distributed systems and networking for accelerated computing. You can reason about collectives, topology, and scale-up versus scale-out performance.

  • BS or MS in Computer Science, Computer Engineering, or a related field, or equivalent experience.

Ways to stand out from the crowd:

  • Contributions to one or more high-performance AI projects. Examples include TensorRT-LLM, vLLM, SGLang, PyTorch, CUDA, Triton, or NCCL.

  • Experience developing AI-agent-supported performance workflows that automatically gather and analyze profiles, identify bottlenecks, explore serving configurations, or produce optimized runtime and kernel code. You validate generated changes through reproducible, human-reviewed tests for performance, model quality, and correctness.

  • Published research, conference presentations, technical talks, or blog posts that clearly explain inference performance methods and results.

  • Delivered advancements for new LLM or VLM architectures, long-context inference, mixture-of-experts models, multimodal pipelines, or large-scale distributed serving.Successfully carrying these out will shape the future of AI inference performance!

With competitive salaries and a generous benefits package ( ), we are widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us and, due to outstanding growth, our best-in-class engineering teams are rapidly growing. If you're a creative and autonomous engineer with a real passion for technology, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 30, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Senior Software Engineer - AI Inference Performance in Santa Clara, CA vacancy
  • $152k - $241.5k

     ...driving advancements in AI and machine learning to...  ...talented and motivated engineers to join our TensorRT...  ...-leading deep learning inference software for NVIDIA AI accelerators. As a Senior Software Engineer in the...  ...Knowledge of close-to-metal performance analysis, optimization... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    7 days ago
  • $152k - $241.5k

    We are now looking for a Senior Software Engineer for Deep Learning Inference! Would you like to make a big impact in Deep...  ...TensorRT, NVIDIA’s SDK for high-performance deep learning inference.Closely...  ...an existing vacancy. NVIDIA uses AI tools in its recruiting processes... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  • $184k - $287.5k

    We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency. You’ll architect and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  • $193.3k - $261.5k

     ...Amazon Neuron, the software development kit used...  ...enabling unparalleled ML inference and training performance.The Inference...  ...software boundary, our engineers build systematic infrastructure...  ...what's possible in AI acceleration.As part...  ...and mentorship. Our senior members enjoy one-on... 
    Senior
    Performance
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    4 days ago
  • $152k - $241.5k

     ...learning and eager to work on cutting-edge AI technology for safety-critical applications? Join NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of technology, enabling high-performance AI inference solutions for automotive safety and other... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $165k - $242k

    Apply for the Senior Software Engineer II, Inference role at CoreWeave. CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform...  ...CoreWeave combines superior infrastructure performance with deep technical expertise to... 
    Senior
    Performance
    Permanent employment
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours
    Shift work

    CoreWeave

    Sunnyvale, CA
    5 days ago
  • Cerebras Systems, Inc. is looking for a Senior Performance Engineer to enhance the performance...  ...competitive pricing models for their AI chip. The ideal candidate will have extensive experience with open-source inference frameworks and an understanding of ML... 
    Senior
    Performance

    Cerebras Systems, Inc.

    Sunnyvale, CA
    4 days ago
  • $193.3k - $261.5k

    We develop AWS Neuron, the complete software stack for Trainium, Amazon's custom cloud...  ....As a Sr. Software Development Engineer on the Inference Model Enablement team, you will onboard...  ...Key job responsibilities* Deliver high-performance models using distributed inference... 
    Senior
    Performance
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    5 days ago
  • NVIDIA Corporation is seeking a Senior Software Engineer for the TensorRT Edge-LLM team in the US. You will develop a high-performance inference framework in modern C++ that extends TensorRT for autoregressive model serving, including speculative decoding and KV cache... 
    Senior
    Performance

    NVIDIA

    Santa Clara, CA
    2 days ago
  • NVIDIA is seeking a Senior Agentic AI Software Engineer (Finance) to build agentic systems and advance AI inference workloads in real-world finance contexts. You will design and...  ...drives experimental agents, optimize performance, and contribute to cutting-edge AI research... 
    Senior
    Performance

    Nvidia Corporation in

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

     ...deep learning ignited modern AI — the next era of...  ...& Deep Learning Compiler Engineer. NVIDIA is hiring software engineers for its Deep Learning...  ...the backbone of NVIDIA’s inference engine, spanning across data...  ...deliver leading inference performance, fast build time, reduced... 
    Senior
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    5 days ago
  •  ...headquartered in Santa Clara, CA, seeks a Principal System Software Engineer for AI Inference Execution. You will join the software team to...  ...proficiency in C/C++/Python on Linux, and experience with distributed, high-performance software. #J-18808-Ljbffr Jobleads-US
    Senior
    Performance

    Jobleads-US

    Santa Clara, CA
    2 days ago
  • $152k - $204k

     ...is The Essential Cloud for AI™. Built for pioneers by pioneers...  ...superior infrastructure performance with deep technical expertise...  ...more at What You'll Do: Senior engineers are area owners who lead designs...  ...our Kubernetes-native inference platform and meet strict P99... 
    Senior
    Performance
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours
    Shift work

    CoreWeave

    Sunnyvale, CA
    9 days ago
  • $229.9k - $262.4k

    Senior Lead AI Engineer (FM Hosting, LLM Inference) Overview: At Capital One, we are creating responsible and reliable...  ...experiences and scalable, high-performance AI infrastructure. At Capital One...  ...develop, test, deploy, and support AI software components including foundation... 
    Senior
    Performance
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    1 day ago
  • $184k - $287.5k

     ...unlimited potential of AI to define the next era...  ...Large-scale model inference architectureContribute...  ...Deep knowledge of E2E AV software integration from perception...  ...management, and performance tuning.What We Need To...  ...See:We’re looking for engineers who can build systems... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

    NVIDIA GPU Architecture Group is seeking a senior software engineer to automate and optimize performance analysis workflows for AI training and inference workloads. You will not only perform analysis but also reshape how it's done, building tools and workflows that scale... 
    Senior
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    5 days ago
  • $152k - $241.5k

    NVIDIA's high-performance computing platforms are powering the AI revolution across many applications...  ...industries. Within our software stack, CUTLASS stands...  ...deep learning models’ inference and training passes to...  ...Computer Science, Computer Engineering, or related field (or... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

     ...large language model inference? Join NVIDIA’s TensorRT...  ...next generation of edge AI for automotive and robotics. We build the software stack that enables...  ...robotics to deliver high-performance, production-ready...  ..., Electrical/Computer Engineering, or a closely related... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  • $184k - $287.5k

    We are now looking for a Senior Software Engineer for AI Resiliency!At NVIDIA, we are pushing the boundaries...  ...-level C++ and Python code. Enhance performance for AI workloads running on...  ...seamless operation of AI training and inference workloads.What We Need to See:You've... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $224k - $356.5k

     ...Team is building the software foundation for scalable, high-performance vehicle computing...  ...for exceptional engineers who thrive on solving...  ....We are seeking a Senior Software Engineer...  ...performance, AI model optimization...  ...platform, deep learning inference, TensorRT and... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...the unlimited potential of AI to define the next era of...  ...looking for outstanding Senior Deep Learning Software Engineers to develop and productize...  ...accelerate deep learning inference on NVIDIA hardware...  ...level kernel development and performance optimization.Develop workflows... 
    Senior
    Performance
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    5 days ago
  • $184k - $287.5k

    We're looking for outstanding AI systems engineers to develop groundbreaking technologies in the inference systems software stack! We build innovative AI systems software to accelerate...  ...experience in GPU kernel development and performance optimizations (especially using CUDA C/... 
    Senior
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    5 days ago
  • $152k - $241.5k

    We are now looking for a Senior Software Engineer for Quantized Inference! NVIDIA is seeking software engineers to accelerate...  ...into functionally correct, performant code, e.g., writing Triton kernels...  ..., well-tested code; fluent with AI-assisted toolingExperience with ML... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    6 days ago
  • $152k - $241.5k

     ...We are looking for a driven Software Engineer to bring ground breaking models...  ...the future of Physical AI!What you'll be doing:Design,...  ...and optimize containerized inference execution for the latest 3D...  ...validate the models accuracy and performance (latency, throughput,... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  • $184k - $287.5k

     ...Artificial Intelligence, High Performance Computing and...  ...motivated Deep Learning engineer to bring advanced CUDA...  ...technologies into AI stacks, including PyTorch...  ...scales up to 100K GPUs to inference down at microsecond...  ...(aka systems software fundamentals)Adaptability... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  •  ...Superintelligence Cloud, is a leader in AI cloud infrastructure...  ...the RoleWe are seeking a Senior Software Engineer to join our Managed...  ...generation of AI training and inference at scale.As a Senior Engineer...  ...systems that are reliable, performant, and elegantly simple for... 
    Senior
    Performance
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    3 days ago
  • $224k - $356.5k

    We are now looking for a Senior System Software Engineer to work on Dynamo. NVIDIA is...  ...to power a revolution in AI, enabling breakthroughs in...  ...team building Generative AI inference platform to make design...  ...build robust, scalable, high performance software components to... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

    We are now looking for a Senior Software Engineer for Agent Simulation & Evaluation...  ...the unlimited potential of AI to define the next era of...  ...their best work.Demand for inference is growing rapidly, driven...  ...model architectures with performance analysis across GPU and LPU... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    19 hours ago
  • $152k - $241.5k

     ...Artificial Intelligence, High-Performance Computing and Visualization...  ...for highly motivated Senior Software Engineers to join our Fabric Networking...  ...recovery, and large-scale AI infrastructure, contributing...  ..., and distributed training/inference systems.Experience with server... 
    Senior
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

    We are now looking for a Senior Agentic AI Software Engineer! Today, NVIDIA is tapping into the unlimited...  ...inspired to do their best work.Demand for inference is growing rapidly, and coding and...  ...with how to evaluate performance of AI agents Experience using coding... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Software Engineer - AI Inference Performance. Be the first to apply!