Senior Software Engineer - AI Inference Performance
$184k - $287.5kNVIDIA
NVIDIA is the platform upon which every new AI-powered application is built. We are seeking a Senior Software Engineer – AI Inference Performance to advance innovative LLM and VLM inference. You will push workloads toward practical performance limits on NVIDIA GPU-accelerated systems. Your work will span models, serving software, distributed runtimes, communication, CUDA kernels, and GPU architecture. Deliver measurable gains in latency, throughput, efficiency, and scale.This is a hands-on role for an engineer who turns performance models and profiler data into working code. You will collaborate with model, framework, kernel, networking, and GPU architecture teams. You will contribute improvements to open-source inference engines and develop methods that others can reproduce. Your work will improve production deployments and help build future NVIDIA platforms.What you'll be doing:Lead end-to-end analysis of LLM/VLM inference processes. Define representative prefill and decode workloads. Optimize time to first token, inter-token latency, P99 end-to-end latency, processing efficiency, and key-value (KV) cache capacity. For multimodal models, isolate preprocessing, encoder, and decoder costs.Build speed-of-light and roofline models to quantify performance headroom. Connect arithmetic intensity, bandwidth, occupancy, memory hierarchy, and communication costs to clear optimization hypotheses.Profile workloads using NVIDIA Nsight Systems, Nsight Compute, PyTorch Profiler, and custom instrumentation. Eliminate bottlenecks in host code, CUDA kernels, memory, communication, and scheduling.Tune serving hyperparameters and techniques such as batching, KV-cache management, quantization, speculative decoding, CUDA Graphs, and model parallelism. Choose them based on workload, hardware, model quality, and service-level objectives.Build and optimize performance-critical kernels, including attention, matrix multiplication, mixture-of-experts routing, quantization, and data movement. Use CUDA, CUTLASS, Triton, or related technologies.Establish repeatable benchmarks, canonical run records, and performance regression gates. Manage aspects such as model, precision, hardware, topology, software, features, and workload; Balance between performance and accuracy. Collaborate across with various teams and contribute high-quality upgrades to TensorRT-LLM, vLLM, SGLang, or associated projects.What we need to see:More than 6 years of experience in full-stack LLM/VLM inference performance involving models, serving, distributed runtimes, kernels, and hardware. Your efforts result in measurable gains in production or production-representative environments.Strong programming skills in Python, Rust and/or C++, plus hands-on experience with CUDA or another GPU programming environment.Demonstrated expertise in speed-of-light analysis, roofline models, microbenchmarks, and tools including NVIDIA Nsight Systems and Nsight Compute. You convert profiles into testable hypotheses and validated progress.Deep understanding of GPU architecture, including Tensor Cores, memory hierarchy, caches, occupancy, synchronization, and numerical formats across hardware generations.Practical experience optimizing inference servers and model execution. You can choose techniques for the workload, including batching, scheduling, KV-cache management, quantization, speculative decoding, and various parallelism strategiesUnderstanding of distributed systems and networking for accelerated computing. You can reason about collectives, topology, and scale-up versus scale-out performance.BS or MS in Computer Science, Computer Engineering, or a related field, or equivalent experience.Ways to stand out from the crowd:Contributions to one or more high-performance AI projects. Examples include TensorRT-LLM, vLLM, SGLang, PyTorch, CUDA, Triton, or NCCL.Experience developing AI-agent-supported performance workflows that automatically gather and analyze profiles, identify bottlenecks, explore serving configurations, or produce optimized runtime and kernel code. You validate generated changes through reproducible, human-reviewed tests for performance, model quality, and correctness.Published research, conference presentations, technical talks, or blog posts that clearly explain inference performance methods and results.Delivered advancements for new LLM or VLM architectures, long-context inference, mixture-of-experts models, multimodal pipelines, or large-scale distributed serving.Successfully carrying these out will shape the future of AI inference performance!With competitive salaries and a generous benefits package ( ), we are widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us and, due to outstanding growth, our best-in-class engineering teams are rapidly growing. If you're a creative and autonomous engineer with a real passion for technology, we want to hear from you!Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until October 2, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time
$229.9k - $262.4k
Senior Lead AI Engineer (FM Hosting, LLM Inference) Overview: At Capital One, we are creating responsible and reliable... ...experiences and scalable, high-performance AI infrastructure. At Capital One... ...develop, test, deploy, and support AI software components including foundation...SeniorPerformanceFull timePart timeLocal area$152k - $204k
...is The Essential Cloud for AI™. Built for pioneers by pioneers... ...superior infrastructure performance with deep technical expertise... ...more at What You'll Do: Senior engineers are area owners who lead... ...evolve our Kubernetes-native inference platform and meet strict P99...SeniorPerformancePermanent employmentFull timeTemporary workCasual workWork at officeFlexible hoursShift work$139k - $204k
...is The Essential Cloud for AI™. Built for pioneers by pioneers... ...superior infrastructure performance with deep technical... ...that powers AI training and inference at scale. This is an opportunity... ...AI. What You'll Do As a Senior Software Engineer I (IC3), you will own...SeniorPerformancePermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$182k - $242k
...is The Essential Cloud for AI™. Built for pioneers by pioneers... ...superior infrastructure performance with deep technical expertise... ...role We're looking for a Senior Engineer for CoreWeave's Benchmarking... ...-to-end MLPerf Training and Inference runs, including workload...SeniorPerformancePermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$184k - $287.5k
...forefront of the generative AI revolution, building the software and systems that power... .... We are looking for a Senior Software Engineer to lead the bring-up,... ...training and inference workloads across NVIDIA... ...scale. You will lead deep performance and reliability investigations...SeniorPerformanceFull timeRemote work$92k - $135k
...CoreWeave is The Essential Cloud for AI™. Built for pioneers by... ...superior infrastructure performance with deep technical expertise... ...What You'll Do: Join the Inference team to ship production features... ...mentorship from experienced engineers. About the role: Implement...PerformancePermanent employmentFull timeTemporary workCasual workInternshipWork at officeFlexible hours$184k - $287.5k
Joining NVIDIA's DGX Cloud AI Efficiency Team means... ..., post-training, inference. Our objective is to deliver... ...an AI infrastructure software engineer to join our team. You'... ...of AI systems.As a senior DGX Cloud AI... ...Artificial Intelligence, High-Performance Computing, and Visualization...SeniorPerformanceFull timeRemote work$165.2k - $223.6k
...builds AWS Neuron, the software development kit used... ...unparalleled ML inference and training performance. The Inference Enablement... ...boundary, our engineers build systematic... ...what's possible in AI acceleration. As... ...and mentorship. Our senior members enjoy one-on...PerformanceWork experience placementInternshipLocal areaFlexible hours$229.9k - $262.4k
...Sr. Lead AI Engineer (Inference Optimization, FM Hosting, AI Platform)At Capital One, we are creating... ...experiences and scalable, high-performance AI infrastructure. At Capital One, you... ...develop, test, deploy, and support AI software components including foundation model...SeniorPerformanceFull timePart time$137.1k - $299.3k
...-throughput batch inference, and fine-tuning on... ...for Generative AI at DoorDash, leading... ...serving and inference engines, fine-tuning and... ...is ideal for a senior engineer who enjoys... ...pushing the cost/performance frontier of GPU inference... ...experience in software engineering ~...SeniorPerformanceHourly payFull timeWork at officeLocal areaRemote workFlexible hours$2,500 per month
...design chips, racks, software, and manufacturing... ...focused on inference . Backed by hundreds... ...staffed by leading engineers, Etched is... ...Etched is seeking a Senior Manufacturing Software... ...for our frontier AI systems. This... ...systems, or other high-performance compute platforms....SeniorPerformanceContract workWork at officeRelocation package- ...Python).RCS & SMS Testing: Automate functional, integration, and performance test scenarios for RCS features (read receipts, typing... ...features, including Google Messages web-pairing, smart replies, and AI-assisted workflows (e.g., Magic Compose, Gemini integration).CI...SeniorPerformance
$184k - $287.5k
...are looking for an experienced engineer to join our Planning and... ...evaluation of our Autonomous Vehicle software.Build compelling, data driven... ...precision and recall of performance metricsFamiliarity with C++Experience... ...vacancy. NVIDIA uses AI tools in its recruiting processes...SeniorPerformanceFull timeRemote work- ...planning frameworks that bridge end-to-end AI driving models with deterministic safety... ...in autonomous systems or safety-critical software. Must possess strong C++ development... ...Real-time Systems Systems Integration Performance Optimization Software Architecture...SeniorPerformance
$152k - $241.5k
...a creative and experienced Software Engineer to help us bring NVIDIA's autonomous... ...self-driving vehicles. As a senior software engineer, your... ...large-scale KPIs and performance measurements for autonomous... ...Experience with or knowledge of AI/ML systems, or a background...SeniorPerformanceFull timeRemote work$152k - $241.5k
...into the unlimited potential of AI to define the next era of... ...container, GPU, and systems engineers. When useful, you will apply... .../prediction) inside existing software workflows.What we need to see... ...PyTorchLinux and HPC / large-scale or performance-sensitive...SeniorPerformanceFull timeRemote work$152k - $241.5k
...tapping into the unlimited potential of AI to define the next era of computing.... ...wants a highly motivated and experienced Senior Software Engineer to join us. At a company celebrated... ...distributed computing environments.Optimize the performance and reliability of cloud applications...SeniorPerformanceFull time$184k - $287.5k
...tapping into the unlimited potential of AI to define the next era of computing.... ...world.We are looking for an outstanding Senior Software Engineer to work on our security team focused... ...securing at scale infrastructure, high-performance computing environments, and AI cluster...SeniorPerformanceFull time$117k - $234k
...are not available. Role Summary:As a Senior Software Engineer at Walmart, you will lead the delivery... ...deployment, while integrating advanced AI/ML capabilities to enhance platform intelligence... ...objectives.Monitor application performance, conduct maintenance, and ensure high-...SeniorPerformanceFull timeTemporary workPart timeRemote work$193.3k - $261.5k
...productivity through innovative generative AI technology. As a AI-powered assistant... ...and enhances collaboration. As a Software Development Engineer on the Quick team, you will play a... ...components. Your work will ensure high performance, scalability, and reliability while...SeniorPerformanceInternshipLocal areaWorldwideFlexible hours$117k - $234k
...OfficeRole summary:Join Walmart as a Senior Software Engineer, leading the delivery of scalable software... ...multiple languages while integrating AI and machine learning components to... ...Kotlin or Swift, architecture patterns, performance, testing, and production-quality app delivery...SeniorPerformanceFull timeTemporary workPart time- ...evaluation frameworks and monitoring systems to ensure the reliability and performance of AI applications. Requirements: Candidates must have 5–8 years of professional software engineering experience with strong fundamentals in system design and distributed systems...SeniorPerformance
- ...and scientific discovery to powering AI and the technologies people rely on... ...career. THE ROLE AMD is seeking a Senior Software Development Engineer to develop and optimize software... ...work across GPU kernel development, performance optimization, AI frameworks, runtime...SeniorPerformance
$152k - $241.5k
...groundbreaking developments in High-Performance Computing, Artificial... ...develops, deploys, and runs engineering infrastructure on large-... ...looking for a highly motivated senior software development engineer to develop... ...vacancy. NVIDIA uses AI tools in its recruiting processes...SeniorPerformanceFull timeWorldwide- ...Description Job Description As a Senior AI Engineer on our Core AI team, you will be a cornerstone... ...'ll Bring ~5+ years of professional software engineering experience with 4+ years... ...techniques for improving AI product performance Experience building AI platform...SeniorPerformance
- ...CAREERS AT NETRIS Shape the AI buildout Join the team shaping the networking... ...world's most demanding AI clouds. Senior Software Engineer About Netris Netris is the... ...as the team grows Take ownership of performance, reliability, and security of the controller...SeniorPerformanceFlexible hours
$184k - $287.5k
...tapping into the unlimited potential of AI to define the next era of computing. An... ...world.We are seeking an outstanding Software Engineer to join our US-based networking software... ...develop innovative, scalable, and high-performance hardware-accelerated software solutions...SeniorPerformanceFull timeShift work$174k - $252k
...testing, maintaining, or launching software products, and 1 year of... .... 3 years of experience with performance, large scale systems data analysis... .... Google's software engineers develop the next-generation technologies... ...experience possible.The AI and Infrastructure team is...SeniorPerformanceWorldwide- ...take you? We are developing AI-native products that bring... ...into the physical world. As a Software Engineer, you will set the technical... ...Background in performance engineering or systems optimization... ...optimization Experience mentoring senior engineers or leading...SeniorPerformanceImmediate start
$152k - $241.5k
We are looking for a software engineer with a strong background in parallel processing... ...architecture to push the limits of performance at the intersection of AI, high-performance computing, and... ...out from the crowd: Experience with inference optimization techniques and...SeniorPerformanceFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Software Engineer - AI Inference Performance. Be the first to apply!
- software engineer internship remote Santa Clara, CA
- senior software engineer ruby on rails Santa Clara, CA
- software developer positions Santa Clara, CA
- agile software developer Santa Clara, CA
- software engineer intern Santa Clara, CA
- rust software engineer Santa Clara, CA
- software engineer internship Santa Clara, CA
- entry level software engineer remote Santa Clara, CA
- senior software engineer remote Santa Clara, CA
- senior software design engineer Santa Clara, CA




