Senior Software Engineer, Deep Learning Inference - TensorRT
$152k - $241.5kNVIDIA
We are now looking for a Senior Software Engineer for Deep Learning Inference! Would you like to make a big impact in Deep Learning by helping build a state-of-the-art inference framework for accelerating Deep Learning models, especially Large Language Models, on NVIDIA GPUs? We are now welcoming exceptional software engineers to apply to Senior Engineering positions in the Deep Learning Inference TensorRT software team.What you’ll be doing:Craft and develop robust inferencing software that can be scaled to multiple platforms for functionality and performanceDevelop components of TensorRT, NVIDIA’s SDK for high-performance deep learning inference.Closely follow academic developments in the field of artificial intelligence and feature update TensorRTUse C++ and Python to build graph parsers, optimizers, and tools for effective deployment of trained deep learning models.Collaborate with teams of deep learning experts, GPU architects and DevOps engineers across diverse teams.What we need to see:A Bachelor's, Master's, PhD or equivalent experience in Computer Science, Computer Engineering, Electrical Engineering or related field.3+ years of software development experience.Strong experience with the latest C++ standardsC++11/C++14/C++17/C++20, etc..Strong grasp of Machine Learning concepts.Experience and knowledge in Computer Architecture, Data Structures, Algorithms.Excellent communication skills, and an aptitude for collaboration and teamwork.Ways to stand out from the crowd:Experience developing System Software.Proficiency in Python as well as Background in GPU kernel programming using CUDA or OpenCL.Experience in software performance benchmarking, profiling, and optimizations.Background in compiler developmentExperience in working with TensorRT, PyTorch, TensorFlow, ONNX Runtime or other ML frameworks.NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative, autonomous and love a challenge, we want to hear from you. Come, join our TensorRT Workflows team and help build the real-time, cost-effective computing platform driving our success in this exciting and quickly growing field.#LI-HybridYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until August 9, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time
$152k - $241.5k
...advancements in AI and machine learning to solve some of the world’s... ...talented and motivated engineers to join our TensorRT team in developing the industry-leading deep learning inference software for NVIDIA AI accelerators. As a Senior Software Engineer in the TensorRT...SeniorFull time$193.3k - $261.5k
...Amazon Neuron, the software development kit... ...used to accelerate deep learning and GenAI... ...unparalleled ML inference and training performance... ...software boundary, our engineers build systematic... ...mentorship. Our senior members enjoy one... ...vLLM, SGLang, TensorRT or similar...SeniorWork experience placementInternshipLocal areaFlexible hours$152k - $241.5k
...passionate about driving innovation in deep learning and eager to work on cutting-edge... ...applications? Join NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of... ...technology, enabling high-performance AI inference solutions for automotive safety...SeniorFull time$152k - $241.5k
...about redefining how software is built in the age of... ...AI? Join NVIDIA’s TensorRT team to help lead a first... ...for out-of-framework inference globally. We are... ...systems-thinking C++ engineer who wants to help scale... ...of state-of-the-art deep learning breakthroughs, and improve...SeniorFull time$184k - $287.5k
We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency... ..., parallel programming, distributed systems, deep learning theories.Knowledgeable and passionate about performance...SeniorFull time- NVIDIA Corporation is seeking a Senior Software Engineer for the TensorRT Edge-LLM team in the US. You will develop a high-performance inference framework in modern C++ that extends TensorRT for autoregressive model serving, including speculative decoding and KV cache...Senior
$224k - $356.5k
We are now looking for a Senior System Software Engineer to work on Dynamo. NVIDIA... ...for its GPU-accelerated deep learning software team. Academic and... ...building Generative AI inference platform to make design and... ...across vLLM, SGLang, and TensorRT-LLM, delivering day-0...SeniorFull time$152k - $241.5k
...time large language model inference? Join NVIDIA’s TensorRT Edge-LLM team and help shape... ...robotics. We build the software stack that enables Large Language... ..., Electrical/Computer Engineering, or a closely related... ...software development experience.Deep understanding of...SeniorFull time$224k - $356.5k
...Team is building the software foundation for scalable... ...for exceptional engineers who thrive on solving... ...computing.We are seeking a Senior Software Engineer for... ...and deployment of deep neural networks that... ...core platform, deep learning inference, TensorRT and related compiler/...SeniorFull time$184k - $287.5k
...are looking for a motivated Deep Learning engineer to bring advanced CUDA... ...scales up to 100K GPUs to inference down at microsecond latency... ...systems principles (aka systems software fundamentals)Adaptability and... ...(e.g., PyTorch, JAX, TensorRT, vLLM, sgLang, Nemo, Megatron...SeniorFull time$224k - $356.5k
...advanced computer vision and deep learning. Our team builds large-... ...-world impact. As a System Software Engineer for Vision AI, you will develop... ...and tuning GPU-accelerated inference pipelines to meet strict... ...acceleration (such as CUDA, TensorRT, or comparable technologies...SeniorFull time$184k - $287.5k
We are now seeking a Senior Infrastructure Software Engineer for NVIDIA TensorRT Edge-LLM!NVIDIA's TensorRT Infrastructure group is seeking excellent software engineers... ...vehicles to Jetson AGX for robotics and edge inference applications. You will work with autonomy to...SeniorFull time- ...Job Overview NVIDIA is seeking an experienced Deep Learning Software Engineer, TensorRT Performance to analyze and improve the performance of NVIDIA’s inference ecosystem, including TensorRT, TensorRT‑EdgeLLM and Torch‑TensorRT. Responsibilities Establish groundbreaking...
$184k - $287.5k
...and highly motivated software professional to work... ...intersection of CUDA and Deep Learning Systems. As the... ...in both training and inference pipelines.Collaborate... ...Computer Science, Computer Engineering, Electrical... ...(e.g., PyTorch, JAX, TensorRT, vLLM, sgLang, Nemo,...SeniorFull time- Develop and evolve a state-of-the-art inference framework in modern C++ that extends TensorRT with autoregressive model serving capabilities.... ...in a relevant field and at least 4 years of software development experience. A deep understanding of transformer models and proficiency...Senior
$184k - $287.5k
...revolution, building the software and systems that... ...looking for a Senior Software Engineer to lead the... ...training and inference workloads across... .... You will lead deep performance and... ...intersection of deep learning systems, GPU... ...NeMo / Megatron, TensorRT-LLM, and adjacent...SeniorFull timeRemote work- ...at the forefront of software and hardware innovation... ...System Software Engineer, AI Inference ExecutionWhat you will... ...architecture, and machine learning... ...frameworks (such as TensorRT-LLM, vLLM, SGLang, etc.)Experience with deep learning frameworks (...3 days per week
$184k - $287.5k
Reinforcement learning post-training is driving... .... RL requires inference, rollout generation... ...an RL Frameworks engineering team to develop the... ...team spans the full software stack, from... ...their need optimizing deep learning frameworks... ...engines (vLLM, SGLang, TensorRT-LLM) into RL...Senior$152k - $241.5k
...computing. More recently, GPU deep learning ignited modern AI - the... .... Local AI seeks a Senior Systems Software Engineer interested in solving client... ...pipelines, and inference runtime features. Identifying... ...ONNX RT, DirectX, PyTorch, TensorRT, Vulkan, llama.cpp. We're...SeniorLocal area$152k - $241.5k
...parallel computing. More recently, GPU deep learning ignited modern AI — the next era of... ...for an AI & Deep Learning Compiler Engineer. NVIDIA is hiring software engineers for its Deep Learning &... ...has been the backbone of NVIDIA’s inference engine, spanning across data...SeniorFull timeRemote work$165k - $242k
Apply for the Senior Software Engineer II, Inference role at CoreWeave. CoreWeave is The Essential... ...performance with deep technical expertise to accelerate... ...frameworks (vLLM, Triton, TensorRT‑LLM, Ray Serve, TorchServe... ..., and we’re constantly learning. Our team cares deeply...SeniorPermanent employmentTemporary workCasual workWork at officeRemote workFlexible hoursShift work$139k - $204k
What You’ll Do Senior engineers are area owners who lead designs, raise engineering... ...our Kubernetes‑native inference platform and meet strict P99... ...or Go (C++ a plus) and deep familiarity with networked systems... ...frameworks (vLLM, Triton, TensorRT‑LLM, Ray Serve, TorchServe)....SeniorPermanent employmentTemporary workCasual workWork at officeRemote workFlexible hoursShift work- Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distributed... ...of hands‑on experience in machine learning, deep learning, and software engineering... ...optimizing models for GPU inference (e.g., TensorRT, Triton Inference Server). Knowledge...Senior
$229.9k - $262.4k
Senior Lead AI Engineer (FM Hosting, LLM Inference) Overview: At Capital One, we are creating responsible... ...leader in using machine learning to create real-time,... ...talent — along with our deep experience in machine... ...deploy, and support AI software components including foundation...SeniorFull timePart timeLocal area$152k - $241.5k
We are now looking for a Senior Software Engineer for Quantized Inference! NVIDIA is seeking software engineers to accelerate the discovery and deployment of efficient inference recipes for LLMs. A recipe defines which operators are transformed into low-precision or sparsified...SeniorFull time$152k - $241.5k
...and industries. Within our software stack, CUTLASS stands out as... ...(GEMM) and related math and deep learning computations on NVIDIA GPUs.... ...the-art deep learning models’ inference and training passes to... ...Computer Science, Computer Engineering, or related field (or equivalent...SeniorFull time$184k - $287.5k
We're looking for outstanding AI systems engineers to develop groundbreaking technologies in the inference systems software stack! We build innovative AI systems software... ...with other engineers at NVIDIA across deep learning frameworks, libraries, kernels, and GPU arch...SeniorFull timeRemote work$184k - $287.5k
...lasting impact on the world.We are looking for outstanding Senior Deep Learning Software Engineers to develop and productize NVIDIA's deep learning... ...Developing compiler technologies to accelerate deep learning inference on NVIDIA hardware platforms for Physical AI.Working...SeniorFull timeWork experience placement$184.7k - $324.8k
Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference Santa Clara, California... ...efficiency, hardware/software codesign, systems... ...Python. Strong knowledge of deep learning architectures including... ...frameworks such as TensorRT-LLM, vLLM, SGLang, TGI,...SeniorWorldwideRelocation$152k - $241.5k
We are now looking for a Senior Infrastructure Software Engineer for Deep Learning Libraries!NVIDIA's Deep Learning Libraries Group is seeking excellent software... ...role spans multiple products, including cuDNN, TensorRT, and CUDA kernel libraries. The mission is to design...SeniorFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Software Engineer, Deep Learning Inference - TensorRT. Be the first to apply!
- cybersecurity software engineer Santa Clara, CA
- graduate software engineer Santa Clara, CA
- software developer fintech Santa Clara, CA
- new graduate software engineer Santa Clara, CA
- senior robotics software engineer Santa Clara, CA
- software engineer visa sponsorship Santa Clara, CA
- software qa engineer Santa Clara, CA
- network software engineer Santa Clara, CA
- software engineer remote Santa Clara, CA
- part time software developer remote Santa Clara, CA

