Senior Software Engineer - TensorRT Edge-LLM
$152k - $241.5kNVIDIA
Are you passionate about pushing the limits of real-time large language model inference? Join NVIDIA’s TensorRT Edge-LLM team and help shape the next generation of edge AI for automotive and robotics. We build the software stack that enables Large Language, Vision-Language, and Multimodal (LLM/VLM/VLA) models to run efficiently on embedded and edge platforms — delivering cutting-edge generative AI experiences directly on-device.What you’ll be doing:Develop and evolve a state-of-the-art inference framework in modern C++ that extends TensorRT with autoregressive model serving capabilities, including speculative decoding, LoRA, MoE, and KV cache management.Design and implement compiler and runtime optimizations tailored for transformer-based models running on constrained, real-time platforms.Collaborate with teams across CUDA, kernel libraries, compilers, and robotics to deliver high-performance, production-ready solutions.Contribute to CUDA kernel and operator development for critical transformer components such as attention, GEMM, and MoE.Benchmark, profile, and optimize inference performance across diverse embedded and automotive environments.Stay ahead of the rapidly evolving LLM/VLM ecosystem and bring emerging techniques into product-grade software.What we need to see:BS, MS, PhD, or equivalent experience in Computer Science, Electrical/Computer Engineering, or a closely related field.4+ years of relevant software development experience.Deep understanding of transformer models and inference optimization techniques (e.g., quantization, tensor parallelism, or memory-efficient scheduling).Proficient programming ability with modern C++ (C++11/14/17 and beyond).Familiarity with popular LLM frameworks and libraries such as TensorRT, TensorRT-LLM, vLLM, SGLang, MLC-LLM, or FlashInfer.A track record of strong software design, execution, and collaboration across fields.Ways to stand out from the crowd:Demonstrated development experience or open-source contributions to LLM inference frameworks and libraries, such as SGLang, vLLM, or FlashInfer.Proficiency with CUDA, including efficient kernel development, performance profiling, and GPU architecture fundamentals.Prior work on autoregressive LLM serving systems, including speculative decoding or KV cache management.Familiarity with compiler infrastructure for large language model inference.Exposure to robotics or embedded AI pipelines, including optimizing for low-latency, resource-constrained systems.NVIDIA is widely considered to be one of the technology world’s most desirable employers. We hire some of the most brilliant and forward-thinking people in the world. If you thrive on innovation, autonomy, and technical excellence, come join us to shape the future of edge AI.#LI-HybridYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until August 2, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, CA, RemoteType: Full time
$184k - $287.5k
We are now seeking a Senior Infrastructure Software Engineer for NVIDIA TensorRT Edge-LLM!NVIDIA's TensorRT Infrastructure group is seeking excellent software engineers to enable the next generation of edge AI. This is an outstanding chance to define the infrastructure...SeniorFull time$227k - $300k
...transformation to AI-enabled software-defined vehicles.... ...production-grade AI on the Edge. We are looking for a great Senior Staff AI Engineer to join our seasoned AI... ...(e.g., Transformers, LLM, CNN, LSTM, Trees) to... ...Experience with NVIDIA TensorRT, Qualcomm SNPE.Sunnyvale...SeniorWork at officeWorldwideFlexible hoursShift work3 days per week$224k - $356.5k
...Team is building the software foundation for... ...looking for exceptional engineers who thrive on... ....We are seeking a Senior Software Engineer... ...learning inference, TensorRT and related... ...TensorRT, TensorRT-LLM, ONNX, PyTorch, CUDA... ...models on embedded, edge, robotics, or automotive...SeniorFull time$152k - $241.5k
We are now looking for a Senior Software Engineer for Deep Learning Inference! Would you like to make a big impact in Deep Learning by helping... ...Senior Engineering positions in the Deep Learning Inference TensorRT software team.What you’ll be doing:Craft and develop robust...SeniorFull time$152k - $241.5k
...a Developer Technology Engineer, you will be at the forefront... ...AI workflows at the edge powered by NVIDIAs RTX... ...performance.Improve LLM & GenAI user experience... ...performance enhancements of OSS software, including but not... ...and SDKs, specifically TensorRT-RTX, cuDNN, NVIDIA...SeniorFull timeLocal area$124k - $195.5k
...We are now looking for a Deep Learning Software Engineer, TensorRT Performance! NVIDIA is seeking an... ...accelerators, from datacenter GPUs to edge SoCs. Implement graph compiler algorithms... ...inference libraries (e.g. TensorRT, TensorRT-LLM, vLLM, SGLang, FlashInfer). ~...Remote work- ...ID: JR2015071 Job Category: Engineering Time Type: Full time... ...accelerated deep learning inference software like TensorRT, DL benchmarking software... ...libraries (e.g. TensorRT, TensorRT-LLM, vLLM, SGLang, FlashInfer).... ....g. Jetson systems or other edge AI accelerators). GPU deep...Full time
$132k - $165k
...shape the future of cybersecurity.RoleWe are looking for a Senior Software Development Engineer-AI Security to join our team. This is a Hybrid role based... ...as virtual memory, multi-threading, system APIs, SLM/LLM models, and excellent debugging and problem-solving skills...SeniorFull timeWork at officeLocal area$272k - $431.25k
...resilient deployment of cutting-edge LLM workloads.We are seeking a Principal Systems Engineer to define the vision and... ...engines (such as vLLM, SGLang, TensorRT-LLM), with a focus on KV-cache... ...accelerators and memory pools.Mentor senior and junior engineers, set technical...Full timeLocal areaRemote work$160k - $200k
Join Fortinet as a Senior Software Developer and play a pivotal role in the entire software... ...features. You will utilize cutting-edge GenAI/LLM technologies to enhance our next-generation... ...of professional software engineering practices, including version control,...SeniorFull timeWorldwide- ...About the Opportunity We're looking for a Software Development Engineer to help build a modern AI-powered... ...systems while building and optimizing LLM-powered workflows for production use.... ...Why Join? Opportunity to build cutting-edge AI-powered products from an early stage...SeniorFull time
$152k - $241.5k
...seeking talented and motivated engineers to join our TensorRT team in developing the industry... ...leading deep learning inference software for NVIDIA AI accelerators. As a Senior Software Engineer in the... ...optimize NVIDIA TensorRT and TensorRT-LLM to supercharge inference...SeniorFull time$184k - $287.5k
...application is built. We are seeking a Senior Software Engineer focused on container and cloud infrastructure... ..., including support for disaggregated LLM inference and other emerging deployment... ...multi-tenant, multi-cluster, or edge/air-gapped container delivery.Contributions...SeniorFull time$184k - $287.5k
...is building an RL Frameworks engineering team to develop the open-source... ...on. The team spans the full software stack, from collaborating... ...areas:Reinforcement learning for LLM post-training (RLHF, PPO,... ...inference engines (vLLM, SGLang, TensorRT-LLM) into RL training loops...SeniorFull time$184k - $287.5k
...for a motivated Deep Learning engineer to bring advanced CUDA... ...stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You... ...systems principles (aka systems software fundamentals)Adaptability and... ...frameworks (e.g., PyTorch, JAX, TensorRT, vLLM, sgLang, Nemo, Megatron...SeniorFull time$152k - $241.5k
Are you passionate about redefining how software is built in the age of Generative AI? Join NVIDIA’s TensorRT team to help lead a first-of-its-kind, AI-native initiative... ...scale.If you are a systems-thinking C++ engineer who wants to help scale out an agentic development...SeniorFull time$152k - $241.5k
...the digital landscape.Local AI seeks a Senior Systems Software Engineer interested in solving client-side AI... ...graphics, web browsers, and edge devices—by driving innovation in both... ...APIs like ONNX RT, DirectX, PyTorch, TensorRT, Vulkan, llama.cpp.We're a top employer...SeniorFull timeLocal area$174k - $252k
...with model researchers, software and hardware teams to... ...Large Language Model (LLM) inference latency and... ...Computer Science, Electrical Engineering, Computer Engineering,... ...by combining cutting-edge technology, infrastructure... ...learning models.As a Senior Performance Co-Design...SeniorWorldwide$136k - $218.5k
...are building a next-generation software platform for semiconductor development... ...to process graphs at a trillion-edge scale. This is an ambitious endeavor, and we need engineers who thrive in a hands-on... ...or agentic tooling, particularly LLM-based code generation.Familiarity...SeniorFull time$224k - $356.5k
...visibility and real-world impact. As a System Software Engineer for Vision AI, you will develop and... ...video, image, and 3D data in both edge and cloud settings.Developing multi-modal... ...experience with GPU acceleration (such as CUDA, TensorRT, or comparable technologies) and low-...SeniorFull time$130k - $180k
...forefront of the AI-powered data engineering revolution. You can read more... ...On: Backend Integration of LLM Architectures : Lead the development... ...: Push the boundaries of software engineering by combining traditional... ...techniques with cutting-edge AI technologies. High Visibility...SeniorWorldwide$185k - $230k
The OpportunityWe are looking for a Senior AI Agent & LLM Engineer who combines strong software engineering capabilities with a deep focus on AI quality. You will help build and improve the AI systems behind Otter’s conversational knowledge engine and AI Chat, spanning...SeniorPermanent employment$200k - $322k
...We are looking for a Senior Technical Marketing Engineer focused on Enterprise AI Software, and accelerating adoption... ...adopt cutting-edge technology, At NVIDIA,... ...AI, RAG, agentic AI, LLM-based applications, inference... ..., NIM, NeMo, TensorRT, Triton Inference Server...SeniorFull time$262k - $365k
...to test/evaluate/deploy across devices (Edge Portal /Model Explorer/Developer Device... ...practical experience.8 years of experience in software development.7 years of experience... ...qualifications:Master’s degree or PhD in Engineering, Computer Science, or a related technical...Senior$193.3k - $261.5k
...builds AWS Neuron, the software development kit used to... ...software boundary, our engineers build systematic infrastructure... ...of a wide variety of LLM model families,... ...sharing and mentorship. Our senior members enjoy one-on-... ...with vLLM, SGLang, TensorRT or similar platforms in...SeniorWork experience placementInternshipLocal areaFlexible hours$152k - $241.5k
...passionate about driving innovation in deep learning and eager to work on cutting-edge AI technology for safety-critical applications? Join NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of technology, enabling high-performance AI inference...SeniorFull time$193.13k - $257.5k
...collaboration, and high standards. Our engineers, product leaders, and go-to-... ..., build, and deploy cutting-edge deep learning models across... ...proven by a track record of software artifacts or academic... ...inference optimization (vLLM, TensorRT-LLM).Desired Skills & Experience:...Work experience placementWork at officeRemote workFlexible hours3 days per week$184k - $287.5k
...AI revolution, building the software and systems that power the... ...workloads. We are looking for a Senior Software Engineer to lead the bring-up,... ...to ensure state-of-the-art LLM workloads run efficiently and... ...PyTorch, NeMo / Megatron, TensorRT-LLM, and adjacent NVIDIA AI...SeniorFull timeRemote work$224k - $356.5k
...Local AI team is building the software stack that makes large language... ...maximum efficiency on NVIDIA edge AI hardware. The AI ecosystem... ...innovations in leading open-source LLM inference frameworks —... ...in Computer Science, Computer Engineering, Electrical Engineering, or equivalent...SeniorFull timeLocal area$174.72k - $295.68k
...transportation through cutting-edge R&D in AI, machine learning,... ...connectivity.You will be a senior engineer on the team building our internal... ...our engineers build and ship software. A core mission is connecting... ...modern AI coding tooling and LLM APIs, including tool use/...SeniorFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Software Engineer - TensorRT Edge-LLM. Be the first to apply!
- ngo software engineer Santa Clara, CA
- software data engineer Santa Clara, CA
- graduate software engineer Santa Clara, CA
- software system engineer Santa Clara, CA
- graduate software developer Santa Clara, CA
- software engineer - early career Santa Clara, CA
- entry level software engineer remote Santa Clara, CA
- software engineer intern Santa Clara, CA
- software developer positions Santa Clara, CA
- junior software developer internship Santa Clara, CA

