Senior Software Engineer - TensorRT Edge-LLM
$152k - $241.5kNVIDIA
Are you passionate about pushing the limits of real-time large language model inference? Join NVIDIA’s TensorRT Edge-LLM team and help shape the next generation of edge AI for automotive and robotics. We build the software stack that enables Large Language, Vision-Language, and Multimodal (LLM/VLM/VLA) models to run efficiently on embedded and edge platforms — delivering cutting-edge generative AI experiences directly on-device.What you’ll be doing:Develop and evolve a state-of-the-art inference framework in modern C++ that extends TensorRT with autoregressive model serving capabilities, including speculative decoding, LoRA, MoE, and KV cache management.Design and implement compiler and runtime optimizations tailored for transformer-based models running on constrained, real-time platforms.Collaborate with teams across CUDA, kernel libraries, compilers, and robotics to deliver high-performance, production-ready solutions.Contribute to CUDA kernel and operator development for critical transformer components such as attention, GEMM, and MoE.Benchmark, profile, and optimize inference performance across diverse embedded and automotive environments.Stay ahead of the rapidly evolving LLM/VLM ecosystem and bring emerging techniques into product-grade software.What we need to see:BS, MS, PhD, or equivalent experience in Computer Science, Electrical/Computer Engineering, or a closely related field.4+ years of relevant software development experience.Deep understanding of transformer models and inference optimization techniques (e.g., quantization, tensor parallelism, or memory-efficient scheduling).Proficient programming ability with modern C++ (C++11/14/17 and beyond).Familiarity with popular LLM frameworks and libraries such as TensorRT, TensorRT-LLM, vLLM, SGLang, MLC-LLM, or FlashInfer.A track record of strong software design, execution, and collaboration across fields.Ways to stand out from the crowd:Demonstrated development experience or open-source contributions to LLM inference frameworks and libraries, such as SGLang, vLLM, or FlashInfer.Proficiency with CUDA, including efficient kernel development, performance profiling, and GPU architecture fundamentals.Prior work on autoregressive LLM serving systems, including speculative decoding or KV cache management.Familiarity with compiler infrastructure for large language model inference.Exposure to robotics or embedded AI pipelines, including optimizing for low-latency, resource-constrained systems.NVIDIA is widely considered to be one of the technology world’s most desirable employers. We hire some of the most brilliant and forward-thinking people in the world. If you thrive on innovation, autonomy, and technical excellence, come join us to shape the future of edge AI.#LI-HybridYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until August 9, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, CA, RemoteType: Full time
$184k - $287.5k
We are now seeking a Senior Infrastructure Software Engineer for NVIDIA TensorRT Edge-LLM!NVIDIA's TensorRT Infrastructure group is seeking excellent software engineers to enable the next generation of edge AI. This is an outstanding chance to define the infrastructure...SeniorFull time- ...inference framework in modern C++ that extends TensorRT with autoregressive model serving... ...a relevant field and at least 4 years of software development experience. A deep understanding... ...are essential. Key Skills C++, TensorRT, LLM, VLM, GEMM, CUDA, Attention, MoE, KV...Senior
- NVIDIA Corporation is seeking a Senior Software Engineer for the TensorRT Edge-LLM team in the US. You will develop a high-performance inference framework in modern C++ that extends TensorRT for autoregressive model serving, including speculative decoding and KV cache...Senior
$227k - $300k
...transformation to AI-enabled software-defined vehicles.... ...production-grade AI on the Edge. We are looking for a great Senior Staff AI Engineer to join our seasoned AI... ...(e.g., Transformers, LLM, CNN, LSTM, Trees) to... ...Experience with NVIDIA TensorRT, Qualcomm SNPE.Sunnyvale...SeniorWork at officeWorldwideFlexible hoursShift work3 days per week$224k - $356.5k
...Team is building the software foundation for... ...looking for exceptional engineers who thrive on... ....We are seeking a Senior Software Engineer... ...learning inference, TensorRT and related... ...TensorRT, TensorRT-LLM, ONNX, PyTorch, CUDA... ...models on embedded, edge, robotics, or automotive...SeniorFull time$152k - $241.5k
We are now looking for a Senior Software Engineer for Deep Learning Inference! Would you like to make a big impact in Deep Learning by helping... ...Senior Engineering positions in the Deep Learning Inference TensorRT software team.What you’ll be doing:Craft and develop robust...SeniorFull time$152k - $241.5k
...a Developer Technology Engineer, you will be at the forefront... ...AI workflows at the edge powered by NVIDIAs RTX... ...performance.Improve LLM & GenAI user experience... ...performance enhancements of OSS software, including but not... ...and SDKs, specifically TensorRT-RTX, cuDNN, NVIDIA...SeniorFull timeLocal area- ...seeking an experienced Deep Learning Software Engineer, TensorRT Performance to analyze and improve the... ...inference libraries (TensorRT, TensorRT‑LLM, vLLM, SGLang, FlashInfer). Experience... ...pipelines (e.g. Jetson systems, other edge AI accelerators). Compensation Base...
$132k - $165k
...shape the future of cybersecurity.RoleWe are looking for a Senior Software Development Engineer-AI Security to join our team. This is a Hybrid role based... ...as virtual memory, multi-threading, system APIs, SLM/LLM models, and excellent debugging and problem-solving skills...SeniorFull timeWork at officeLocal area$272k - $431.25k
...resilient deployment of cutting-edge LLM workloads.We are seeking a Principal Systems Engineer to define the vision and... ...engines (such as vLLM, SGLang, TensorRT-LLM), with a focus on KV-cache... ...accelerators and memory pools.Mentor senior and junior engineers, set technical...Full timeLocal areaRemote work- PACCAR Silicon Valley Innovation Center in Sunnyvale, California, is seeking a Senior Software Engineer to develop and deploy software for truck vehicle technologies, including edge and cloud components. You will guide external resources, maintain code quality, and advance...Senior
$160k - $200k
Join Fortinet as a Senior Software Developer and play a pivotal role in the entire software... ...features. You will utilize cutting-edge GenAI/LLM technologies to enhance our next-generation... ...of professional software engineering practices, including version control,...SeniorFull timeWorldwide- ...Senior Software Engineer In Test We are hiring a hands-on Senior Software Engineer in Test to design... ...Biases, or equivalent) • Experience with LLM or computer vision evaluation is a plus... .../recall, etc.) • Robustness, edge cases, and failure modes • Data quality...SeniorShift work
$152k - $241.5k
...seeking talented and motivated engineers to join our TensorRT team in developing the industry... ...leading deep learning inference software for NVIDIA AI accelerators. As a Senior Software Engineer in the... ...optimize NVIDIA TensorRT and TensorRT-LLM to supercharge inference...SeniorFull time$130k - $180k
...forefront of the AI-powered data engineering revolution. You can read more... ...On: Backend Integration of LLM Architectures : Lead the development... ...: Push the boundaries of software engineering by combining traditional... ...techniques with cutting‑edge AI technologies. High Visibility...SeniorWorldwide$139k - $204k
...What You’ll Do Senior engineers are area owners who lead designs, raise engineering standards, and deliver measurable improvements to latency... ...Contributions to inference frameworks (vLLM, Triton, TensorRT‑LLM, Ray Serve, TorchServe). Experience with CUDA kernels, NCCL/...SeniorPermanent employmentTemporary workCasual workWork at officeRemote workFlexible hoursShift work$136k - $218.5k
...are building a next-generation software platform for semiconductor development... ...to process graphs at a trillion-edge scale. This is an ambitious endeavor, and we need engineers who thrive in a hands-on... ...or agentic tooling, particularly LLM-based code generation.Familiarity...SeniorFull time$152k - $241.5k
Are you passionate about redefining how software is built in the age of Generative AI? Join NVIDIA’s TensorRT team to help lead a first-of-its-kind, AI-native initiative... ...scale.If you are a systems-thinking C++ engineer who wants to help scale out an agentic development...SeniorFull time$184k - $287.5k
...for a motivated Deep Learning engineer to bring advanced CUDA... ...stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You... ...systems principles (aka systems software fundamentals)Adaptability and... ...frameworks (e.g., PyTorch, JAX, TensorRT, vLLM, sgLang, Nemo, Megatron...SeniorFull time$224k - $356.5k
...visibility and real-world impact. As a System Software Engineer for Vision AI, you will develop and... ...video, image, and 3D data in both edge and cloud settings.Developing multi-modal... ...experience with GPU acceleration (such as CUDA, TensorRT, or comparable technologies) and low-...SeniorFull time$185k - $230k
The OpportunityWe are looking for a Senior AI Agent & LLM Engineer who combines strong software engineering capabilities with a deep focus on AI quality. You will help build and improve the AI systems behind Otter’s conversational knowledge engine and AI Chat, spanning...SeniorPermanent employment$200k - $322k
...We are looking for a Senior Technical Marketing Engineer focused on Enterprise AI Software, and accelerating adoption... ...adopt cutting-edge technology, At NVIDIA,... ...AI, RAG, agentic AI, LLM-based applications, inference... ..., NIM, NeMo, TensorRT, Triton Inference Server...SeniorFull time$262k - $365k
...to test/evaluate/deploy across devices (Edge Portal /Model Explorer/Developer Device... ...practical experience.8 years of experience in software development.7 years of experience... ...qualifications:Master’s degree or PhD in Engineering, Computer Science, or a related technical...Senior$182k - $242k
...5. Learn more at About this role We're looking for a Senior Engineer for CoreWeave's Benchmarking & Performance team. You will... ..., NVLink/PCIe, memory bandwidth) or model-serving stacks (llm-d, vLLM, TensorRT-LLM, Megatron-LM). ~ Effective communicator comfortable...SeniorPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$165k - $242k
Apply for the Senior Software Engineer II, Inference role at CoreWeave. CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers... ...Contributions to inference frameworks (vLLM, Triton, TensorRT‑LLM, Ray Serve, TorchServe). Experience with CUDA kernels, NCCL...SeniorPermanent employmentTemporary workCasual workWork at officeRemote workFlexible hoursShift work$152k - $241.5k
...passionate about driving innovation in deep learning and eager to work on cutting-edge AI technology for safety-critical applications? Join NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of technology, enabling high-performance AI inference...SeniorFull time- ...business successful through AI by combining cutting-edge technology, infrastructure, and talent. The TPU... ...models, including LLMs, on our own hardware. As a Senior Performance Co-design Engineer, you will focus on LLM Serving Studies, analyzing and optimizing serving performance...Senior
$193.3k - $261.5k
...builds Amazon Neuron, the software development kit used to... ...software boundary, our engineers build systematic... ...tuning of a wide variety of LLM model families,... ...sharing and mentorship. Our senior members enjoy one-on-one... ...serving with vLLM, SGLang, TensorRT or similar platforms in...SeniorWork experience placementInternshipLocal areaFlexible hours$130k - $180k
...forefront of the AI-powered data engineering revolution. You can read more... ...On: Backend Integration of LLM Architectures : Lead the development... ...: Push the boundaries of software engineering by combining... ...traditional techniques with cutting‑edge AI technologies. High...SeniorWorldwide$182k - $242k
...inference. Our stack is engineered for speed, scale, and... ...product. We’re looking for a Senior Engineer for CoreWeave’... ...the critical path of LLM inference. Optimize... ...-serving stacks (vLLM, TensorRT-LLM, llm-d, SGLang). Implement... ..., GPU/accelerator software, or performance-...SeniorPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Software Engineer - TensorRT Edge-LLM. Be the first to apply!
- cybersecurity software engineer Santa Clara, CA
- graduate software engineer Santa Clara, CA
- software developer fintech Santa Clara, CA
- new graduate software engineer Santa Clara, CA
- senior robotics software engineer Santa Clara, CA
- software engineer visa sponsorship Santa Clara, CA
- software qa engineer Santa Clara, CA
- network software engineer Santa Clara, CA
- software engineer remote Santa Clara, CA
- part time software developer remote Santa Clara, CA


