Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Software Engineer - TensorRT Edge-LLM

$152k - $241.5k

NVIDIA

Are you passionate about pushing the limits of real-time large language model inference? Join NVIDIA’s TensorRT Edge-LLM team and help shape the next generation of edge AI for automotive and robotics. We build the software stack that enables Large Language, Vision-Language, and Multimodal (LLM/VLM/VLA) models to run efficiently on embedded and edge platforms — delivering cutting-edge generative AI experiences directly on-device.What you’ll be doing:Develop and evolve a state-of-the-art inference framework in modern C++ that extends TensorRT with autoregressive model serving capabilities, including speculative decoding, LoRA, MoE, and KV cache management.Design and implement compiler and runtime optimizations tailored for transformer-based models running on constrained, real-time platforms.Collaborate with teams across CUDA, kernel libraries, compilers, and robotics to deliver high-performance, production-ready solutions.Contribute to CUDA kernel and operator development for critical transformer components such as attention, GEMM, and MoE.Benchmark, profile, and optimize inference performance across diverse embedded and automotive environments.Stay ahead of the rapidly evolving LLM/VLM ecosystem and bring emerging techniques into product-grade software.What we need to see:BS, MS, PhD, or equivalent experience in Computer Science, Electrical/Computer Engineering, or a closely related field.4+ years of relevant software development experience.Deep understanding of transformer models and inference optimization techniques (e.g., quantization, tensor parallelism, or memory-efficient scheduling).Proficient programming ability with modern C++ (C++11/14/17 and beyond).Familiarity with popular LLM frameworks and libraries such as TensorRT, TensorRT-LLM, vLLM, SGLang, MLC-LLM, or FlashInfer.A track record of strong software design, execution, and collaboration across fields.Ways to stand out from the crowd:Demonstrated development experience or open-source contributions to LLM inference frameworks and libraries, such as SGLang, vLLM, or FlashInfer.Proficiency with CUDA, including efficient kernel development, performance profiling, and GPU architecture fundamentals.Prior work on autoregressive LLM serving systems, including speculative decoding or KV cache management.Familiarity with compiler infrastructure for large language model inference.Exposure to robotics or embedded AI pipelines, including optimizing for low-latency, resource-constrained systems.NVIDIA is widely considered to be one of the technology world’s most desirable employers. We hire some of the most brilliant and forward-thinking people in the world. If you thrive on innovation, autonomy, and technical excellence, come join us to shape the future of edge AI.#LI-HybridYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until August 2, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, CA, RemoteType: Full time

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Senior Software Engineer - TensorRT Edge-LLM in Santa Clara, CA vacancy
  • $184k - $287.5k

    We are now seeking a Senior Infrastructure Software Engineer for NVIDIA TensorRT Edge-LLM!NVIDIA's TensorRT Infrastructure group is seeking excellent software engineers to enable the next generation of edge AI. This is an outstanding chance to define the infrastructure... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    7 hours ago
  • $227k - $300k

     ...transformation to AI-enabled software-defined vehicles....  ...production-grade AI on the Edge. We are looking for a great Senior Staff AI Engineer to join our seasoned AI...  ...(e.g., Transformers, LLM, CNN, LSTM, Trees) to...  ...Experience with NVIDIA TensorRT, Qualcomm SNPE.Sunnyvale... 
    Senior
    Work at office
    Worldwide
    Flexible hours
    Shift work
    3 days per week

    Sonatus

    Sunnyvale, CA
    4 days ago
  • $224k - $356.5k

     ...Team is building the software foundation for...  ...looking for exceptional engineers who thrive on...  ....We are seeking a Senior Software Engineer...  ...learning inference, TensorRT and related...  ...TensorRT, TensorRT-LLM, ONNX, PyTorch, CUDA...  ...models on embedded, edge, robotics, or automotive... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

    We are now looking for a Senior Software Engineer for Deep Learning Inference! Would you like to make a big impact in Deep Learning by helping...  ...Senior Engineering positions in the Deep Learning Inference TensorRT software team.What you’ll be doing:Craft and develop robust... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

     ...a Developer Technology Engineer, you will be at the forefront...  ...AI workflows at the edge powered by NVIDIAs RTX...  ...performance.Improve LLM & GenAI user experience...  ...performance enhancements of OSS software, including but not...  ...and SDKs, specifically TensorRT-RTX, cuDNN, NVIDIA... 
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    1 day ago
  • $124k - $195.5k

     ...We are now looking for a Deep Learning Software Engineer, TensorRT Performance! NVIDIA is seeking an...  ...accelerators, from datacenter GPUs to edge SoCs. Implement graph compiler algorithms...  ...inference libraries (e.g. TensorRT, TensorRT-LLM, vLLM, SGLang, FlashInfer). ~... 
    Remote work

    NVIDIA

    Santa Clara, CA
    5 days ago
  •  ...ID: JR2015071 Job Category: Engineering Time Type: Full time...  ...accelerated deep learning inference software like TensorRT, DL benchmarking software...  ...libraries (e.g. TensorRT, TensorRT-LLM, vLLM, SGLang, FlashInfer)....  ....g. Jetson systems or other edge AI accelerators). GPU deep... 
    Full time

    NVIDIA AI

    Santa Clara, CA
    4 days ago
  • $132k - $165k

     ...shape the future of cybersecurity.RoleWe are looking for a Senior Software Development Engineer-AI Security to join our team. This is a Hybrid role based...  ...as virtual memory, multi-threading, system APIs, SLM/LLM models, and excellent debugging and problem-solving skills... 
    Senior
    Full time
    Work at office
    Local area

    Zscaler

    San Jose, CA
    4 days ago
  • $272k - $431.25k

     ...resilient deployment of cutting-edge LLM workloads.We are seeking a Principal Systems Engineer to define the vision and...  ...engines (such as vLLM, SGLang, TensorRT-LLM), with a focus on KV-cache...  ...accelerators and memory pools.Mentor senior and junior engineers, set technical... 
    Full time
    Local area
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $160k - $200k

    Join Fortinet as a Senior Software Developer and play a pivotal role in the entire software...  ...features. You will utilize cutting-edge GenAI/LLM technologies to enhance our next-generation...  ...of professional software engineering practices, including version control,... 
    Senior
    Full time
    Worldwide

    Fortinet

    Sunnyvale, CA
    4 days ago
  •  ...About the Opportunity We're looking for a Software Development Engineer to help build a modern AI-powered...  ...systems while building and optimizing LLM-powered workflows for production use....  ...Why Join? Opportunity to build cutting-edge AI-powered products from an early stage... 
    Senior
    Full time

    Petals Careers Private Limited

    Sunnyvale, CA
    3 days ago
  • $152k - $241.5k

     ...seeking talented and motivated engineers to join our TensorRT team in developing the industry...  ...leading deep learning inference software for NVIDIA AI accelerators. As a Senior Software Engineer in the...  ...optimize NVIDIA TensorRT and TensorRT-LLM to supercharge inference... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...application is built. We are seeking a Senior Software Engineer focused on container and cloud infrastructure...  ..., including support for disaggregated LLM inference and other emerging deployment...  ...multi-tenant, multi-cluster, or edge/air-gapped container delivery.Contributions... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...is building an RL Frameworks engineering team to develop the open-source...  ...on. The team spans the full software stack, from collaborating...  ...areas:Reinforcement learning for LLM post-training (RLHF, PPO,...  ...inference engines (vLLM, SGLang, TensorRT-LLM) into RL training loops... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...for a motivated Deep Learning engineer to bring advanced CUDA...  ...stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You...  ...systems principles (aka systems software fundamentals)Adaptability and...  ...frameworks (e.g., PyTorch, JAX, TensorRT, vLLM, sgLang, Nemo, Megatron... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

    Are you passionate about redefining how software is built in the age of Generative AI? Join NVIDIA’s TensorRT team to help lead a first-of-its-kind, AI-native initiative...  ...scale.If you are a systems-thinking C++ engineer who wants to help scale out an agentic development... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

     ...the digital landscape.Local AI seeks a Senior Systems Software Engineer interested in solving client-side AI...  ...graphics, web browsers, and edge devices—by driving innovation in both...  ...APIs like ONNX RT, DirectX, PyTorch, TensorRT, Vulkan, llama.cpp.We're a top employer... 
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    4 days ago
  • $174k - $252k

     ...with model researchers, software and hardware teams to...  ...Large Language Model (LLM) inference latency and...  ...Computer Science, Electrical Engineering, Computer Engineering,...  ...by combining cutting-edge technology, infrastructure...  ...learning models.As a Senior Performance Co-Design... 
    Senior
    Worldwide

    Google

    Sunnyvale, CA
    2 days ago
  • $136k - $218.5k

     ...are building a next-generation software platform for semiconductor development...  ...to process graphs at a trillion-edge scale. This is an ambitious endeavor, and we need engineers who thrive in a hands-on...  ...or agentic tooling, particularly LLM-based code generation.Familiarity... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $224k - $356.5k

     ...visibility and real-world impact. As a System Software Engineer for Vision AI, you will develop and...  ...video, image, and 3D data in both edge and cloud settings.Developing multi-modal...  ...experience with GPU acceleration (such as CUDA, TensorRT, or comparable technologies) and low-... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $130k - $180k

     ...forefront of the AI-powered data engineering revolution. You can read more...  ...On: Backend Integration of LLM Architectures : Lead the development...  ...: Push the boundaries of software engineering by combining traditional...  ...techniques with cutting-edge AI technologies. High Visibility... 
    Senior
    Worldwide

    PA Early Stage Partners

    Sunnyvale, CA
    3 days ago
  • $185k - $230k

    The OpportunityWe are looking for a Senior AI Agent & LLM Engineer who combines strong software engineering capabilities with a deep focus on AI quality. You will help build and improve the AI systems behind Otter’s conversational knowledge engine and AI Chat, spanning... 
    Senior
    Permanent employment

    Otter.ai

    Mountain View, CA
    1 day ago
  • $200k - $322k

     ...We are looking for a Senior Technical Marketing Engineer focused on Enterprise AI Software, and accelerating adoption...  ...adopt cutting-edge technology, At NVIDIA,...  ...AI, RAG, agentic AI, LLM-based applications, inference...  ..., NIM, NeMo, TensorRT, Triton Inference Server... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $262k - $365k

     ...to test/evaluate/deploy across devices (Edge Portal /Model Explorer/Developer Device...  ...practical experience.8 years of experience in software development.7 years of experience...  ...qualifications:Master’s degree or PhD in Engineering, Computer Science, or a related technical... 
    Senior

    Google

    Sunnyvale, CA
    1 day ago
  • $193.3k - $261.5k

     ...builds AWS Neuron, the software development kit used to...  ...software boundary, our engineers build systematic infrastructure...  ...of a wide variety of LLM model families,...  ...sharing and mentorship. Our senior members enjoy one-on-...  ...with vLLM, SGLang, TensorRT or similar platforms in... 
    Senior
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    4 days ago
  • $152k - $241.5k

     ...passionate about driving innovation in deep learning and eager to work on cutting-edge AI technology for safety-critical applications? Join NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of technology, enabling high-performance AI inference... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $193.13k - $257.5k

     ...collaboration, and high standards. Our engineers, product leaders, and go-to-...  ..., build, and deploy cutting-edge deep learning models across...  ...proven by a track record of software artifacts or academic...  ...inference optimization (vLLM, TensorRT-LLM).Desired Skills & Experience:... 
    Work experience placement
    Work at office
    Remote work
    Flexible hours
    3 days per week

    Eightfold

    Santa Clara, CA
    7 hours ago
  • $184k - $287.5k

     ...AI revolution, building the software and systems that power the...  ...workloads. We are looking for a Senior Software Engineer to lead the bring-up,...  ...to ensure state-of-the-art LLM workloads run efficiently and...  ...PyTorch, NeMo / Megatron, TensorRT-LLM, and adjacent NVIDIA AI... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $224k - $356.5k

     ...Local AI team is building the software stack that makes large language...  ...maximum efficiency on NVIDIA edge AI hardware. The AI ecosystem...  ...innovations in leading open-source LLM inference frameworks —...  ...in Computer Science, Computer Engineering, Electrical Engineering, or equivalent... 
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    3 days ago
  • $174.72k - $295.68k

     ...transportation through cutting-edge R&D in AI, machine learning,...  ...connectivity.You will be a senior engineer on the team building our internal...  ...our engineers build and ship software. A core mission is connecting...  ...modern AI coding tooling and LLM APIs, including tool use/... 
    Senior
    Full time

    XPENG Motors

    Santa Clara, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Software Engineer - TensorRT Edge-LLM. Be the first to apply!