Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Edge AI Engineer: LLM Inference (TensorRT)

NVIDIA

NVIDIA Corporation is seeking a Senior Software Engineer for the TensorRT Edge-LLM team in the US. You will develop a high-performance inference framework in modern C++ that extends TensorRT for autoregressive model serving, including speculative decoding and KV cache management, within embedded and edge platforms. You will collaborate across CUDA and robotics teams, optimize transformer components, and contribute to kernel development while staying ahead of LLM/VLM trends. #J-18808-Ljbffr NVIDIA

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Senior Edge AI Engineer: LLM Inference (TensorRT) in Santa Clara, CA vacancy
  • $152k - $241.5k

     ...limits of real-time large language model inference? Join NVIDIA’s TensorRT Edge-LLM team and help shape the next generation of edge AI for automotive and robotics. We build the...  ...in Computer Science, Electrical/Computer Engineering, or a closely related field.4+ years of relevant... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

    We are now seeking a Senior Infrastructure Software Engineer for NVIDIA TensorRT Edge-LLM!NVIDIA's TensorRT Infrastructure group is...  ...enable the next generation of edge AI. This is an outstanding chance...  ...AGX for robotics and edge inference applications. You will work with... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • NVIDIA in Santa Clara, CA seeks a senior software engineer focusing on GPU computing and ML inference to optimize LLM workloads on edge AI hardware. You will track open-source inference frameworks, map architectures to NVIDIA GPUs, and report performance metrics. With... 
    Senior

    NVIDIA

    Santa Clara, CA
    4 days ago
  • $227k - $300k

     ...transformation to AI-enabled software-defined...  ...-grade AI on the Edge. We are looking for a great Senior Staff AI Engineer to join our...  ...g., Transformers, LLM, CNN, LSTM, Trees)...  ...C++ (C++14/17 for inference).Deep proficiency...  ...Experience with NVIDIA TensorRT, Qualcomm SNPE.... 
    Senior
    Work at office
    Worldwide
    Flexible hours
    Shift work
    3 days per week

    Sonatus

    Sunnyvale, CA
    1 day ago
  • Intel in Santa Clara, CA, seeks an experienced software engineer to optimize local inference for edge devices. You will work on llama.cpp, vLLM, and...  ...source engines while advancing hardware-aware optimizations for efficient AI on local devices. #J-18808-Ljbffr Intel
    Senior
    Local area

    Intel

    Santa Clara, CA
    10 hours ago
  • $169.78k - $338.69k

     ...experienced and mission-driven Senior/Staff AI Infrastructure Engineer, Inference & Optimization to lead the...  ...scheduling of cutting-edge AI models across on-vehicle...  ...mainstream inference engines (e.g., TensorRT, ONNX Runtime) and specialized LLM inference/serving frameworks... 
    Senior
    Full time

    DiDi Labs

    San Jose, CA
    2 days ago
  • NVIDIA seeks a Senior Systems Software Engineer to tackle client-side AI challenges on Windows and Linux PCs with limited...  ...AI models, data pipelines, and inference runtimes for performance on next...  ...experience with ONNX RT, PyTorch, TensorRT, llama.cpp, and vLLM. #J-18808-... 
    Senior
    Local area

    NVIDIA

    Santa Clara, CA
    2 days ago
  • NVIDIA seeks a Senior Software Engineer for Local AI on Windows and Linux PCs. You will partner with software...  ...on end-to-end optimization of inference runtimes and edge deployment. The role requires a...  ...using ONNX RT, PyTorch, and TensorRT, plus excellent communication.... 
    Senior
    Local area

    Nvidia Corporation in

    Santa Clara, CA
    1 day ago
  •  ...Santa Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design and...  ..., build new abstractions for LLM serving engines, and contribute to...  ...on CUDA C/C++, Triton, and cutting-edge MLIR-based tooling, and participate... 
    Senior

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $151.8k - $332.2k

     ...expect We are looking for an AI Inference Engineer with a solid background in...  ...will work on the most cutting edge speech modeling and...  ...speech recognition, speech-llm or AI model inference.Display...  ...accelerating AI models using CUDA, TensorRT, and mixed-precision computation... 
    Full time
    Work at office
    Remote work

    Zoom

    San Jose, CA
    1 day ago
  •  ...potential of generative AI to power the...  ...sits at the leading edge of what’s possible with LLM inference on heterogeneous hardware...  ...research and engineering team that moves fast...  ...frameworks (vLLM, SGLang, TensorRT-LLM, etc.) and can...  ...hardware.• Small, senior team with high... 
    Senior

    d-Matrix

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

     ...software is built in the age of Generative AI? Join NVIDIA’s TensorRT team to help lead a first-of-its-kind...  ...entry point for out-of-framework inference globally. We are moving beyond...  ...scale.If you are a systems-thinking C++ engineer who wants to help scale out an agentic... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    19 hours ago
  • $152k - $241.5k

    We are now looking for a Senior Software Engineer for Deep Learning Inference! Would you like to make a big impact in Deep...  ...in the Deep Learning Inference TensorRT software team.What you’ll be doing...  ...an existing vacancy. NVIDIA uses AI tools in its recruiting processes.... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

    NVIDIA is seeking an experienced AI systems engineer to innovate and develop cutting-edge technologies in AI inference systems. You will design and optimize kernel technologies to accelerate workloads for NVIDIA's hardware architecture. The ideal candidate holds a Master... 
    Senior

    NVIDIA

    Santa Clara, CA
    1 day ago
  •  ...Corporation in Santa Clara, CA is seeking outstanding AI systems engineers to develop groundbreaking inference technologies for the hardware-accelerated stack. You...  ..., code generators, and GPU kernel innovations for LLM workloads. Join a team that designs extensible abstractions... 
    Senior

    NVIDIA Corporation

    Santa Clara, CA
    3 days ago
  • NVIDIA is seeking a Senior Software Engineer for Deep Learning Inference to help build a state-of-the-art inference framework...  ...models. You will join the TensorRT Workflows team and tackle scalable...  ...and ML researchers to push TensorRT forward. #J-18808-Ljbffr NVIDIA AI
    Senior

    NVIDIA AI

    Santa Clara, CA
    4 days ago
  •  ...Corporation in Santa Clara, CA seeks a Senior Software Engineer to advance Deep Learning Inference within TensorRT. You will build scalable...  ...C++, Python, and cutting-edge NVIDIA GPUs. This role offers impact...  ...opportunity to shape next-generation AI acceleration. #J-18808-Ljbffr... 
    Senior

    NVIDIA Corporation

    Santa Clara, CA
    3 days ago
  • $182.5k - $260.5k

     ...for the cloud and AI era. We secure and...  ...platform, its Zero Trust Engine, and the powerful...  ...are available at Senior Staff and above....  ..., you own the inference and optimization layer...  ...AI.Cutting-edge, unusual stack. The...  ...runtimes (vLLM/SGLang, TensorRT-LLM, ONNX Runtime,... 
    Senior

    Netskope

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...unlimited potential of AI to define the next...  ...Technology Engineer, you will be at the...  ...AI workflows at the edge powered by NVIDIAs...  ...performance.Improve LLM & GenAI user experience...  ...GPU-accelerated AI inference driven by NVIDIA APIs...  ...SDKs, specifically TensorRT-RTX, cuDNN, NVIDIA... 
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

     ...are looking for a software engineer with a strong background in...  ...performance at the intersection of AI, high-performance computing,...  ...the crowd: Experience with inference optimization techniques and...  ...production.Experience with TensorRT, TensorRT-LLM, and cuTile.Experience... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    19 hours ago
  • $152k - $241.5k

     ...driving advancements in AI and machine learning...  ...and motivated engineers to join our TensorRT team in developing the...  ...leading deep learning inference software for NVIDIA AI accelerators. As a Senior Software Engineer in...  ...TensorRT and TensorRT-LLM to supercharge inference... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $272k - $431.25k

     ...throughput, low-latency inference framework for serving generative AI and reasoning...  ...deployment of cutting-edge LLM workloads.We are...  ...Principal Systems Engineer to define the vision...  ...such as vLLM, SGLang, TensorRT-LLM), with a focus...  ...pools.Mentor senior and junior engineers... 
    Full time
    Local area
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $160k - $198k

     ...members.What You’ll DoAs a Senior AI Systems Engineer, you will architect, deploy,...  ...scale AI model training and inference. You will ensure our machine...  ...scheduling alongside advanced LLM serving engines.Cross-...  ...Engineers to productionize cutting-edge models, establish monitoring... 
    Senior
    Local area

    Archer Aviation

    San Jose, CA
    3 days ago
  • $200k - $322k

    NVIDIA is seeking a dedicated Senior AI Product Engineer to join our ambitious AI...  ...identify and resolve rough edges before customer impact, and...  ...Demonstrated experience in shipping LLM-backed features, including...  ..., distillation, or serving/inference optimization.Demonstrated... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • Cerebras Systems, Inc. is looking for a Senior Performance Engineer to enhance the performance...  ...competitive pricing models for their AI chip. The ideal candidate will have extensive experience with open-source inference frameworks and an understanding of ML... 
    Senior

    Cerebras Systems, Inc.

    Sunnyvale, CA
    10 hours ago
  • $193.3k - $261.5k

     ...enabling unparalleled ML inference and training...  ...boundary, our engineers build systematic infrastructure...  ...'s possible in AI acceleration.As...  ...a wide variety of LLM model families,...  ...mentorship. Our senior members enjoy one-...  ...with vLLM, SGLang, TensorRT or similar platforms... 
    Senior
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $229.9k - $262.4k

     ...Overview Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform) Overview: At Capital One, we are creating responsible...  ..., and more. ~ Invent and introduce state-of-the-art LLM optimization techniques to improve the performance — scalability... 
    Senior
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    more than 2 months ago
  •  ...patients worldwide.We’re a team of engineers, clinicians, and innovators united...  ...robotic platforms. As a Senior AI/ML Research Engineer (Computer Vision...  ...active-learning loops.Real-time / edge inference optimization (e.g., TensorRT, NVIDIA Jetson).Fine-grained interaction... 
    Senior
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    Sunnyvale, CA
    4 days ago
  • $152k - $241.5k

     ...deep learning and eager to work on cutting-edge AI technology for safety-critical applications? Join NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of technology, enabling high-performance AI inference solutions for automotive safety and... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • Lendistry is seeking a Senior AI Engineer to lead the delivery of AI strategies, focusing on end-to-end LLM features including document intelligence and risk assessment workflows. This role involves collaborating with senior leaders and mentoring junior engineers, ensuring... 
    Senior

    Lendistry

    Santa Clara, CA
    10 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Edge AI Engineer: LLM Inference (TensorRT). Be the first to apply!