Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Edge AI Engineer: LLM Inference (TensorRT)

NVIDIA Corporation

NVIDIA Corporation is seeking a Senior Software Engineer for the TensorRT Edge-LLM team in the US. You will develop a high-performance inference framework in modern C++ that extends TensorRT for autoregressive model serving, including speculative decoding and KV cache management, within embedded and edge platforms.

You will collaborate across CUDA and robotics teams, optimize transformer components, and contribute to kernel development while staying ahead of LLM/VLM trends.

#J-18808-Ljbffr
Vacancy posted 4 hours ago
Similar jobs that could be interesting for youBased on the Senior Edge AI Engineer: LLM Inference (TensorRT) in Santa Clara, CA vacancy
  • $152k - $241.5k

     ...limits of real-time large language model inference? Join NVIDIA’s TensorRT Edge-LLM team and help shape the next generation of edge AI for automotive and robotics. We build the...  ...in Computer Science, Electrical/Computer Engineering, or a closely related field.4+ years of relevant... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

    We are now seeking a Senior Infrastructure Software Engineer for NVIDIA TensorRT Edge-LLM!NVIDIA's TensorRT Infrastructure group is...  ...enable the next generation of edge AI. This is an outstanding chance...  ...AGX for robotics and edge inference applications. You will work with... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $227k - $300k

     ...transformation to AI-enabled software-defined...  ...-grade AI on the Edge. We are looking for a great Senior Staff AI Engineer to join our...  ...g., Transformers, LLM, CNN, LSTM, Trees)...  ...C++ (C++14/17 for inference).Deep proficiency...  ...Experience with NVIDIA TensorRT, Qualcomm SNPE.... 
    Senior
    Work at office
    Worldwide
    Flexible hours
    Shift work
    3 days per week

    Sonatus

    Sunnyvale, CA
    2 days ago
  • Intel in Santa Clara, CA, seeks an experienced software engineer to optimize local inference for edge devices. You will work on llama.cpp, vLLM, and...  ...source engines while advancing hardware-aware optimizations for efficient AI on local devices. #J-18808-Ljbffr Intel
    Senior
    Local area

    Intel

    Santa Clara, CA
    6 days ago
  • Intel is seeking a seasoned software engineer to accelerate AI inference on edge hardware. You will optimize llama.cpp/vLLM, tune KV cache, batching and scheduling, and push quantization strategies to balance speed and quality. This role focuses on low-latency, privacy-... 
    Senior

    PVH (Tommy Hilfiger/Calvin Klein)

    Santa Clara, CA
    6 days ago
  •  ...NVIDIA seeks a Senior Systems Software Engineer to tackle client-side AI challenges on Windows and Linux PCs with limited...  ...AI models, data pipelines, and inference runtimes for performance on next...  ...experience with ONNX RT, PyTorch, TensorRT, llama.cpp, and vLLM. #J-18808... 
    Senior
    Local area

    NVIDIA

    Santa Clara, CA
    4 hours ago
  •  ...Santa Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design and...  ..., build new abstractions for LLM serving engines, and contribute to...  ...on CUDA C/C++, Triton, and cutting-edge MLIR-based tooling, and participate... 
    Senior

    NVIDIA

    Santa Clara, CA
    4 hours ago
  • $151.8k - $332.2k

     ...expect We are looking for an AI Inference Engineer with a solid background in...  ...will work on the most cutting edge speech modeling and...  ...speech recognition, speech-llm or AI model inference.Display...  ...accelerating AI models using CUDA, TensorRT, and mixed-precision computation... 
    Full time
    Work at office
    Remote work

    Zoom

    San Jose, CA
    2 days ago
  •  ...potential of generative AI to power the...  ...sits at the leading edge of what’s possible with LLM inference on heterogeneous hardware...  ...research and engineering team that moves fast...  ...frameworks (vLLM, SGLang, TensorRT-LLM, etc.) and can...  ...hardware.• Small, senior team with high... 
    Senior

    d-Matrix

    Santa Clara, CA
    5 days ago
  • $152k - $241.5k

     ...software is built in the age of Generative AI? Join NVIDIA’s TensorRT team to help lead a first-of-its-kind...  ...entry point for out-of-framework inference globally. We are moving beyond...  ...scale.If you are a systems-thinking C++ engineer who wants to help scale out an agentic... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

    We are now looking for a Senior Software Engineer for Deep Learning Inference! Would you like to make a big impact in Deep...  ...in the Deep Learning Inference TensorRT software team.What you’ll be doing...  ...an existing vacancy. NVIDIA uses AI tools in its recruiting processes.... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...AMD is seeking a Senior Physical AI Software Engineer in San Jose, CA to develop next-generation software for intelligent edge systems that blend embedded computing, real-time processing...  ...and RTOS platforms, integrating AI inference, computer vision, and sensor processing... 
    Senior

    AMD

    San Jose, CA
    4 hours ago
  •  ...NVIDIA is seeking a Senior Software Engineer for Deep Learning Inference to help build a state-of-the-art inference framework on NVIDIA GPUs, accelerating large language models. You will join the TensorRT Workflows team and tackle scalable, real-time inferencing challenges... 
    Senior

    NVIDIA AI

    Santa Clara, CA
    4 hours ago
  •  ...Corporation in Santa Clara, CA is seeking outstanding AI systems engineers to develop groundbreaking inference technologies for the hardware-accelerated stack. You...  ..., code generators, and GPU kernel innovations for LLM workloads. Join a team that designs extensible abstractions... 
    Senior

    NVIDIA Corporation

    Santa Clara, CA
    4 days ago
  •  ...NVIDIA is seeking a Senior Agentic AI Software Engineer (Finance) to build agentic systems and advance AI inference workloads in real-world finance contexts. You will design and implement...  ...performance, and contribute to cutting-edge AI research at scale. This role... 
    Senior

    Nvidia Corporation in

    Santa Clara, CA
    4 hours ago
  •  ...NVIDIA is seeking a Senior Agentic AI Software Engineer to advance agentic AI systems and workloads from scalable...  ...build agentic components, analyze inference dynamics, and collaborate with teams...  ...agents is essential, with a strong edge in inference or agent architectures... 
    Senior

    NVIDIA

    Santa Clara, CA
    4 hours ago
  • $250k - $300k

     ...the only vertically integrated AI infrastructure company built...  ...production. That means owning the inference stack end to end: profiling...  ...also work directly with customer engineering teams to tailor deployments to...  .... Comfort with modern LLM serving frameworks such as vLLM... 
    Senior
    Temporary work

    Crusoe

    Sunnyvale, CA
    15 days ago
  •  ...Corporation in Santa Clara, CA seeks a Senior Software Engineer to advance Deep Learning Inference within TensorRT. You will build scalable...  ...C++, Python, and cutting-edge NVIDIA GPUs. This role offers impact...  ...opportunity to shape next-generation AI acceleration. #J-18808-Ljbffr
    Senior

    NVIDIA Corporation

    Santa Clara, CA
    5 hours ago
  • $182.5k - $260.5k

     ...for the cloud and AI era. We secure and...  ...platform, its Zero Trust Engine, and the powerful...  ...are available at Senior Staff and above....  ..., you own the inference and optimization layer...  ...AI.Cutting-edge, unusual stack. The...  ...runtimes (vLLM/SGLang, TensorRT-LLM, ONNX Runtime,... 
    Senior

    Netskope

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

     ...are looking for a software engineer with a strong background in...  ...performance at the intersection of AI, high-performance computing,...  ...the crowd: Experience with inference optimization techniques and...  ...production.Experience with TensorRT, TensorRT-LLM, and cuTile.Experience... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

     ...unlimited potential of AI to define the next...  ...Technology Engineer, you will be at the...  ...AI workflows at the edge powered by NVIDIAs...  ...performance.Improve LLM & GenAI user experience...  ...GPU-accelerated AI inference driven by NVIDIA APIs...  ...SDKs, specifically TensorRT-RTX, cuDNN, NVIDIA... 
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

     ...driving advancements in AI and machine learning...  ...and motivated engineers to join our TensorRT team in developing the...  ...leading deep learning inference software for NVIDIA AI accelerators. As a Senior Software Engineer in...  ...TensorRT and TensorRT-LLM to supercharge inference... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $272k - $431.25k

     ...throughput, low-latency inference framework for serving generative AI and reasoning...  ...deployment of cutting-edge LLM workloads.We are...  ...Principal Systems Engineer to define the vision...  ...such as vLLM, SGLang, TensorRT-LLM), with a focus...  ...pools.Mentor senior and junior engineers... 
    Full time
    Local area
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $160k - $198k

     ...members.What You’ll DoAs a Senior AI Systems Engineer, you will architect, deploy,...  ...scale AI model training and inference. You will ensure our machine...  ...scheduling alongside advanced LLM serving engines.Cross-...  ...Engineers to productionize cutting-edge models, establish monitoring... 
    Senior
    Local area

    Archer Aviation

    San Jose, CA
    4 days ago
  • $200k - $322k

    NVIDIA is seeking a dedicated Senior AI Product Engineer to join our ambitious AI...  ...identify and resolve rough edges before customer impact, and...  ...Demonstrated experience in shipping LLM-backed features, including...  ..., distillation, or serving/inference optimization.Demonstrated... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • Cerebras Systems, Inc. is looking for a Senior Performance Engineer to enhance the performance...  ...competitive pricing models for their AI chip. The ideal candidate will have extensive experience with open-source inference frameworks and an understanding of ML... 
    Senior

    Cerebras Systems, Inc.

    Sunnyvale, CA
    5 days ago
  • $193.3k - $261.5k

     ...enabling unparalleled ML inference and training...  ...boundary, our engineers build systematic infrastructure...  ...'s possible in AI acceleration.As...  ...a wide variety of LLM model families,...  ...mentorship. Our senior members enjoy one-...  ...with vLLM, SGLang, TensorRT or similar platforms... 
    Senior
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    15 hours ago
  •  ...patients worldwide.We’re a team of engineers, clinicians, and innovators united...  ...robotic platforms. As a Senior AI/ML Research Engineer (Computer Vision...  ...active-learning loops.Real-time / edge inference optimization (e.g., TensorRT, NVIDIA Jetson).Fine-grained interaction... 
    Senior
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    Sunnyvale, CA
    5 days ago
  • $189k - $301k

     ...Conductor in San Jose, CA is seeking a seasoned engineer to lead co-design efforts for optimizing AI model inference performance. The role requires a deep understanding of AI infrastructure, covering everything from model definition to serving. The ideal candidate... 
    Senior

    Conductor

    San Jose, CA
    4 hours ago
  • $152k - $241.5k

     ...deep learning and eager to work on cutting-edge AI technology for safety-critical applications? Join NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of technology, enabling high-performance AI inference solutions for automotive safety and... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Edge AI Engineer: LLM Inference (TensorRT). Be the first to apply!