Senior Edge AI Engineer: LLM Inference (TensorRT)
NVIDIA Corporation
NVIDIA Corporation is seeking a Senior Software Engineer for the TensorRT Edge-LLM team in the US. You will develop a high-performance inference framework in modern C++ that extends TensorRT for autoregressive model serving, including speculative decoding and KV cache management, within embedded and edge platforms.
You will collaborate across CUDA and robotics teams, optimize transformer components, and contribute to kernel development while staying ahead of LLM/VLM trends.
#J-18808-Ljbffr$152k - $241.5k
...limits of real-time large language model inference? Join NVIDIA’s TensorRT Edge-LLM team and help shape the next generation of edge AI for automotive and robotics. We build the... ...in Computer Science, Electrical/Computer Engineering, or a closely related field.4+ years of relevant...SeniorFull time$184k - $287.5k
We are now seeking a Senior Infrastructure Software Engineer for NVIDIA TensorRT Edge-LLM!NVIDIA's TensorRT Infrastructure group is... ...enable the next generation of edge AI. This is an outstanding chance... ...AGX for robotics and edge inference applications. You will work with...SeniorFull time$227k - $300k
...transformation to AI-enabled software-defined... ...-grade AI on the Edge. We are looking for a great Senior Staff AI Engineer to join our... ...g., Transformers, LLM, CNN, LSTM, Trees)... ...C++ (C++14/17 for inference).Deep proficiency... ...Experience with NVIDIA TensorRT, Qualcomm SNPE....SeniorWork at officeWorldwideFlexible hoursShift work3 days per week- Intel in Santa Clara, CA, seeks an experienced software engineer to optimize local inference for edge devices. You will work on llama.cpp, vLLM, and... ...source engines while advancing hardware-aware optimizations for efficient AI on local devices. #J-18808-Ljbffr IntelSeniorLocal area
- Intel is seeking a seasoned software engineer to accelerate AI inference on edge hardware. You will optimize llama.cpp/vLLM, tune KV cache, batching and scheduling, and push quantization strategies to balance speed and quality. This role focuses on low-latency, privacy-...Senior
- ...NVIDIA seeks a Senior Systems Software Engineer to tackle client-side AI challenges on Windows and Linux PCs with limited... ...AI models, data pipelines, and inference runtimes for performance on next... ...experience with ONNX RT, PyTorch, TensorRT, llama.cpp, and vLLM. #J-18808...SeniorLocal area
- ...Santa Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design and... ..., build new abstractions for LLM serving engines, and contribute to... ...on CUDA C/C++, Triton, and cutting-edge MLIR-based tooling, and participate...Senior
$151.8k - $332.2k
...expect We are looking for an AI Inference Engineer with a solid background in... ...will work on the most cutting edge speech modeling and... ...speech recognition, speech-llm or AI model inference.Display... ...accelerating AI models using CUDA, TensorRT, and mixed-precision computation...Full timeWork at officeRemote work- ...potential of generative AI to power the... ...sits at the leading edge of what’s possible with LLM inference on heterogeneous hardware... ...research and engineering team that moves fast... ...frameworks (vLLM, SGLang, TensorRT-LLM, etc.) and can... ...hardware.• Small, senior team with high...Senior
$152k - $241.5k
...software is built in the age of Generative AI? Join NVIDIA’s TensorRT team to help lead a first-of-its-kind... ...entry point for out-of-framework inference globally. We are moving beyond... ...scale.If you are a systems-thinking C++ engineer who wants to help scale out an agentic...SeniorFull time$152k - $241.5k
We are now looking for a Senior Software Engineer for Deep Learning Inference! Would you like to make a big impact in Deep... ...in the Deep Learning Inference TensorRT software team.What you’ll be doing... ...an existing vacancy. NVIDIA uses AI tools in its recruiting processes....SeniorFull time- ...AMD is seeking a Senior Physical AI Software Engineer in San Jose, CA to develop next-generation software for intelligent edge systems that blend embedded computing, real-time processing... ...and RTOS platforms, integrating AI inference, computer vision, and sensor processing...Senior
- ...NVIDIA is seeking a Senior Software Engineer for Deep Learning Inference to help build a state-of-the-art inference framework on NVIDIA GPUs, accelerating large language models. You will join the TensorRT Workflows team and tackle scalable, real-time inferencing challenges...Senior
- ...Corporation in Santa Clara, CA is seeking outstanding AI systems engineers to develop groundbreaking inference technologies for the hardware-accelerated stack. You... ..., code generators, and GPU kernel innovations for LLM workloads. Join a team that designs extensible abstractions...Senior
- ...NVIDIA is seeking a Senior Agentic AI Software Engineer (Finance) to build agentic systems and advance AI inference workloads in real-world finance contexts. You will design and implement... ...performance, and contribute to cutting-edge AI research at scale. This role...Senior
- ...NVIDIA is seeking a Senior Agentic AI Software Engineer to advance agentic AI systems and workloads from scalable... ...build agentic components, analyze inference dynamics, and collaborate with teams... ...agents is essential, with a strong edge in inference or agent architectures...Senior
$250k - $300k
...the only vertically integrated AI infrastructure company built... ...production. That means owning the inference stack end to end: profiling... ...also work directly with customer engineering teams to tailor deployments to... .... Comfort with modern LLM serving frameworks such as vLLM...SeniorTemporary work- ...Corporation in Santa Clara, CA seeks a Senior Software Engineer to advance Deep Learning Inference within TensorRT. You will build scalable... ...C++, Python, and cutting-edge NVIDIA GPUs. This role offers impact... ...opportunity to shape next-generation AI acceleration. #J-18808-LjbffrSenior
$182.5k - $260.5k
...for the cloud and AI era. We secure and... ...platform, its Zero Trust Engine, and the powerful... ...are available at Senior Staff and above.... ..., you own the inference and optimization layer... ...AI.Cutting-edge, unusual stack. The... ...runtimes (vLLM/SGLang, TensorRT-LLM, ONNX Runtime,...Senior$152k - $241.5k
...are looking for a software engineer with a strong background in... ...performance at the intersection of AI, high-performance computing,... ...the crowd: Experience with inference optimization techniques and... ...production.Experience with TensorRT, TensorRT-LLM, and cuTile.Experience...SeniorFull time$152k - $241.5k
...unlimited potential of AI to define the next... ...Technology Engineer, you will be at the... ...AI workflows at the edge powered by NVIDIAs... ...performance.Improve LLM & GenAI user experience... ...GPU-accelerated AI inference driven by NVIDIA APIs... ...SDKs, specifically TensorRT-RTX, cuDNN, NVIDIA...SeniorFull timeLocal area$152k - $241.5k
...driving advancements in AI and machine learning... ...and motivated engineers to join our TensorRT team in developing the... ...leading deep learning inference software for NVIDIA AI accelerators. As a Senior Software Engineer in... ...TensorRT and TensorRT-LLM to supercharge inference...SeniorFull time$272k - $431.25k
...throughput, low-latency inference framework for serving generative AI and reasoning... ...deployment of cutting-edge LLM workloads.We are... ...Principal Systems Engineer to define the vision... ...such as vLLM, SGLang, TensorRT-LLM), with a focus... ...pools.Mentor senior and junior engineers...Full timeLocal areaRemote work$160k - $198k
...members.What You’ll DoAs a Senior AI Systems Engineer, you will architect, deploy,... ...scale AI model training and inference. You will ensure our machine... ...scheduling alongside advanced LLM serving engines.Cross-... ...Engineers to productionize cutting-edge models, establish monitoring...SeniorLocal area$200k - $322k
NVIDIA is seeking a dedicated Senior AI Product Engineer to join our ambitious AI... ...identify and resolve rough edges before customer impact, and... ...Demonstrated experience in shipping LLM-backed features, including... ..., distillation, or serving/inference optimization.Demonstrated...SeniorFull time- Cerebras Systems, Inc. is looking for a Senior Performance Engineer to enhance the performance... ...competitive pricing models for their AI chip. The ideal candidate will have extensive experience with open-source inference frameworks and an understanding of ML...Senior
$193.3k - $261.5k
...enabling unparalleled ML inference and training... ...boundary, our engineers build systematic infrastructure... ...'s possible in AI acceleration.As... ...a wide variety of LLM model families,... ...mentorship. Our senior members enjoy one-... ...with vLLM, SGLang, TensorRT or similar platforms...SeniorWork experience placementInternshipLocal areaFlexible hours- ...patients worldwide.We’re a team of engineers, clinicians, and innovators united... ...robotic platforms. As a Senior AI/ML Research Engineer (Computer Vision... ...active-learning loops.Real-time / edge inference optimization (e.g., TensorRT, NVIDIA Jetson).Fine-grained interaction...SeniorLocal areaWorldwideFlexible hours
$189k - $301k
...Conductor in San Jose, CA is seeking a seasoned engineer to lead co-design efforts for optimizing AI model inference performance. The role requires a deep understanding of AI infrastructure, covering everything from model definition to serving. The ideal candidate...Senior$152k - $241.5k
...deep learning and eager to work on cutting-edge AI technology for safety-critical applications? Join NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of technology, enabling high-performance AI inference solutions for automotive safety and...SeniorFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Edge AI Engineer: LLM Inference (TensorRT). Be the first to apply!
- ai developer Santa Clara, CA
- ai engineer Santa Clara, CA
- ai prompt engineer Santa Clara, CA
- senior ai engineer Santa Clara, CA
- ai engineer remote Santa Clara, CA
- senior lead project manager Santa Clara, CA
- senior robotics software engineer Santa Clara, CA
- senior devops engineer remote Santa Clara, CA
- senior sas administrator Santa Clara, CA
- senior IT manager Santa Clara, CA


