Senior Edge AI Engineer: LLM Inference (TensorRT)
NVIDIA
NVIDIA Corporation is seeking a Senior Software Engineer for the TensorRT Edge-LLM team in the US. You will develop a high-performance inference framework in modern C++ that extends TensorRT for autoregressive model serving, including speculative decoding and KV cache management, within embedded and edge platforms. You will collaborate across CUDA and robotics teams, optimize transformer components, and contribute to kernel development while staying ahead of LLM/VLM trends. #J-18808-Ljbffr NVIDIA
$152k - $241.5k
...limits of real-time large language model inference? Join NVIDIA’s TensorRT Edge-LLM team and help shape the next generation of edge AI for automotive and robotics. We build the... ...in Computer Science, Electrical/Computer Engineering, or a closely related field.4+ years of relevant...SeniorFull time$184k - $287.5k
We are now seeking a Senior Infrastructure Software Engineer for NVIDIA TensorRT Edge-LLM!NVIDIA's TensorRT Infrastructure group is... ...enable the next generation of edge AI. This is an outstanding chance... ...AGX for robotics and edge inference applications. You will work with...SeniorFull time- NVIDIA in Santa Clara, CA seeks a senior software engineer focusing on GPU computing and ML inference to optimize LLM workloads on edge AI hardware. You will track open-source inference frameworks, map architectures to NVIDIA GPUs, and report performance metrics. With...Senior
$227k - $300k
...transformation to AI-enabled software-defined... ...-grade AI on the Edge. We are looking for a great Senior Staff AI Engineer to join our... ...g., Transformers, LLM, CNN, LSTM, Trees)... ...C++ (C++14/17 for inference).Deep proficiency... ...Experience with NVIDIA TensorRT, Qualcomm SNPE....SeniorWork at officeWorldwideFlexible hoursShift work3 days per week- Intel in Santa Clara, CA, seeks an experienced software engineer to optimize local inference for edge devices. You will work on llama.cpp, vLLM, and... ...source engines while advancing hardware-aware optimizations for efficient AI on local devices. #J-18808-Ljbffr IntelSeniorLocal area
$169.78k - $338.69k
...experienced and mission-driven Senior/Staff AI Infrastructure Engineer, Inference & Optimization to lead the... ...scheduling of cutting-edge AI models across on-vehicle... ...mainstream inference engines (e.g., TensorRT, ONNX Runtime) and specialized LLM inference/serving frameworks...SeniorFull time- NVIDIA seeks a Senior Systems Software Engineer to tackle client-side AI challenges on Windows and Linux PCs with limited... ...AI models, data pipelines, and inference runtimes for performance on next... ...experience with ONNX RT, PyTorch, TensorRT, llama.cpp, and vLLM. #J-18808-...SeniorLocal area
- NVIDIA seeks a Senior Software Engineer for Local AI on Windows and Linux PCs. You will partner with software... ...on end-to-end optimization of inference runtimes and edge deployment. The role requires a... ...using ONNX RT, PyTorch, and TensorRT, plus excellent communication....SeniorLocal area
- ...Santa Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design and... ..., build new abstractions for LLM serving engines, and contribute to... ...on CUDA C/C++, Triton, and cutting-edge MLIR-based tooling, and participate...Senior
$151.8k - $332.2k
...expect We are looking for an AI Inference Engineer with a solid background in... ...will work on the most cutting edge speech modeling and... ...speech recognition, speech-llm or AI model inference.Display... ...accelerating AI models using CUDA, TensorRT, and mixed-precision computation...Full timeWork at officeRemote work- ...potential of generative AI to power the... ...sits at the leading edge of what’s possible with LLM inference on heterogeneous hardware... ...research and engineering team that moves fast... ...frameworks (vLLM, SGLang, TensorRT-LLM, etc.) and can... ...hardware.• Small, senior team with high...Senior
$152k - $241.5k
...software is built in the age of Generative AI? Join NVIDIA’s TensorRT team to help lead a first-of-its-kind... ...entry point for out-of-framework inference globally. We are moving beyond... ...scale.If you are a systems-thinking C++ engineer who wants to help scale out an agentic...SeniorFull time$152k - $241.5k
We are now looking for a Senior Software Engineer for Deep Learning Inference! Would you like to make a big impact in Deep... ...in the Deep Learning Inference TensorRT software team.What you’ll be doing... ...an existing vacancy. NVIDIA uses AI tools in its recruiting processes....SeniorFull time$184k - $287.5k
NVIDIA is seeking an experienced AI systems engineer to innovate and develop cutting-edge technologies in AI inference systems. You will design and optimize kernel technologies to accelerate workloads for NVIDIA's hardware architecture. The ideal candidate holds a Master...Senior- ...Corporation in Santa Clara, CA is seeking outstanding AI systems engineers to develop groundbreaking inference technologies for the hardware-accelerated stack. You... ..., code generators, and GPU kernel innovations for LLM workloads. Join a team that designs extensible abstractions...Senior
- NVIDIA is seeking a Senior Software Engineer for Deep Learning Inference to help build a state-of-the-art inference framework... ...models. You will join the TensorRT Workflows team and tackle scalable... ...and ML researchers to push TensorRT forward. #J-18808-Ljbffr NVIDIA AISenior
- ...Corporation in Santa Clara, CA seeks a Senior Software Engineer to advance Deep Learning Inference within TensorRT. You will build scalable... ...C++, Python, and cutting-edge NVIDIA GPUs. This role offers impact... ...opportunity to shape next-generation AI acceleration. #J-18808-Ljbffr...Senior
$182.5k - $260.5k
...for the cloud and AI era. We secure and... ...platform, its Zero Trust Engine, and the powerful... ...are available at Senior Staff and above.... ..., you own the inference and optimization layer... ...AI.Cutting-edge, unusual stack. The... ...runtimes (vLLM/SGLang, TensorRT-LLM, ONNX Runtime,...Senior$152k - $241.5k
...unlimited potential of AI to define the next... ...Technology Engineer, you will be at the... ...AI workflows at the edge powered by NVIDIAs... ...performance.Improve LLM & GenAI user experience... ...GPU-accelerated AI inference driven by NVIDIA APIs... ...SDKs, specifically TensorRT-RTX, cuDNN, NVIDIA...SeniorFull timeLocal area$152k - $241.5k
...are looking for a software engineer with a strong background in... ...performance at the intersection of AI, high-performance computing,... ...the crowd: Experience with inference optimization techniques and... ...production.Experience with TensorRT, TensorRT-LLM, and cuTile.Experience...SeniorFull time$152k - $241.5k
...driving advancements in AI and machine learning... ...and motivated engineers to join our TensorRT team in developing the... ...leading deep learning inference software for NVIDIA AI accelerators. As a Senior Software Engineer in... ...TensorRT and TensorRT-LLM to supercharge inference...SeniorFull time$272k - $431.25k
...throughput, low-latency inference framework for serving generative AI and reasoning... ...deployment of cutting-edge LLM workloads.We are... ...Principal Systems Engineer to define the vision... ...such as vLLM, SGLang, TensorRT-LLM), with a focus... ...pools.Mentor senior and junior engineers...Full timeLocal areaRemote work$160k - $198k
...members.What You’ll DoAs a Senior AI Systems Engineer, you will architect, deploy,... ...scale AI model training and inference. You will ensure our machine... ...scheduling alongside advanced LLM serving engines.Cross-... ...Engineers to productionize cutting-edge models, establish monitoring...SeniorLocal area$200k - $322k
NVIDIA is seeking a dedicated Senior AI Product Engineer to join our ambitious AI... ...identify and resolve rough edges before customer impact, and... ...Demonstrated experience in shipping LLM-backed features, including... ..., distillation, or serving/inference optimization.Demonstrated...SeniorFull time- Cerebras Systems, Inc. is looking for a Senior Performance Engineer to enhance the performance... ...competitive pricing models for their AI chip. The ideal candidate will have extensive experience with open-source inference frameworks and an understanding of ML...Senior
$193.3k - $261.5k
...enabling unparalleled ML inference and training... ...boundary, our engineers build systematic infrastructure... ...'s possible in AI acceleration.As... ...a wide variety of LLM model families,... ...mentorship. Our senior members enjoy one-... ...with vLLM, SGLang, TensorRT or similar platforms...SeniorWork experience placementInternshipLocal areaFlexible hours$229.9k - $262.4k
...Overview Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform) Overview: At Capital One, we are creating responsible... ..., and more. ~ Invent and introduce state-of-the-art LLM optimization techniques to improve the performance — scalability...SeniorFull timePart timeLocal area- ...patients worldwide.We’re a team of engineers, clinicians, and innovators united... ...robotic platforms. As a Senior AI/ML Research Engineer (Computer Vision... ...active-learning loops.Real-time / edge inference optimization (e.g., TensorRT, NVIDIA Jetson).Fine-grained interaction...SeniorLocal areaWorldwideFlexible hours
$152k - $241.5k
...deep learning and eager to work on cutting-edge AI technology for safety-critical applications? Join NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of technology, enabling high-performance AI inference solutions for automotive safety and...SeniorFull time- Lendistry is seeking a Senior AI Engineer to lead the delivery of AI strategies, focusing on end-to-end LLM features including document intelligence and risk assessment workflows. This role involves collaborating with senior leaders and mentoring junior engineers, ensuring...Senior
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Edge AI Engineer: LLM Inference (TensorRT). Be the first to apply!
- ai engineer Santa Clara, CA
- senior ai engineer Santa Clara, CA
- ai prompt engineer Santa Clara, CA
- ai engineer remote Santa Clara, CA
- ai developer Santa Clara, CA
- senior technology project manager Santa Clara, CA
- senior c++ developer Santa Clara, CA
- remote senior business analyst Santa Clara, CA
- senior director fp&a Santa Clara, CA
- senior manager clinical operations Santa Clara, CA

