Edge AI LLM Inference Architect (C++, TensorRT)
NVIDIA AI
NVIDIA AI in Santa Clara is seeking a highly capable software engineer to advance an advanced inference framework using modern C++. The role focuses on extending TensorRT with autoregressive model serving capabilities and requires collaboration across CUDA, kernel libraries, compilers, and robotics teams to deliver high-performance, production-ready solutions. The candidate should hold a BS/MS/PhD (or equivalent) and have at least four years of software development experience with a deep #J-18808-Ljbffr NVIDIA AI
- Develop and evolve a state-of-the-art inference framework in modern C++ that extends TensorRT with autoregressive model serving... .... Key Skills C++, TensorRT, LLM, VLM, GEMM, CUDA, Attention, MoE,... ...Infrastructure, Robotics, Embedded AI Benefits Equity, Health Insurance...C++
- Job Description: Edge AI Architect - CUDA / C++ / Computer Vision Experience Level: 10+ Years Department: Edge AI & Embedded Systems About the Role... ...optimization tools. AI/ML Acumen (Crucial) Required: Strong LLM, RAG, agentic architecture understanding. Understanding...C++Contract work
$174.05k - $278.45k
Distinguished Technologist, Edge AI Architect Description - Job Summary The... ...security and OS trust foundation, inference serving, agentic runtimes and... ...Certification (Python, C++, Rust, Java, or similar). Cloud... ...a plus. Knowledge & Skills LLM, vLM, and multi-modal model...C++Full timeContract workTemporary workWork experience placementLocal areaFlexible hoursShift work$272k - $431.25k
...group is solving some of AI’s hardest... ...interconnects.This Principal Architect role leads the research... ...such as vLLM, SGLang, and TensorRT-LLM.Publishing findings, representing... ...training and inference patterns.Proficiency in... ...programming languages such as C, C++, Rust and Python.Ways...C++Full timeRemote work$272k - $431.25k
...throughput, low-latency inference framework for serving generative AI and reasoning... ...deployment of cutting-edge LLM workloads.We are... ...scale LLM inference.Architect and implement deep... ...such as vLLM, SGLang, TensorRT-LLM), with a focus... ...in C/C++ and Python, with a...C++Full timeLocal areaRemote work- ...potential of generative AI to power the... ...sits at the leading edge of what’s possible with LLM inference on heterogeneous hardware... ...(vLLM, SGLang, TensorRT-LLM, etc.) and can extend... ...with hardware architects to provide firmware... ...proficiency in Python and C/C++.• Hands-on...C++
- codvo-team is seeking an Edge AI Architect to lead design and deployment of AI models on edge devices for medical applications. The role focuses... ...candidate has 10+ years in AI engineering, hands-on CUDA/C++ experience, and familiarity with OpenCV, TensorFlow, and PyTorch...C++Remote work
$152k - $241.5k
...unlimited potential of AI to define the next... ...workflows at the edge powered by NVIDIAs... ...performance.Improve LLM & GenAI user... ...Strong proficiency in C/C++, Python, software design... ...GPU-accelerated AI inference driven by NVIDIA... ...SDKs, specifically TensorRT-RTX, cuDNN, NVIDIA...C++Full timeLocal area$151.8k - $332.2k
...We are looking for an AI Inference Engineer with a solid background... ...on the most cutting edge speech modeling and... ...recognition, speech-llm or AI model inference.Display... ..., shell scripts, C/C++; familiarity with ML frameworks... ...AI models using CUDA, TensorRT, and mixed-precision...C++Full timeWork at officeRemote work$182.5k - $260.5k
...networking for the cloud and AI era. We secure and... ...Scientist, you own the inference and optimization layer... ...of agentic AI.Cutting-edge, unusual stack. The hard... ...runtimes (vLLM/SGLang, TensorRT-LLM, ONNX Runtime, llama.cpp... ...comfort reaching into C++ for low-level interop is...C++$197.5k - $272k
...the transformation to AI-enabled software-defined... ...-grade AI on the Edge. We are looking for a... ...(e.g., Transformers, LLM, CNN, LSTM, Trees) to... ...working knowledge of modern C++ (C++14/17 for inference).Deep proficiency with... ...Experience with NVIDIA TensorRT, Qualcomm SNPE....C++Work at officeWorldwideFlexible hoursShift work3 days per week$124k - $195.5k
...Software Engineer, TensorRT PerformanceNVIDIA is... ...of NVIDIA's inference ecosystem! NVIDIA is... ...areas like Generative AI, Recommenders and Vision... ...datacenter GPUs to edge SoCs. Implement... ...experience.Strong C++, Python programming... ...TensorRT, TensorRT-LLM, vLLM, SGLang, FlashInfer...C++$152k - $241.5k
...Engineer, Deep Learning Inference -... ...Deep Learning Inference TensorRT software team.**What you... ...feature update TensorRT* Use C++ and Python to build graph... ...learning experts, GPU architects and DevOps engineers across... ...existing vacancy.NVIDIA uses AI tools in its recruiting...C++$197.3k - $225.1k
Lead AI Engineer (FM Hosting, LLM Inference) Overview At Capital One, we are creating responsible and reliable AI systems, changing banking for good.... ...Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, or Golang ~ Experience developing and applying...C++Full timePart timeLocal area$184k - $287.5k
...unlimited potential of AI to define the... ...Communication Architect. We scale the... ...models and training/inference frameworks to... ...and optimizing LLM training and inference... ...on cutting-edge hardware.Deep understanding... ...as PyTorch, TensorRT-LLM, vLLM,... ...skills in C++ and Python.Familiarity...C++Full timeWork experience placement- ...Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design... ..., build new abstractions for LLM serving engines, and contribute... ...collaborate across teams, work on CUDA C/C++, Triton, and cutting-edge MLIR-based tooling, and...C++
- ...involves working with advanced technologies in AI and autonomous driving. Applicants should... ...degree and strong programming skills in C++ and Python, along with an interest in deep... ...environment and the chance to work on cutting-edge technologies. #J-18808-Ljbffr XpengmotorsC++Internship
$184.7k - $324.8k
...Foundation Models Inference - Cloud OS & Inference... ...Learning and AI We are the Foundation... ...partners to bring cutting-edge model architectures... ...on experience with LLM inference stacks.... ...frameworks such as TensorRT-LLM, vLLM, SGLang,... ...kernels using CUDA C++ or OpenAI Triton....C++WorldwideRelocation- ...Engineer to advance Deep Learning Inference within TensorRT. You will build scalable inferencing... ...the-art in DL inference, leveraging C++, Python, and cutting-edge NVIDIA GPUs. This role offers... ...opportunity to shape next-generation AI acceleration. #J-18808-Ljbffr NVIDIA...C++
$206.4k - $379.1k
...impressive content. The AI Foundations team... ...looking for a Principal Architect to build and implement... ...spanning model orchestration, inference systems, data pipelines... ..., data pipelines, LLM orchestration layers, in... ...engineeringProficiency in Python, Java, C++, or Go, with an...C++Full timeTemporary workLocal areaWorldwideFlexible hours- ...computing experiences—from AI and data centers, to PCs... ...a Principal GenAI Inference Optimization Engineer to... ...and cost efficiency for LLM and multimodal model serving... ...vLLM, SGLang, Triton, TensorRT-LLM, or similar.- Experience... ...one systems language (C++/CUDA/HIP).- Experience...C++
$207k - $300k
Analyze and optimize AI inference workloads across the application... ...in Python and C++, including navigating,... ...Experience with real world LLM inference serving... ...frameworks (e.g., vLLM, TensorRT-LLM, SGLang, Dynamo).Experience... ...with the cutting edge AI agents developed by...C++- ...unleashing the potential of generative AI to power the transformation of... ...System Software Engineer, AI Inference ExecutionWhat you will do:The... ...fundamentalsProficient in C/C++/Python development in Linux... ...serving frameworks (such as TensorRT-LLM, vLLM, SGLang, etc.)Experience...C++3 days per week
$184k - $287.5k
...platform upon which every new AI-powered application is... ...Software Engineer - AI Inference Performance to advance innovative LLM and VLM inference. You will... ...high-quality upgrades to TensorRT-LLM, vLLM, SGLang, or... ...skills in Python, Rust and/or C++, plus hands-on experience...C++Full time$152k - $241.5k
...and eager to work on cutting-edge AI technology for safety-... ...applications? Join NVIDIA's TensorRT team as a Senior Software Engineer... ...high-performance AI inference solutions for automotive safety... ...automotive applications using modern C++Orchestrate the integration of new...C++Full time$219k - $351k
...Title: Principal engineer, AI Serving Framework Architect (Software)The Architecture... ...methodologies for maximizing AI inference performance in multi-rack... ...a Large Language Model (LLM) Inference Software Stack on... ...skillsets: PyTorch, Python, and C++ A collaborative mindset,...C++Work at officeFlexible hours$195.2k - $361.2k
...journey is to transform AI into something safer, more... ...user's machine (AI PC, edge, on-prem, and beyond),... ...actually own. You optimize inference engines (llama.cpp, vLLM... ...backgroundStrong in C++ and/or Python; comfortable... ...level codeExperience with LLM inference. (attention,...C++Full timeInternshipLocal areaImmediate startShift work- ...can thrive.Job DescriptionThe AI Inference Engineer plays a critical... ...centers to resource-constrained edge devices—with a strong emphasis... ...including NVIDIA GPUs (CUDA/TensorRT), Apple Silicon (CoreML), and... ...programming languages such as Python, C++, Rust, or Golang specifically...C++Full timeLocal areaImmediate start
$184k - $287.5k
We are now looking for a Senior Deep Learning Architect, LLM Inference!NVIDIA is at the forefront of the generative AI revolution. The Inference Benchmarking (IB) team specifically focuses on inference server performance optimization for Large Language Models (LLMs). If...Full time$124k - $195.5k
...NVIDIA TensorRT Team Software EngineerAre you passionate about driving... ...eager to work on cutting-edge AI technology? Join NVIDIA's... ...contributing to high-performance AI inference solutions for specialized... ...software using modern C++Collaborate with teams across the hardware...C++Internship
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Edge AI LLM Inference Architect (C++, TensorRT). Be the first to apply!

