Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Edge AI LLM Inference Architect (C++, TensorRT)

NVIDIA AI

NVIDIA AI in Santa Clara is seeking a highly capable software engineer to advance an advanced inference framework using modern C++. The role focuses on extending TensorRT with autoregressive model serving capabilities and requires collaboration across CUDA, kernel libraries, compilers, and robotics teams to deliver high-performance, production-ready solutions. The candidate should hold a BS/MS/PhD (or equivalent) and have at least four years of software development experience with a deep #J-18808-Ljbffr NVIDIA AI

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Edge AI LLM Inference Architect (C++, TensorRT) in Santa Clara, CA vacancy
  • Develop and evolve a state-of-the-art inference framework in modern C++ that extends TensorRT with autoregressive model serving...  .... Key Skills C++, TensorRT, LLM, VLM, GEMM, CUDA, Attention, MoE,...  ...Infrastructure, Robotics, Embedded AI Benefits Equity, Health Insurance... 
    C++

    NVIDIA AI

    Santa Clara, CA
    4 days ago
  • Job Description: Edge AI Architect - CUDA / C++ / Computer Vision Experience Level: 10+ Years Department: Edge AI & Embedded Systems About the Role...  ...optimization tools. AI/ML Acumen (Crucial) Required: Strong LLM, RAG, agentic architecture understanding. Understanding... 
    C++
    Contract work

    codvo-team

    Santa Clara, CA
    2 days ago
  • $174.05k - $278.45k

    Distinguished Technologist, Edge AI Architect Description - Job Summary The...  ...security and OS trust foundation, inference serving, agentic runtimes and...  ...Certification (Python, C++, Rust, Java, or similar). Cloud...  ...a plus. Knowledge & Skills LLM, vLM, and multi-modal model... 
    C++
    Full time
    Contract work
    Temporary work
    Work experience placement
    Local area
    Flexible hours
    Shift work

    HP Inc.

    Palo Alto, CA
    1 day ago
  • $272k - $431.25k

     ...group is solving some of AI’s hardest...  ...interconnects.This Principal Architect role leads the research...  ...such as vLLM, SGLang, and TensorRT-LLM.Publishing findings, representing...  ...training and inference patterns.Proficiency in...  ...programming languages such as C, C++, Rust and Python.Ways... 
    C++
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $272k - $431.25k

     ...throughput, low-latency inference framework for serving generative AI and reasoning...  ...deployment of cutting-edge LLM workloads.We are...  ...scale LLM inference.Architect and implement deep...  ...such as vLLM, SGLang, TensorRT-LLM), with a focus...  ...in C/C++ and Python, with a... 
    C++
    Full time
    Local area
    Remote work

    Nvidia

    Santa Clara, CA
    18 hours ago
  •  ...potential of generative AI to power the...  ...sits at the leading edge of what’s possible with LLM inference on heterogeneous hardware...  ...(vLLM, SGLang, TensorRT-LLM, etc.) and can extend...  ...with hardware architects to provide firmware...  ...proficiency in Python and C/C++.• Hands-on... 
    C++

    d-Matrix

    Santa Clara, CA
    3 days ago
  • codvo-team is seeking an Edge AI Architect to lead design and deployment of AI models on edge devices for medical applications. The role focuses...  ...candidate has 10+ years in AI engineering, hands-on CUDA/C++ experience, and familiarity with OpenCV, TensorFlow, and PyTorch... 
    C++
    Remote work

    codvo-team

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...unlimited potential of AI to define the next...  ...workflows at the edge powered by NVIDIAs...  ...performance.Improve LLM & GenAI user...  ...Strong proficiency in C/C++, Python, software design...  ...GPU-accelerated AI inference driven by NVIDIA...  ...SDKs, specifically TensorRT-RTX, cuDNN, NVIDIA... 
    C++
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    2 days ago
  • $151.8k - $332.2k

     ...We are looking for an AI Inference Engineer with a solid background...  ...on the most cutting edge speech modeling and...  ...recognition, speech-llm or AI model inference.Display...  ..., shell scripts, C/C++; familiarity with ML frameworks...  ...AI models using CUDA, TensorRT, and mixed-precision... 
    C++
    Full time
    Work at office
    Remote work

    Zoom

    San Jose, CA
    18 hours ago
  • $182.5k - $260.5k

     ...networking for the cloud and AI era. We secure and...  ...Scientist, you own the inference and optimization layer...  ...of agentic AI.Cutting-edge, unusual stack. The hard...  ...runtimes (vLLM/SGLang, TensorRT-LLM, ONNX Runtime, llama.cpp...  ...comfort reaching into C++ for low-level interop is... 
    C++

    Netskope

    Santa Clara, CA
    1 day ago
  • $197.5k - $272k

     ...the transformation to AI-enabled software-defined...  ...-grade AI on the Edge. We are looking for a...  ...(e.g., Transformers, LLM, CNN, LSTM, Trees) to...  ...working knowledge of modern C++ (C++14/17 for inference).Deep proficiency with...  ...Experience with NVIDIA TensorRT, Qualcomm SNPE.... 
    C++
    Work at office
    Worldwide
    Flexible hours
    Shift work
    3 days per week

    Sonatus

    Sunnyvale, CA
    18 hours ago
  • $124k - $195.5k

     ...Software Engineer, TensorRT PerformanceNVIDIA is...  ...of NVIDIA's inference ecosystem! NVIDIA is...  ...areas like Generative AI, Recommenders and Vision...  ...datacenter GPUs to edge SoCs. Implement...  ...experience.Strong C++, Python programming...  ...TensorRT, TensorRT-LLM, vLLM, SGLang, FlashInfer... 
    C++

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

     ...Engineer, Deep Learning Inference -...  ...Deep Learning Inference TensorRT software team.**What you...  ...feature update TensorRT* Use C++ and Python to build graph...  ...learning experts, GPU architects and DevOps engineers across...  ...existing vacancy.NVIDIA uses AI tools in its recruiting... 
    C++

    NVIDIA Corporation

    Santa Clara, CA
    2 days ago
  • $197.3k - $225.1k

    Lead AI Engineer (FM Hosting, LLM Inference) Overview At Capital One, we are creating responsible and reliable AI systems, changing banking for good....  ...Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, or Golang ~ Experience developing and applying... 
    C++
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    3 days ago
  • $184k - $287.5k

     ...unlimited potential of AI to define the...  ...Communication Architect. We scale the...  ...models and training/inference frameworks to...  ...and optimizing LLM training and inference...  ...on cutting-edge hardware.Deep understanding...  ...as PyTorch, TensorRT-LLM, vLLM,...  ...skills in C++ and Python.Familiarity... 
    C++
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    18 hours ago
  •  ...Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design...  ..., build new abstractions for LLM serving engines, and contribute...  ...collaborate across teams, work on CUDA C/C++, Triton, and cutting-edge MLIR-based tooling, and... 
    C++

    NVIDIA

    Santa Clara, CA
    18 hours ago
  •  ...involves working with advanced technologies in AI and autonomous driving. Applicants should...  ...degree and strong programming skills in C++ and Python, along with an interest in deep...  ...environment and the chance to work on cutting-edge technologies. #J-18808-Ljbffr Xpengmotors
    C++
    Internship

    Xpengmotors

    Santa Clara, CA
    4 days ago
  • $184.7k - $324.8k

     ...Foundation Models Inference - Cloud OS & Inference...  ...Learning and AI We are the Foundation...  ...partners to bring cutting-edge model architectures...  ...on experience with LLM inference stacks....  ...frameworks such as TensorRT-LLM, vLLM, SGLang,...  ...kernels using CUDA C++ or OpenAI Triton.... 
    C++
    Worldwide
    Relocation

    Apple Inc.

    Santa Clara, CA
    1 day ago
  •  ...Engineer to advance Deep Learning Inference within TensorRT. You will build scalable inferencing...  ...the-art in DL inference, leveraging C++, Python, and cutting-edge NVIDIA GPUs. This role offers...  ...opportunity to shape next-generation AI acceleration. #J-18808-Ljbffr NVIDIA... 
    C++

    NVIDIA Corporation

    Santa Clara, CA
    2 days ago
  • $206.4k - $379.1k

     ...impressive content. The AI Foundations team...  ...looking for a Principal Architect to build and implement...  ...spanning model orchestration, inference systems, data pipelines...  ..., data pipelines, LLM orchestration layers, in...  ...engineeringProficiency in Python, Java, C++, or Go, with an... 
    C++
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours

    Adobe Systems

    San Jose, CA
    18 hours ago
  •  ...computing experiences—from AI and data centers, to PCs...  ...a Principal GenAI Inference Optimization Engineer to...  ...and cost efficiency for LLM and multimodal model serving...  ...vLLM, SGLang, Triton, TensorRT-LLM, or similar.- Experience...  ...one systems language (C++/CUDA/HIP).- Experience... 
    C++

    AMD

    San Jose, CA
    4 days ago
  • $207k - $300k

    Analyze and optimize AI inference workloads across the application...  ...in Python and C++, including navigating,...  ...Experience with real world LLM inference serving...  ...frameworks (e.g., vLLM, TensorRT-LLM, SGLang, Dynamo).Experience...  ...with the cutting edge AI agents developed by... 
    C++

    Google

    Mountain View, CA
    2 days ago
  •  ...unleashing the potential of generative AI to power the transformation of...  ...System Software Engineer, AI Inference ExecutionWhat you will do:The...  ...fundamentalsProficient in C/C++/Python development in Linux...  ...serving frameworks (such as TensorRT-LLM, vLLM, SGLang, etc.)Experience... 
    C++
    3 days per week

    d-Matrix

    Santa Clara, CA
    18 hours ago
  • $184k - $287.5k

     ...platform upon which every new AI-powered application is...  ...Software Engineer - AI Inference Performance to advance innovative LLM and VLM inference. You will...  ...high-quality upgrades to TensorRT-LLM, vLLM, SGLang, or...  ...skills in Python, Rust and/or C++, plus hands-on experience... 
    C++
    Full time

    Nvidia

    Santa Clara, CA
    14 hours ago
  • $152k - $241.5k

     ...and eager to work on cutting-edge AI technology for safety-...  ...applications? Join NVIDIA's TensorRT team as a Senior Software Engineer...  ...high-performance AI inference solutions for automotive safety...  ...automotive applications using modern C++Orchestrate the integration of new... 
    C++
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $219k - $351k

     ...Title: Principal engineer, AI Serving Framework Architect (Software)The Architecture...  ...methodologies for maximizing AI inference performance in multi-rack...  ...a Large Language Model (LLM) Inference Software Stack on...  ...skillsets: PyTorch, Python, and C++ A collaborative mindset,... 
    C++
    Work at office
    Flexible hours

    Samsung Semiconductor

    San Jose, CA
    18 hours ago
  • $195.2k - $361.2k

     ...journey is to transform AI into something safer, more...  ...user's machine (AI PC, edge, on-prem, and beyond),...  ...actually own. You optimize inference engines (llama.cpp, vLLM...  ...backgroundStrong in C++ and/or Python; comfortable...  ...level codeExperience with LLM inference. (attention,... 
    C++
    Full time
    Internship
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    18 hours ago
  •  ...can thrive.Job DescriptionThe AI Inference Engineer plays a critical...  ...centers to resource-constrained edge devices—with a strong emphasis...  ...including NVIDIA GPUs (CUDA/TensorRT), Apple Silicon (CoreML), and...  ...programming languages such as Python, C++, Rust, or Golang specifically... 
    C++
    Full time
    Local area
    Immediate start

    F5 Networks

    San Jose, CA
    14 hours ago
  • $184k - $287.5k

    We are now looking for a Senior Deep Learning Architect, LLM Inference!NVIDIA is at the forefront of the generative AI revolution. The Inference Benchmarking (IB) team specifically focuses on inference server performance optimization for Large Language Models (LLMs). If... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $124k - $195.5k

     ...NVIDIA TensorRT Team Software EngineerAre you passionate about driving...  ...eager to work on cutting-edge AI technology? Join NVIDIA's...  ...contributing to high-performance AI inference solutions for specialized...  ...software using modern C++Collaborate with teams across the hardware... 
    C++
    Internship

    NVIDIA

    Santa Clara, CA
    19 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Edge AI LLM Inference Architect (C++, TensorRT). Be the first to apply!