Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Deep Learning Architect, LLM Inference

$184k - $287.5k

NVIDIA

We are now looking for a Senior Deep Learning Architect, LLM Inference!NVIDIA is at the forefront of the generative AI revolution. The Inference Benchmarking (IB) team specifically focuses on inference server performance optimization for Large Language Models (LLMs). If you're passionate about pushing the boundaries of GPU hardware and software performance and understand terms like disaggregated serving, data parallel attention, MoE, Qwen3.5, DeepSeek, GPT-OSS, then this is a great role for you!What you'll be doing:You will do workload characterization of the latest LLMs and inference servers like vLLM, SGLang and TRT-LLM to ensure NVIDIA maintains its leadership position.Join forces with the performance marketing team to build engaging content, including blog posts and updates to InferenceX to highlight NVIDIA's outstanding inference achievements.Collaborate with engineers from AI startup companies to establish standard benchmarking methodologies.Develop a constantly evolving inference performance data results website.Invent E2E profiling and analysis tools that you will use to keep up with the rapid pace of Generative AI.Contribute to deep learning software projects, such as PyTorch, TRT-LLM, vLLM, and SGLang to drive advancements in the field.Verify that new GPU product launches produce industry leading performance.Collaborate across the company to guide the direction of inference serving, working with software, research, and product teams to ensure best-in-class performance.Use the latest coding agents and inference technology to improve team efficiency.What we need to see:Master's or PhD degree in Computer Science, Computer Engineering, related fields, or equivalent experience. 6+ years of relevant software development experience.Detailed knowledge of deep learning inference serving, PyTorch programming, profiling, and compiler optimizations.Experience developing client server LLM applications with OpenAI API or MCP and identifying performance bottlenecks.Solid understanding of CPU and GPU microarchitecture and performance characteristics.Experience with complex software projects like frameworks, compilers, or operating systems.Demonstrated proficiency with the latest AI coding agents like Claude Code, Codex, and CursorExcellent written and verbal communication skills and the ability to work independently and collaboratively in a fast-paced environment.Ways to stand out from the crowd:Demonstrate a drive to continuously improve software and hardware performance.Showcase examples of novel use cases for agentic AI tools in the workplace.Experience with databases and visualization tools will set you apart.NVIDIA is widely considered to be one of the technology world's most desirable employers. We have a team of highly skilled and motivated individuals who excel in their work. If you have a proactive and independent approach, we want to hear from you!Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until August 15, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior Deep Learning Architect, LLM Inference in Santa Clara, CA vacancy
  • $208k - $327.75k

     ....We are looking for a Senior AI Architect to help define the next...  ...decisions through deep workload characterization...  ...Engineering, Machine Learning, Robotics, or related...  ...systems, scaling laws, and inference optimization...  ..., Triton, or TensorRT-LLM.Experience influencing... 
    Senior
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...architecture group at NVIDIA has openings for a Deep Learning Communication Architect. We scale the DNN models and training/inference frameworks to systems with hundreds of...  ...Experience in evaluating, analyzing, and optimizing LLM training and inference performance of state-... 
    Senior
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...and we’re seeking a visionary Product Architect with strong expertise in systems architecture...  ...agentic & RAG-based workflows, inference at scale, large scale training & fine-tuning...  ...certifications or publications in AI, deep learning, or related fieldsExperience leading... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

    We are now looking for a Senior GPU & Deep Learning Architect!The NVIDIA GPU Architecture group is looking for world class architects and software...  ...especially for deep learning workloads, both training and inference, and maintain our leadership by developing new parallel... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

    We are now looking for a Senior Deep Learning Computer Architect! NVIDIA is seeking architects like you to help design hardware accelerator and processor...  ...;Performance analysis and optimization;Experience with LLM workloads, including performance tuning considerations such... 
    Senior
    Full time
    Night shift

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

    We are now seeking a Senior Deep Learning Performance Architect!NVIDIA is looking for outstanding Performance Architects with a background in performance...  ...crowd:Background with deep neural network training, inference and optimization in leading frameworks (e.g. Pytorch,... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $210.87k - $329.86k

     ...a high-impact, visionary Senior Principal AI Engineer to architect, design, and oversee the next...  ...-Tuned Enterprise LLM," transforming massive, highly...  ...Thermal Engineers to encode deep domain expertise, design rules...  ...NLP), or advanced Machine Learning architectures.Hardware... 
    Senior
    Temporary work
    Work at office
    Local area
    Worldwide
    Shift work

    Celestica

    San Jose, CA
    3 days ago
  • NVIDIA is seeking a senior leader to shape the global strategy for scaled-out AI inference. You will architect high-throughput, low-latency distributed pipelines and model serving strategies for massive scale and reliability on NVIDIA hardware. You will drive the technical... 
    Senior

    NVIDIA

    Santa Clara, CA
    1 day ago
  • Advanced Micro Devices is seeking a Senior GPU Inference Performance Engineer to own end-to-end profiling of GPU-accelerated AI inference workloads. You will analyze workloads across AMD Instinct and NVIDIA GPUs, profile AI serving frameworks, and explain performance gaps... 
    Senior

    Advanced Micro Devices

    Santa Clara, CA
    26 minutes ago
  • $184k - $356.5k

    NVIDIA is seeking a Senior High-Performance LLM Training Engineer to enhance the efficiency of LLM training workloads. Focused on optimizing NVIDIA...  ...PhD or equivalent degree with substantial experience in deep learning, GPU architecture, and performance optimization. The base... 
    Senior

    NVIDIA

    Santa Clara, CA
    24 minutes ago
  •  ...innovative infrastructure company is seeking a Member of Technical Staff — CI Engineer to improve CI reliability for their open-source LLM inference engine. The role requires 3+ years' experience in CI/CD, knowledge of Linux and GPU computing, as well as strong skills in Bash... 
    Senior

    RadixArk

    Palo Alto, CA
    3 days ago
  • $168k - $258.75k

    Inference is the fastest growing and most competitive area in Generative...  ...deployment techniques. As a Senior Product Manager for AI...  ...group. We focus on enabling deep learning across all GPU use cases and...  ...SGLang, FlashInfer, TensorRT-LLM, Triton, Dynamo, TorchAO, etc... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • Lead the end-to-end performance analysis and optimization of LLM and VLM inference on NVIDIA GPU-accelerated systems. Develop performance-critical kernels and contribute improvements to open-source inference engines to reduce latency and increase throughput. Requires over... 
    Senior

    NVIDIA AI

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...such as large language models (LLM) and diffusion models for maximal inference efficiency using techniques ranging...  ....We are now looking for a Senior Deep Learning Software Engineer to develop and...  ...market.Play a pivotal role in architecting and designing a modular and scalable... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $272k - $431.25k

     ...with application developers to architect and implement specialized...  ...computing (HPC) or distributed deep learning.Parallelism Expertise: Deep understanding...  ...InfiniBand verbs is required.Inference & Serving: Advanced knowledge...  ..., specifically TensorRT-LLM, vLLM, SGLang, and NVIDIA... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $208k - $327.75k

     ...running on NVIDIA hardware. Every inference deployment — from a single-GPU...  ...to run inference by turning deep optimization techniques into...  ...team driving the company’s Deep Learning and Generative AI strategy. We...  ...land across TensorRT-LLM, vLLM, SGLang, and NVIDIA Dynamo... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $174.72k - $295.68k

     ...cutting-edge R&D in AI, machine learning, and smart connectivity.Our...  ...to build strong foundation for LLM deployment and quality sign-off...  ...tuning, PTQ, QAT, on-vehicle inference and related fields.Key ResponsibilitiesDevelop...  ...quantizing or deploying deep learning models in production.... 
    Senior
    Full time

    XPENG Motors

    Santa Clara, CA
    1 day ago
  • d-Matrix in Santa Clara, CA is seeking a Sr. Staff ML Researcher to advance LLM algorithmic optimization on our DNN accelerators. You will design and implement efficient inference algorithms, collaborating with mathematicians, ML researchers and engineers on high-impact... 
    Senior
    3 days per week

    Entrada Ventures

    Santa Clara, CA
    1 day ago
  •  ...seeking a highly capable software engineer to advance an advanced inference framework using modern C++. The role focuses on extending...  ...a BS/MS/PhD (or equivalent) and have at least four years of software development experience with a deep #J-18808-Ljbffr NVIDIA AI

    NVIDIA AI

    Santa Clara, CA
    22 minutes ago
  • $148k - $235.75k

    We are looking for a Senior Technical Product Marketing Manager....  ...business and pivotal in our inference marketing. You will be focused...  ...preferred.6+ years of experience in LLM, AI/ML development in an...  ..., distributed inference, deep learning frameworks (PyTorch, TensorFlow... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...We are seeking an expert Solutions Architect to assist customers in building AI/ML...  ...aspects related to tasks like large scale LLM training and inference.Conducting regular technical...  ...the crowd:Hands-on experience with Deep Learning frameworks (PyTorch, JAX, etc.), compilers... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $224k - $356.5k

    We are now looking for a Senior System Software Engineer to work...  ...engineers for its GPU-accelerated deep learning software team. Academic and...  ...team building Generative AI inference platform to make design and...  ...inference engines (vLLM, SGLang, TRT-LLM) and expand these... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...what’s possible with LLM inference on heterogeneous hardware...  ...patterns to deep optimization of inference...  ...closely with hardware architects to provide firmware and...  ...real hardware.• Small, senior team with high autonomy...  ...embrace challenges and learn together every day.d-Matrix... 
    Senior

    d-Matrix

    Santa Clara, CA
    4 days ago
  • $224k - $356.5k

     ...technology for architectures used for artificial intelligence (AI) / deep learning (DL), high-performance computing (HPC), cloud service...  ...socket CPU and CPU/GPU systems.Work with CPU and interconnect architects to improve future CPU and system designs based on your... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

    We are now looking for a Senior AI Training Performance ArchitectNVIDIA is seeking a senior engineer who is obsessed with performance...  ...-quality software across multiple layers of NVIDIA's deep learning platform stack, from drivers to DL frameworks.Build and support... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...engineers to join us and build AI inference systems that serve large-...  ...extreme efficiency. You’ll architect and implement high-performance...  ...programming, distributed systems, deep learning theories.Knowledgeable and...  ...building and optimizing LLM inference engines (e.g., vLLM... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $168k - $258.75k

     ...graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing. NVIDIA is a “...  ...creativity and intelligence.NVIDIA is looking for a Product Architect to help define & design SoC- and GPU-based products and drive... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

     ...are now looking for a Senior Performance Architect for Nemotron! At NVIDIA...  ...of AI systems through deep model-system-hardware co...  ...Decoding, Agentic Pipelines, Inference-time compute scaling,...  ....Experience with deep learning frameworks like PyTorch, TRT-LLM, VLLM, SGLangA Growth... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $224k - $356.5k

     ...world.As an AI Storage Platform Architect at NVIDIA, this position will...  ...for disaggregated inference (aligned with NVIDIA Dynamo),...  ...field (or equivalent experience).Deep expertise in AI infrastructure...  ...disaggregated inference architectures, LLM training pipelines, and autonomous... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

     ...science. Today, NVIDIA’s GPU simulates human intelligence, running deep learning algorithms and acting as the brain of computers, robots and...  ...our team! NVIDIA Architecture Modeling group is looking for Architects, Functional Modeling Engineers, and Simulation experts to... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Deep Learning Architect, LLM Inference. Be the first to apply!