Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Deep Learning Architect, LLM Inference

$184k - $287.5k

Nvidia

We are now looking for a Senior Deep Learning Architect, LLM Inference!NVIDIA is at the forefront of the generative AI revolution. The Inference Benchmarking (IB) team specifically focuses on inference server performance optimization for Large Language Models (LLMs). If you're passionate about pushing the boundaries of GPU hardware and software performance and understand terms like disaggregated serving, data parallel attention, MoE, Qwen3.5, DeepSeek, GPT-OSS, then this is a great role for you!What you'll be doing:You will do workload characterization of the latest LLMs and inference servers like vLLM, SGLang and TRT-LLM to ensure NVIDIA maintains its leadership position.Join forces with the performance marketing team to build engaging content, including blog posts and updates to InferenceX to highlight NVIDIA's outstanding inference achievements.Collaborate with engineers from AI startup companies to establish standard benchmarking methodologies.Develop a constantly evolving inference performance data results website.Invent E2E profiling and analysis tools that you will use to keep up with the rapid pace of Generative AI.Contribute to deep learning software projects, such as PyTorch, TRT-LLM, vLLM, and SGLang to drive advancements in the field.Verify that new GPU product launches produce industry leading performance.Collaborate across the company to guide the direction of inference serving, working with software, research, and product teams to ensure best-in-class performance.Use the latest coding agents and inference technology to improve team efficiency.What we need to see:Master's or PhD degree in Computer Science, Computer Engineering, related fields, or equivalent experience. 6+ years of relevant software development experience.Detailed knowledge of deep learning inference serving, PyTorch programming, profiling, and compiler optimizations.Experience developing client server LLM applications with OpenAI API or MCP and identifying performance bottlenecks.Solid understanding of CPU and GPU microarchitecture and performance characteristics.Experience with complex software projects like frameworks, compilers, or operating systems.Demonstrated proficiency with the latest AI coding agents like Claude Code, Codex, and CursorExcellent written and verbal communication skills and the ability to work independently and collaboratively in a fast-paced environment.Ways to stand out from the crowd:Demonstrate a drive to continuously improve software and hardware performance.Showcase examples of novel use cases for agentic AI tools in the workplace.Experience with databases and visualization tools will set you apart.NVIDIA is widely considered to be one of the technology world's most desirable employers. We have a team of highly skilled and motivated individuals who excel in their work. If you have a proactive and independent approach, we want to hear from you!Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until August 15, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior Deep Learning Architect, LLM Inference in Santa Clara, CA vacancy
  • $208k - $327.75k

     ....We are looking for a Senior AI Architect to help define the next...  ...decisions through deep workload characterization...  ...Engineering, Machine Learning, Robotics, or related...  ...systems, scaling laws, and inference optimization...  ..., Triton, or TensorRT-LLM.Experience influencing... 
    Senior
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...architecture group at NVIDIA has openings for a Deep Learning Communication Architect. We scale the DNN models and training/inference frameworks to systems with hundreds of...  ...Experience in evaluating, analyzing, and optimizing LLM training and inference performance of state-... 
    Senior
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

    We are now looking for a Senior Deep Learning Computer Architect! NVIDIA is seeking architects like you to help design hardware accelerator and processor...  ...;Performance analysis and optimization;Experience with LLM workloads, including performance tuning considerations such... 
    Senior
    Full time
    Night shift

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

    We are now looking for a Senior GPU & Deep Learning Architect!The NVIDIA GPU Architecture group is looking for world class architects and software...  ...especially for deep learning workloads, both training and inference, and maintain our leadership by developing new parallel... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

    We are now looking for a Senior Deep Learning Performance Architect! NVIDIA is seeking outstanding Performance Analysis Architects to help analyze and...  ...Learning ASIC architecture evaluation for training and/or inference.Strong programming skills in Python and C++. Ways to... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $168k - $258.75k

    Inference is the fastest growing and most competitive area in Generative...  ...deployment techniques. As a Senior Product Manager for AI...  ...group. We focus on enabling deep learning across all GPU use cases and...  ...SGLang, FlashInfer, TensorRT-LLM, Triton, Dynamo, TorchAO, etc... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

     ...advancements in AI and machine learning to solve some of the...  ...the industry-leading deep learning inference software for NVIDIA AI accelerators. As a Senior Software Engineer in...  ...and TensorRT-LLM to supercharge inference...  ...experts and GPU architects throughout the company... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $272k - $431.25k

     ...with application developers to architect and implement specialized...  ...computing (HPC) or distributed deep learning.Parallelism Expertise: Deep understanding...  ...InfiniBand verbs is required.Inference & Serving: Advanced knowledge...  ..., specifically TensorRT-LLM, vLLM, SGLang, and NVIDIA... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $148k - $235.75k

    We are looking for a Senior Technical Product Marketing Manager....  ...business and pivotal in our inference marketing. You will be focused...  ...preferred.6+ years of experience in LLM, AI/ML development in an...  ..., distributed inference, deep learning frameworks (PyTorch, TensorFlow... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $224k - $356.5k

    We are now looking for a Senior System Software Engineer to work...  ...engineers for its GPU-accelerated deep learning software team. Academic and...  ...team building Generative AI inference platform to make design and...  ...inference engines (vLLM, SGLang, TRT-LLM) and expand these... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...what’s possible with LLM inference on heterogeneous hardware...  ...patterns to deep optimization of inference...  ...closely with hardware architects to provide firmware and...  ...real hardware.• Small, senior team with high autonomy...  ...embrace challenges and learn together every day.d-Matrix... 
    Senior

    d-Matrix

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...We are seeking an expert Solutions Architect to assist customers in building AI/ML...  ...aspects related to tasks like large scale LLM training and inference.Conducting regular technical...  ...the crowd:Hands-on experience with Deep Learning frameworks (PyTorch, JAX, etc.), compilers... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $224k - $356.5k

     ...technology for architectures used for artificial intelligence (AI) / deep learning (DL), high-performance computing (HPC), cloud service...  ...socket CPU and CPU/GPU systems.Work with CPU and interconnect architects to improve future CPU and system designs based on your... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

    We are now looking for a Senior AI Training Performance ArchitectNVIDIA is seeking a senior engineer who is obsessed with performance...  ...-quality software across multiple layers of NVIDIA's deep learning platform stack, from drivers to DL frameworks.Build and support... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $168k - $258.75k

     ...graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing. NVIDIA is a “...  ...creativity and intelligence.NVIDIA is looking for a Product Architect to help define & design SoC- and GPU-based products and drive... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

     ...are now looking for a Senior Performance Architect for Nemotron! At NVIDIA...  ...of AI systems through deep model-system-hardware co...  ...Decoding, Agentic Pipelines, Inference-time compute scaling,...  ....Experience with deep learning frameworks like PyTorch, TRT-LLM, VLLM, SGLangA Growth... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...engineers to join us and build AI inference systems that serve large-...  ...extreme efficiency. You’ll architect and implement high-performance...  ...programming, distributed systems, deep learning theories.Knowledgeable and...  ...building and optimizing LLM inference engines (e.g., vLLM... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $224k - $356.5k

     ...world.As an AI Storage Platform Architect at NVIDIA, this position will...  ...for disaggregated inference (aligned with NVIDIA Dynamo),...  ...field (or equivalent experience).Deep expertise in AI infrastructure...  ...disaggregated inference architectures, LLM training pipelines, and autonomous... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

     ...science. Today, NVIDIA’s GPU simulates human intelligence, running deep learning algorithms and acting as the brain of computers, robots and...  ...our team! NVIDIA Architecture Modeling group is looking for Architects, Functional Modeling Engineers, and Simulation experts to... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing. NVIDIA is...  ...team today!We are looking for an outstanding hands-on architect/engineer for a Senior HPC architect role to support deployment and bringup of... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $169.78k - $338.69k

     ...experienced and mission-driven Senior/Staff AI Infrastructure Engineer, Inference & Optimization to lead...  ...gap between frontier deep learning algorithms and real-...  ...anomalies. Architect and scale service-oriented...  ...Runtime) and specialized LLM inference/serving... 
    Senior
    Full time

    DiDi Labs

    San Jose, CA
    2 days ago
  • $193.3k - $261.5k

     ...kit used to accelerate deep learning and GenAI workloads on...  ...unparalleled ML inference and training performance...  ...acceleration technologyYou will architect and implement business...  ...of a wide variety of LLM model families,...  ...sharing and mentorship. Our senior members enjoy one-on-... 
    Senior
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  •  ...Senior AI Architect Nexxa.ai is building artificial super intelligence for heavy industries —...  ...environments. Our mission is to translate deep technical breakthroughs into...  ...experience in software engineering, machine learning, data science, or closely related technical... 
    Senior

    Nexxa.ai

    Sunnyvale, CA
    5 days ago
  •  ...Principal System Software Engineer, AI Inference ExecutionWhat you will do:The role...  ...architecture, and machine learning fundamentalsProficient in C/C++/Python...  ...frameworks (such as TensorRT-LLM, vLLM, SGLang, etc.)Experience with deep learning frameworks (such as PyTorch... 
    3 days per week

    d-Matrix

    Santa Clara, CA
    1 day ago
  •  ...NVIDIA Corporation is seeking a Senior Software Architect to enhance communication libraries crucial for scaling Deep Learning and HPC applications. You will investigate performance bottlenecks and design innovative communication technologies. The ideal candidate holds... 
    Senior
    Remote job

    Jobleads-US

    Santa Clara, CA
    2 days ago
  • $229.9k - $262.4k

    Senior Lead AI Engineer (LLM Gateway, FM Hosting) Overview: At Capital One, we are...  ...industry leader in using machine learning to create real-time,...  ...class talent — along with our deep experience in machine...  ...training, large language model inference, similarity search, guardrails... 
    Senior
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    3 days ago
  •  ...We are seeking a Robotics AI Architect to define and scale next-generation...  ..., you will synthesize learnings from real-world deployments and...  ...refinementsLead deep technical engagements, including...  ...stakeholdersDeep understanding of:AI inference runtimes and deployment tradeoffsSystem... 
    Senior

    AMD

    San Jose, CA
    2 days ago
  • $182.5k - $260.5k

     ....Visit Careers at Netskope to learn more. Follow us on LinkedIn and...  ....Positions are available at Senior Staff and above. Candidates are...  ...Learning Scientist, you own the inference and optimization layer that...  ...runtimes (vLLM/SGLang, TensorRT-LLM, ONNX Runtime, llama.cpp, or... 
    Senior

    Netskope

    Santa Clara, CA
    2 days ago
  •  ...you will play a pivotal role in optimizing and developing deep learning frameworks for AMD GPUs. Your work will be instrumental...  ...deep learning models, and enabling RL training and SOTA LLM and Multimodal inference at scale across multi-GPU and multi-node systems. You... 
    Senior

    AMD

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

     ...passionate about driving innovation in deep learning and eager to work on cutting-edge AI technology...  ...? Join NVIDIA's TensorRT team as a Senior Software Engineer, and be at the...  ...technology, enabling high-performance AI inference solutions for automotive safety and other... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Deep Learning Architect, LLM Inference. Be the first to apply!