Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Engineering Manager, Deep Learning Inference

$224k - $356.5k

NVIDIA

NVIDIA is seeking an exceptional Manager, Deep Learning Inference Software, to lead a world-class engineering team advancing the state of AI model deployment. You will shape the software powering today’s most sophisticated AI systems — from large language models to multimodal generative AI — all accelerated on NVIDIA GPUs. The Deep Learning Inference team develops and optimizes open-source frameworks that make AI deployment scalable, efficient, and accessible — including SGLang, vLLM, and FlashInfer. Our work enables developers worldwide to harness NVIDIA accelerators for real-time inference at every scale, from datacenter clusters to edge devices.What you'll be doing:Lead, mentor, and scale a high-performing engineering team focused on deep learning inference and GPU-accelerated software.Guide the strategy, roadmap, and execution of NVIDIA's OSS inference frameworks engineering.Partner with internal compiler, libraries, and research teams to deliver end-to-end optimized inference pipelines across NVIDIA accelerators.Oversee performance tuning, profiling, and optimization of large-scale models for LLM, multimodal, and generative AI applications.Guide engineers in adopting best practices for CUDA, Triton, CUTLASS, and multi-GPU communications (NIXL, NCCL, NVSHMEM).Represent the team in roadmap and planning discussions, ensuring alignment with NVIDIA’s broader AI and software strategies.Foster a culture of technical excellence, open collaboration, and continuous innovation.What we need to see:MS, PhD, or equivalent experience in Computer Science, Electrical/Computer Engineering, or a related field.6+ overall years of software development experience, including 3+ years in technical leadership or engineering management.Strong background in C/C++ software design and development; proficiency in Python is a plus.Hands-on experience with GPU programming (CUDA, Triton, CUTLASS) and performance optimization.Proven record of deploying or optimizing deep learning models in production environments.Experience leading teams using Agile or collaborative software development practices.Ways to Stand out from The Crowd:Significant open-source contributions to deep learning or inference frameworks such as PyTorch, vLLM, SGLang, Triton, or TensorRT-LLM.Deep understanding of multi-GPU communications (NIXL, NCCL, NVSHMEM) and distributed inference architectures.Expertise in performance modeling, profiling, and system-level optimization across CPU and GPU platforms.Proven ability to mentor engineers, guide architectural decisions, and deliver complex projects with measurable impact.Publications, patents, or talks on LLM serving, model optimization, or GPU performance engineering.With highly competitive salaries and a comprehensive benefits package, NVIDIA is widely considered to be one of the technology world's most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us and, our rapid growth means endless opportunities for career advancement.If you’re a passionate technical leader ready to shape the future of AI inference frameworks — and build the software that powers the world’s most advanced models — we’d love to hear from you.#LI-Hybrid Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD for Level 3, and 272,000 USD - 431,250 USD for Level 4.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until August 1, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, GA, Remote; US, DC, Remote; US, IL, Remote; US, CA, Remote; US, MA, RemoteType: Full time

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Engineering Manager, Deep Learning Inference in Santa Clara, CA vacancy
  • $224k - $356.5k

     ...serving performance across various inference frameworks. Hyperscalers, cloud providers...  ..., and scaling. As Technical Lead Manager, you will lead the engineering team within NVIDIA’s Dynamo...  ...tech lead, TLM, or engineering manager.Deep understanding of LLM inference mechanics... 
    Suggested
    Full time
    Local area
    Remote work
    Worldwide

    Nvidia

    Santa Clara, CA
    4 days ago
  • $168k - $270.25k

     ...scale, turning rapidly evolving deep learning models into highly optimized...  ...programs for training and inference. As AI models, GPU...  ...reasoning, and large-scale systems engineering. To address these complex challenges...  ...are seeking an Engineering Manager to spearhead our strategy... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    8 hours ago
  • $207k - $301k

    People Management and Talent Development: Lead, mentor,...  ...team of systems and ML engineers. Drive a culture of...  ...safety, and continuous learning. Guide career paths, define...  ...experience utilizing deep-dive ML profiling...  ...Distributed Cloud (DSC) AI Inference Platform team operates... 
    Suggested

    Google

    Sunnyvale, CA
    2 days ago
  • $272k - $431.25k

     ...continues to increase, we are seeking outstanding engineers to join our team and help shape the future of LLM inference.Our team is dedicated to pushing the...  ...equivalent experience).15+ years of experience in deep learning and deep learning systems design.Proficiency in... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $272k - $431.25k

     ...intelligence—computers that can learn, reason, and interact with...  ...every industry. GPU-accelerated deep learning provides the foundation...  ...Principal Perception Engineer to lead the design and productization...  ...development and optimizing training or inference pipelines through custom CUDA... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    8 hours ago
  • $224k - $356.5k

    NVIDIA is seeking an Engineering Manager to lead the development of an agentic platform for observing...  ...visibility into model behavior, inference performance, reliability, and cost...  ...functional role for someone who understands deep learning systems, observability, and large-... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $272k - $431.25k

     ...and Benchmarking R&D group requires a senior software engineer. In this exciting role, you will profile, analyze, and...  ...large-scale GPU and CPU clusters used for distributed Deep Learning LLM training and inference. Your primary focus will be collectives communication and... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...us as we shape the future of AI and beyond. Together, we advance your career. THE ROLEWe are seeking a Principal GenAI Inference Optimization Engineer to join our Models and Applications team. This role focuses on improving performance, efficiency, and scalability of generative... 

    AMD

    San Jose, CA
    1 day ago
  • $320k

     ...the globe. This team is the execution engine behind NVIDIA’s Vision AI strategy—owning...  ...Engineering, who is hands-on with deep learning and comfortable reading/modeling code,...  ...record of delivering robust, low-latency inference at scale. You have led teams that turn... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $272k - $431.25k

     ...graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU...  ...advancement.Are you a motivated system software engineer with a deep understanding of device drivers who has phenomenal... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $272k - $431.25k

    At NVIDIA, we are seeking exceptional engineers to join our autonomous driving team to design...  ...generative, imitation, and reinforcement learning—to improve the planning and reasoning...  ...coder passionate about autonomous systems.Deep understanding of modern deep learning architectures... 
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    4 days ago
  • $201.3k - $352.3k

     ...DescriptionIt all started when engineer Fred Luddy wrote code...  ..., and product managers with a dual mission. We...  ...efficiency, and inference costs.Lead a High-Performing...  ...Core Product, Machine Learning Platforms, and...  ...frontier AI SDKs and deep experience deploying or... 
    Work experience placement
    Work at office
    Immediate start
    Remote work
    Flexible hours
    Shift work

    ServiceNow

    Santa Clara, CA
    3 days ago
  • $224k - $356.5k

    We are now looking for an AI Developer Technology Engineering Manager:Join our global Developer Technology (DevTech) team at NVIDIA, where...  ...engineers accelerating end-to-end performance of real-world Deep Learning and Machine Learning applications and developing novel... 
    Full time
    Temporary work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $272k - $431.25k

     ...them. NVIDIA seeks a Senior Engineering Manager to define and drive NVIDIA's...  ...across training, post-training, inference, and robotics, bridging new...  ..., and hardware teamsBuild deep partnerships with key open-...  ...runtime infrastructure for deep learning frameworks (JAX, PyTorch,... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...industry-leading training and inference speeds; over 10 times faster...  ...RoleWe're hiring a Principal Engineer for our Inference Cloud Platform...  ...or cloud infrastructure. Deep expertise in distributed systems...  ...best work through continuous learning, growth and support of those... 

    Cerebras Systems

    Sunnyvale, CA
    2 days ago
  • $336.4k - $381.6k

     ...spearhead our new Application Engineering team, a self-sufficient and...  ...ground up. As a founding manager, you'll lead a small but mighty...  ...working across robotics, machine learning, and systems integration. You...  ...essential. If you also bring deep expertise in machine learning... 
    Full time
    Local area
    Work from home

    Wayve

    Sunnyvale, CA
    18 hours ago
  • Achronix is looking for a dynamic engineer focused on PCB bring up, functional validation, measurements of PCBs, and lab management. These PCBs are used for AI inference applications and are integrated into Achronix data centers. The candidate will be defining and running... 
    Work at office
    Remote work

    Achronix Semiconductor Corporation

    Santa Clara, CA
    4 days ago
  • $224k - $356.5k

     ...and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning — the next era of computing —...  ...thoughtful people in the world.We are looking for an excellent engineering manager to own and deliver an end to end manageability stack for... 
    Full time
    Work at office

    Nvidia

    Santa Clara, CA
    2 days ago
  • $224k - $356.5k

     ...accelerating it. We are accelerating LLM inference across the stack and across all open...  ...re seeking a highly skilled and driven Engineering Manager to take the lead in accelerating the next...  ...leadership role at the intersection of deep technical expertise and world-class... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $190k

     ...entertainment. We are looking for a Manager to lead the Content Personalization Algorithms Engineering team. You will lead the way for a team of machine learning engineers and researchers to...  ...Search, or Recommender Systems. ~ Deep Learning, Ranking, LLMs, or... 
    Hourly pay
    Full time
    Immediate start
    Flexible hours

    Netflix

    Los Gatos, CA
    3 days ago
  • $272k - $431.25k

     ...NVIDIA is the engine of modern AI, and robotics is where...  ...engineering leader and manager to head up Isaac for...  ...grasping, motion, and learned skills they need to do...  ...training, and on-robot inference on Jetson and edge platforms...  ....Learned manipulation. Deep experience with... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $224k - $356.5k

     ...seeking a deeply technical software manager to lead production AI inference for NVIDIA Inference Microservices (...  ...stack, combining optimized inference engines, model profiles/recipes, validated...  ...research, security, and operations. ~ Deep understanding of AI/ML fundamentals,... 

    Socket.dev

    Santa Clara, CA
    3 days ago
  • $210.87k - $329.86k

     ...visionary Senior Principal AI Engineer to architect, design, and oversee...  ...autonomous AI agents that can manage multi-tool workflows across...  ...and Thermal Engineers to encode deep domain expertise, design rules...  ...(NLP), or advanced Machine Learning architectures.Hardware Lifecycle... 
    Temporary work
    Work at office
    Local area
    Worldwide
    Shift work

    Celestica

    San Jose, CA
    4 days ago
  • $268.6k - $395k

     ...discovery experiences. As a Principal Engineer, you will lead the technical direction...  ...techniques such as sequence modeling, deep learning, and large language models (LLMs). Your...  ...iteration. Partner closely with product managers, data scientists, and designers to... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash Usa

    Sunnyvale, CA
    12 hours ago
  • $272k - $431.25k

     ...infrastructure that stores, manages, and serves exabytes of data...  ...infrastructure, enabling researchers and engineers to reliably store massive...  ...accelerating training and inference pipelines.We are seeking a...  ...production services at scale.Deep technical background in distributed... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $278.1k - $347.6k

     ...within that runtime. As our Principal Engineer for On-Device AI Inference & Systems, you will be the foremost...  ...the render path, and frame-budget management alongside the renderer. Architect...  ...limits, and binding model. Equivalent deep experience with a native GPU/compute... 
    Work at office
    Worldwide
    Relocation package

    Unity

    Mountain View, CA
    2 days ago
  • $150k - $160k

     ...work in a high performing global company where employees collaborate and strive for excellence. Job Description As an Engineering Program Manager, you will work closely with an internal inter-disciplinary team to drive key aspects of product execution, test, and delivery... 
    Full time
    Temporary work
    Flexible hours

    XP Power

    San Jose, CA
    5 hours ago
  • $272k - $431.25k

     ...and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning - the next era of computing -...  ...people in the world.We are looking for an excellent Senior Engineering Manager to lead a large firmware engineering organization delivering... 
    Full time

    Nvidia

    Santa Clara, CA
    8 hours ago
  • $272k - $431.25k

     ...NVIDIA is seeking a Senior MLOps Engineering Manager to join our Autonomous Driving organization in Santa Clara, CA. This role offers an...  ...following domains: Autonomous Vehicles, Robotics, Computer Vision, Deep Learning, or GPU‑accelerated computing.Excellent communication and... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $224k - $356.5k

     ...and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning — the next era of computing —...  ...team, and we are looking for a highly motivated, creative Engineering Manager to drive Factory System Software and Diagnostics Integration... 
    Full time
    Shift work

    Nvidia

    Santa Clara, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Engineering Manager, Deep Learning Inference. Be the first to apply!