Engineering Manager, Deep Learning Inference
$224k - $356.5kNVIDIA
NVIDIA is seeking an exceptional Manager, Deep Learning Inference Software, to lead a world-class engineering team advancing the state of AI model deployment. You will shape the software powering today’s most sophisticated AI systems — from large language models to multimodal generative AI — all accelerated on NVIDIA GPUs. The Deep Learning Inference team develops and optimizes open-source frameworks that make AI deployment scalable, efficient, and accessible — including SGLang, vLLM, and FlashInfer. Our work enables developers worldwide to harness NVIDIA accelerators for real-time inference at every scale, from datacenter clusters to edge devices.What you'll be doing:Lead, mentor, and scale a high-performing engineering team focused on deep learning inference and GPU-accelerated software.Guide the strategy, roadmap, and execution of NVIDIA's OSS inference frameworks engineering.Partner with internal compiler, libraries, and research teams to deliver end-to-end optimized inference pipelines across NVIDIA accelerators.Oversee performance tuning, profiling, and optimization of large-scale models for LLM, multimodal, and generative AI applications.Guide engineers in adopting best practices for CUDA, Triton, CUTLASS, and multi-GPU communications (NIXL, NCCL, NVSHMEM).Represent the team in roadmap and planning discussions, ensuring alignment with NVIDIA’s broader AI and software strategies.Foster a culture of technical excellence, open collaboration, and continuous innovation.What we need to see:MS, PhD, or equivalent experience in Computer Science, Electrical/Computer Engineering, or a related field.6+ overall years of software development experience, including 3+ years in technical leadership or engineering management.Strong background in C/C++ software design and development; proficiency in Python is a plus.Hands-on experience with GPU programming (CUDA, Triton, CUTLASS) and performance optimization.Proven record of deploying or optimizing deep learning models in production environments.Experience leading teams using Agile or collaborative software development practices.Ways to Stand out from The Crowd:Significant open-source contributions to deep learning or inference frameworks such as PyTorch, vLLM, SGLang, Triton, or TensorRT-LLM.Deep understanding of multi-GPU communications (NIXL, NCCL, NVSHMEM) and distributed inference architectures.Expertise in performance modeling, profiling, and system-level optimization across CPU and GPU platforms.Proven ability to mentor engineers, guide architectural decisions, and deliver complex projects with measurable impact.Publications, patents, or talks on LLM serving, model optimization, or GPU performance engineering.With highly competitive salaries and a comprehensive benefits package, NVIDIA is widely considered to be one of the technology world's most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us and, our rapid growth means endless opportunities for career advancement.If you’re a passionate technical leader ready to shape the future of AI inference frameworks — and build the software that powers the world’s most advanced models — we’d love to hear from you.#LI-Hybrid Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD for Level 3, and 272,000 USD - 431,250 USD for Level 4.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until August 1, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, GA, Remote; US, DC, Remote; US, IL, Remote; US, CA, Remote; US, MA, RemoteType: Full time
$224k - $356.5k
...serving performance across various inference frameworks. Hyperscalers, cloud providers... ..., and scaling. As Technical Lead Manager, you will lead the engineering team within NVIDIA’s Dynamo... ...tech lead, TLM, or engineering manager.Deep understanding of LLM inference mechanics...SuggestedFull timeLocal areaRemote workWorldwide$168k - $270.25k
...scale, turning rapidly evolving deep learning models into highly optimized... ...programs for training and inference. As AI models, GPU... ...reasoning, and large-scale systems engineering. To address these complex challenges... ...are seeking an Engineering Manager to spearhead our strategy...SuggestedFull time$207k - $301k
People Management and Talent Development: Lead, mentor,... ...team of systems and ML engineers. Drive a culture of... ...safety, and continuous learning. Guide career paths, define... ...experience utilizing deep-dive ML profiling... ...Distributed Cloud (DSC) AI Inference Platform team operates...Suggested$272k - $431.25k
...continues to increase, we are seeking outstanding engineers to join our team and help shape the future of LLM inference.Our team is dedicated to pushing the... ...equivalent experience).15+ years of experience in deep learning and deep learning systems design.Proficiency in...SuggestedFull time$272k - $431.25k
...intelligence—computers that can learn, reason, and interact with... ...every industry. GPU-accelerated deep learning provides the foundation... ...Principal Perception Engineer to lead the design and productization... ...development and optimizing training or inference pipelines through custom CUDA...SuggestedFull time$224k - $356.5k
NVIDIA is seeking an Engineering Manager to lead the development of an agentic platform for observing... ...visibility into model behavior, inference performance, reliability, and cost... ...functional role for someone who understands deep learning systems, observability, and large-...Full time$272k - $431.25k
...and Benchmarking R&D group requires a senior software engineer. In this exciting role, you will profile, analyze, and... ...large-scale GPU and CPU clusters used for distributed Deep Learning LLM training and inference. Your primary focus will be collectives communication and...Full timeRemote work- ...us as we shape the future of AI and beyond. Together, we advance your career. THE ROLEWe are seeking a Principal GenAI Inference Optimization Engineer to join our Models and Applications team. This role focuses on improving performance, efficiency, and scalability of generative...
$320k
...the globe. This team is the execution engine behind NVIDIA’s Vision AI strategy—owning... ...Engineering, who is hands-on with deep learning and comfortable reading/modeling code,... ...record of delivering robust, low-latency inference at scale. You have led teams that turn...Full time$272k - $431.25k
...graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU... ...advancement.Are you a motivated system software engineer with a deep understanding of device drivers who has phenomenal...Full time$272k - $431.25k
At NVIDIA, we are seeking exceptional engineers to join our autonomous driving team to design... ...generative, imitation, and reinforcement learning—to improve the planning and reasoning... ...coder passionate about autonomous systems.Deep understanding of modern deep learning architectures...Full timeWork experience placement$201.3k - $352.3k
...DescriptionIt all started when engineer Fred Luddy wrote code... ..., and product managers with a dual mission. We... ...efficiency, and inference costs.Lead a High-Performing... ...Core Product, Machine Learning Platforms, and... ...frontier AI SDKs and deep experience deploying or...Work experience placementWork at officeImmediate startRemote workFlexible hoursShift work$224k - $356.5k
We are now looking for an AI Developer Technology Engineering Manager:Join our global Developer Technology (DevTech) team at NVIDIA, where... ...engineers accelerating end-to-end performance of real-world Deep Learning and Machine Learning applications and developing novel...Full timeTemporary work$272k - $431.25k
...them. NVIDIA seeks a Senior Engineering Manager to define and drive NVIDIA's... ...across training, post-training, inference, and robotics, bridging new... ..., and hardware teamsBuild deep partnerships with key open-... ...runtime infrastructure for deep learning frameworks (JAX, PyTorch,...Full time- ...industry-leading training and inference speeds; over 10 times faster... ...RoleWe're hiring a Principal Engineer for our Inference Cloud Platform... ...or cloud infrastructure. Deep expertise in distributed systems... ...best work through continuous learning, growth and support of those...
$336.4k - $381.6k
...spearhead our new Application Engineering team, a self-sufficient and... ...ground up. As a founding manager, you'll lead a small but mighty... ...working across robotics, machine learning, and systems integration. You... ...essential. If you also bring deep expertise in machine learning...Full timeLocal areaWork from home- Achronix is looking for a dynamic engineer focused on PCB bring up, functional validation, measurements of PCBs, and lab management. These PCBs are used for AI inference applications and are integrated into Achronix data centers. The candidate will be defining and running...Work at officeRemote work
$224k - $356.5k
...and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning — the next era of computing —... ...thoughtful people in the world.We are looking for an excellent engineering manager to own and deliver an end to end manageability stack for...Full timeWork at office$224k - $356.5k
...accelerating it. We are accelerating LLM inference across the stack and across all open... ...re seeking a highly skilled and driven Engineering Manager to take the lead in accelerating the next... ...leadership role at the intersection of deep technical expertise and world-class...Full time$190k
...entertainment. We are looking for a Manager to lead the Content Personalization Algorithms Engineering team. You will lead the way for a team of machine learning engineers and researchers to... ...Search, or Recommender Systems. ~ Deep Learning, Ranking, LLMs, or...Hourly payFull timeImmediate startFlexible hours$272k - $431.25k
...NVIDIA is the engine of modern AI, and robotics is where... ...engineering leader and manager to head up Isaac for... ...grasping, motion, and learned skills they need to do... ...training, and on-robot inference on Jetson and edge platforms... ....Learned manipulation. Deep experience with...Full time$224k - $356.5k
...seeking a deeply technical software manager to lead production AI inference for NVIDIA Inference Microservices (... ...stack, combining optimized inference engines, model profiles/recipes, validated... ...research, security, and operations. ~ Deep understanding of AI/ML fundamentals,...$210.87k - $329.86k
...visionary Senior Principal AI Engineer to architect, design, and oversee... ...autonomous AI agents that can manage multi-tool workflows across... ...and Thermal Engineers to encode deep domain expertise, design rules... ...(NLP), or advanced Machine Learning architectures.Hardware Lifecycle...Temporary workWork at officeLocal areaWorldwideShift work$268.6k - $395k
...discovery experiences. As a Principal Engineer, you will lead the technical direction... ...techniques such as sequence modeling, deep learning, and large language models (LLMs). Your... ...iteration. Partner closely with product managers, data scientists, and designers to...Hourly payWork at officeLocal areaRemote workFlexible hours$272k - $431.25k
...infrastructure that stores, manages, and serves exabytes of data... ...infrastructure, enabling researchers and engineers to reliably store massive... ...accelerating training and inference pipelines.We are seeking a... ...production services at scale.Deep technical background in distributed...Full time$278.1k - $347.6k
...within that runtime. As our Principal Engineer for On-Device AI Inference & Systems, you will be the foremost... ...the render path, and frame-budget management alongside the renderer. Architect... ...limits, and binding model. Equivalent deep experience with a native GPU/compute...Work at officeWorldwideRelocation package$150k - $160k
...work in a high performing global company where employees collaborate and strive for excellence. Job Description As an Engineering Program Manager, you will work closely with an internal inter-disciplinary team to drive key aspects of product execution, test, and delivery...Full timeTemporary workFlexible hours$272k - $431.25k
...and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning - the next era of computing -... ...people in the world.We are looking for an excellent Senior Engineering Manager to lead a large firmware engineering organization delivering...Full time$272k - $431.25k
...NVIDIA is seeking a Senior MLOps Engineering Manager to join our Autonomous Driving organization in Santa Clara, CA. This role offers an... ...following domains: Autonomous Vehicles, Robotics, Computer Vision, Deep Learning, or GPU‑accelerated computing.Excellent communication and...Full time$224k - $356.5k
...and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning — the next era of computing —... ...team, and we are looking for a highly motivated, creative Engineering Manager to drive Factory System Software and Diagnostics Integration...Full timeShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Engineering Manager, Deep Learning Inference. Be the first to apply!



