Engineering Manager, LLM Inference & Deployment at Scale
$224k - $356.5kNVIDIA
NVIDIA is looking for an Engineering Manager to lead the team responsible for deploying and serving Large Language Models (LLMs) and Vision-Language Models (VLMs) at scale.Our team builds and operates an AI inference platform that enables customers to deploy and run brand new generative AI models efficiently across NVIDIA GPU platforms. The platform operates at the intersection of model optimization, inference systems, distributed computing, and production infrastructure. In this role, you will manage a team of engineers responsible for taking modern models from research and experimentation into highly optimized, reliable, and scalable production deployments. You will work closely with research scientists, software engineers, and hardware specialists to push the boundaries of AI inference performance and deliver a world-class model deployment platform.What you will be doing:Lead, mentor, and grow a high-performing team building and operating a platform for deploying GenAI models at scale.Drive the deployment and optimization of LLMs and VLMs for low-latency, high-throughput, and cost-efficient inference.Analyze, profile, and optimize end-to-end deep learning workloads across the model, inference stack, distributed systems, and GPU hardware.Work closely with research teams and model developers to bring new architectures and models from prototype to production.Drive technical strategy and roadmap for model deployment, inference optimization, scalability, reliability, and performance.Establish engineering standards for benchmarking, profiling, production deployment, and continuous performance optimization.Collaborate with internal and external partners to enable seamless deployment of rapidly evolving GenAI models..What we want to see:Master’s or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.8+ years overall experience, including 3 years of management/leadership.Strong hands-on experience with LLMs and/or VLMs and a solid understanding of modern deep learning architectures.Deep expertise in inference optimization techniques such as quantization, speculative decoding, continuous batching, prefix caching, and KV-cache optimization.Experience with disaggregated inference/serving, distributed inference, multi-node deployments, and GPU cluster orchestration.Strong technical leadership, communication, and people-management skills.Ways to stand out from the crowd:Experience building or operating AI inference, model serving, or model deployment platforms.Practical experience in working with TensorRT, TensorRT-LLM, vLLM, SGLang, or comparable inference/serving frameworks.Experience building highly available, scalable, and observable production services for AI workloads.NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you!#LI-HybridYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD for Level 3, and 272,000 USD - 431,250 USD for Level 4.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until September 1, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time
$224k - $356.5k
...standard for assessing LLM serving performance across various inference frameworks. Hyperscalers... ...improving efficiency, and scaling. As Technical Lead Manager, you will lead the engineering team within NVIDIA’s... ..., and Kubernetes-native deployment.Taking ownership for the...SuggestedFull timeLocal areaRemote workWorldwide$224k - $356.5k
...seeking an exceptional Manager, Deep Learning Inference Software, to lead a world-class engineering team advancing the state of AI model deployment. You will shape the... ...time inference at every scale, from datacenter clusters... ...-scale models for LLM, multimodal, and generative...SuggestedFull timeRemote workWorldwide$224k - $356.5k
...accelerating it. We are accelerating LLM inference across the stack and across... ...highly skilled and driven Engineering Manager to take the lead in... ...experience for LLM deployment. Lead software development... ...Proven ability to lead and scale high-performing engineering...SuggestedFull time- ...ROLEWe are seeking a Principal GenAI Inference Optimization Engineer to join our Models and... ...and cost efficiency for real-world deployment of large-scale models, working across the software... ...throughput, and cost efficiency for LLM and multimodal model serving in production...Suggested
$201.3k - $352.3k
...all started when engineer Fred Luddy wrote code... ..., and product managers with a dual mission... ...Design, own, and scale automated testing... ...window efficiency, and inference costs.Lead a High-... ...deep experience deploying or testing agentic... ...ROUGE, BLEU, G-Eval, LLM-as-a-judge...SuggestedWork experience placementWork at officeImmediate startRemote workFlexible hoursShift work- A leading technology company in California is seeking an Engineering Manager to lead the development of cutting-edge LLM/VLM technologies. In this hands-on leadership role, you will manage a team responsible for optimizing runtime and frameworks, while collaborating with...
$201k - $402k
...OpportunityWe are looking for a Senior Engineering Manager to lead the design, development, and scaling of a next-generation... ...capabilities in on-prem and hybrid deployments.You will lead a team... ...GPU scheduling, training, and inference systemsExecution & Operational...Work at officeLocal areaRemote workRelocation package3 days per week$154k - $312k
...platform, its Zero Trust Engine, and the powerful... ...research and enterprise-scale business impact. Leveraging... ...the architecture and deployment of large-scale, production... ...Production-Grade Inference Systems: Design, optimize... ...leveraging the latest LLM serving technologies such...- ...deliver industry-leading training and inference speeds; over 10 times faster than... ...-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra... ...The RoleWe're hiring a Principal Engineer for our Inference Cloud Platform. This...
$224k - $356.5k
...systems.We are looking for an Engineering Manager to build and lead the team... ...spanning agent runtimes, local inference, Windows integration,... ...software, and local or hybrid deployment, with the judgment to guide... ...patterns.Experience with TensorRT-LLM, Ollama, llama.cpp, vLLM,...Full timeLocal area$171.6k - $245k
...The team works across product management, engineering, applied AI/ML, data,... ...customer-ready capabilities at scale. This is a highly multi-functional... ...continuous integration and deployment.Communicate program status,... ...at least one generative AI, LLM-based, or agentic capability...Full timeTemporary workLocal areaFlexible hours$270k - $300k
...Senior Engineering Manager, MLOpsPalo Alto, California, United... ...company that is rapidly scaling while maintaining a... ...to automated deployment and real-time monitoring... ...infrastructure cost per inference.Champion Operational... ...Monitor emerging trends in LLM-ops, vector databases...Seasonal workLocal area$240.1k - $420.2k
...DescriptionIt all started when engineer Fred Luddy wrote code that... ...agentic AI and enterprise-scale search systems that power... ...autonomous systems safe to deploy at scale. Retrieval and grounding... ...or retrieval. Exposure to LLM fine-tuning or inference optimization in...Work experience placementWork at officeImmediate startRemote workFlexible hours$272k - $431.25k
...Machine Learning (ML) Engineer to join the GPU accelerated... ...for running large scale workloads for ETL, SQL,... .../DL model training and inference pipelines, spanning many... ...partners and customers on the deployment of complex machine... ..., including LLM/GenAI, reinforcement learning...Full time$272k - $425.5k
A leading technology firm in Santa Clara seeks a Principal Software Engineer for Large-Scale LLM Memory and Storage Systems. The role involves designing a unified memory layer for large-scale inference, integrating with LLM serving engines, and mentoring engineers. Candidates...Remote job$195k - $285k
...compute. As a Principal Hardware Design Engineer, you will be a cornerstone of our... ...solve the industry's most massive LLM inference challenges.Required Qualifications•... ...Compute Project) standards or large-scale data center deployments.• Knowledge of liquid cooling integration...3 days per week$262.5k - $394k
...Principal Data Engineering Lead - Services Special... ...at a larger scale, coordinating and... ...lineage, metadata management, schema governance... ...development, testing, deployment, and operations.Align... ...ML model inference - including LLMs and... ...TensorRT/TensorRT-LLM), and serving frameworks...Relocation- ...Tuesday.About the RoleFleet Engineering owns the full lifecycle of Lambda... ...— new product introduction, deployment, operation, and reliability... ...provisioning, firmware management, out-of-band access, and power... ...while driving efficiency at scale.We are hiring multiple...Work at officeLocal areaWork from homeFlexible hours
$185.1k - $284.1k
...everyday drivers at unprecedented scale. Join us to help deliver the next generation... ...results into clear feedback for engineering and leadership, and help accelerate validated AV deployment at scale. We are looking for an Engineering Manager to lead a team building the...Full timeLocal areaRemote workWork from homeFlexible hours$240k - $290k
...fast-growing company with the scale and impact of an established... ...and proven by global deployment, we’re solving some of the most... ...looking for an experienced Senior Engineering Manager to build and lead our AI... ...well as our cloud-only AI and LLM-based products. You will...Work at officeWorldwideFlexible hoursShift work$272k - $431.25k
...exceptional Principal Perception Engineer to lead the design and... ...using techniques such as large-scale pretraining, distillation, and... ...robustness, and are ready for deployment at scale.Provide technical... ...development and optimizing training or inference pipelines through custom CUDA...Full time$225k - $325k
...impact on a global scale. At Goodwin, we... ...peers. Director of Engineering The Director of... ..., and production deployment of AI-driven... ...system latency, model inference performance, and infrastructure... ..., product management, and business... ...system architectures, LLM orchestration...$141k - $221k
...Country: USA SummaryThe Senior Staff Engineer, Program/Project Management is accountable for planning and... ...customer feedbackCoordinate site-wide deployment efforts.Implement change as directed... ...development - from drawing board to full-scale production and after-market services...Temporary workWork at officeLocal areaRemote workRelocation- ...THIS FEATURED OPPORTUNITY The Engineering Program Manager will join the Channel Strategy &... ...teams and ensure on-time delivery and deployment of key initiatives Define milestones... ...managing projects involving LLM-based products or AI-driven systems...Local areaFlexible hours
$250k - $300k
.... We're hiring a strong engineering leader to own the Cloud &... ...Tech Lead ready to step into management OR a Manager who still codes... ...for cloud infrastructure, deployments, and platform capabilities... ...environments and production systems at scale Drive improvements in...Work at officeShift work$140k - $215k
...AI-native platform. We work on large scale distributed systems, processing almost... ...CrowdStrike is seeking an experienced SDET Engineering Manager to lead our quality engineering... ...and integrating automated testing into deployment workflowsExperience developing and deploying...Full timeContract workWork experience placementWork at officeLocal areaWorldwide2 days per week3 days per week$160k - $240k
...Senior Engineering Manager (Backend & Platform)Fiserv is a global leader in payments and financial... ...processing financial transactions at scale.Establish operational excellence,... ...guarantee platform reliability, zero-downtime deployments, and fast incident recovery.Direct...Contract workWork at officeWorldwideMonday to Friday$237.6k - $356.4k
...Engineering Manager, FileMaker ServerAt Apple, great ideas have a way of becoming great products... ...platform technologies that power enterprise-scale solutions and directly impact customer... ...(APIs), and cloud or on-premises deployments.Review designs and code to ensure...Relocation$204k - $343k
...flexibility and trust our employees to manage their schedules responsibly.... .... About the role As an Engineering Manager on the Data... ...cutting‑edge AI techniques to scale these capabilities, setting technical... ...development, safety, and deployment milestones. At Applied Intuition...Full timeFor contractorsFor subcontractorCasual workWork at officeRemote workDay shift- ...CrowdStrike SDET Engineering ManagerAs a global leader in cybersecurity... ...platform. We work on large scale distributed systems, processing... ...experienced SDET Engineering Manager to lead our quality... ...integrating automated testing into deployment workflowsExperience developing...Contract workWork at officeWorldwide2 days per week3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Engineering Manager, LLM Inference & Deployment at Scale. Be the first to apply!

