Engineering Manager, LLM Inference & Deployment at Scale
$224k - $356.5kNVIDIA
NVIDIA is looking for an Engineering Manager to lead the team responsible for deploying and serving Large Language Models (LLMs) and Vision-Language Models (VLMs) at scale.Our team builds and operates an AI inference platform that enables customers to deploy and run brand new generative AI models efficiently across NVIDIA GPU platforms. The platform operates at the intersection of model optimization, inference systems, distributed computing, and production infrastructure. In this role, you will manage a team of engineers responsible for taking modern models from research and experimentation into highly optimized, reliable, and scalable production deployments. You will work closely with research scientists, software engineers, and hardware specialists to push the boundaries of AI inference performance and deliver a world-class model deployment platform.What you will be doing:Lead, mentor, and grow a high-performing team building and operating a platform for deploying GenAI models at scale.Drive the deployment and optimization of LLMs and VLMs for low-latency, high-throughput, and cost-efficient inference.Analyze, profile, and optimize end-to-end deep learning workloads across the model, inference stack, distributed systems, and GPU hardware.Work closely with research teams and model developers to bring new architectures and models from prototype to production.Drive technical strategy and roadmap for model deployment, inference optimization, scalability, reliability, and performance.Establish engineering standards for benchmarking, profiling, production deployment, and continuous performance optimization.Collaborate with internal and external partners to enable seamless deployment of rapidly evolving GenAI models..What we want to see:Master’s or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.8+ years overall experience, including 3 years of management/leadership.Strong hands-on experience with LLMs and/or VLMs and a solid understanding of modern deep learning architectures.Deep expertise in inference optimization techniques such as quantization, speculative decoding, continuous batching, prefix caching, and KV-cache optimization.Experience with disaggregated inference/serving, distributed inference, multi-node deployments, and GPU cluster orchestration.Strong technical leadership, communication, and people-management skills.Ways to stand out from the crowd:Experience building or operating AI inference, model serving, or model deployment platforms.Practical experience in working with TensorRT, TensorRT-LLM, vLLM, SGLang, or comparable inference/serving frameworks.Experience building highly available, scalable, and observable production services for AI workloads.NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you!#LI-HybridYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD for Level 3, and 272,000 USD - 431,250 USD for Level 4.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until September 1, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time
$224k - $356.5k
NVIDIA is seeking an Engineering Manager to lead the development of an agentic... ...optimizing GenAI models deployed at scale. In this role, you will... ...visibility into model behavior, inference performance, reliability,... ...signals across large scale LLM and VLM deployments. It will...SuggestedFull time$224k - $356.5k
...seeking an exceptional Manager, Deep Learning Inference Software, to lead a world-class engineering team advancing the state of AI model deployment. You will shape the... ...time inference at every scale, from datacenter clusters... ...-scale models for LLM, multimodal, and generative...SuggestedFull timeRemote workWorldwide$224k - $356.5k
...standard for assessing LLM serving performance across various inference frameworks. Hyperscalers... ...improving efficiency, and scaling. As Technical Lead Manager, you will lead the engineering team within NVIDIA’s... ..., and Kubernetes-native deployment.Taking ownership for the...SuggestedFull timeLocal areaRemote workWorldwide$224k - $356.5k
...accelerating it. We are accelerating LLM inference across the stack and across... ...highly skilled and driven Engineering Manager to take the lead in... ...experience for LLM deployment. Lead software development... ...Proven ability to lead and scale high-performing engineering...SuggestedFull time- ...ROLEWe are seeking a Principal GenAI Inference Optimization Engineer to join our Models and... ...and cost efficiency for real-world deployment of large-scale models, working across the software... ...throughput, and cost efficiency for LLM and multimodal model serving in production...Suggested
$207k - $300k
...infrastructure (e.g., model deployment, model evaluation,... ...in a people management or team leadership role... ...Master’s degree or PhD in Engineering, Computer Science, or... ...distributed computing, large-scale system design,... ...supporting training and inference across Alphabet. You will...Immediate startWorldwide- ...deliver industry-leading training and inference speeds; over 10 times faster than... ...-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra... ...The RoleWe're hiring a Principal Engineer for our Inference Cloud Platform. This...
$201k - $402k
...OpportunityWe are looking for a Senior Engineering Manager to lead the design, development, and scaling of a next-generation... ...capabilities in on-prem and hybrid deployments.You will lead a team... ...GPU scheduling, training, and inference systemsExecution & Operational...Work at officeLocal areaRemote workRelocation package3 days per week$220k - $300k
...Principal Machine Learning Engineer San Jose,... ...and operations at scale. SambaNova Suite... ...training and fine-tuning, inference optimization,... ...bridges advanced LLM research and practical deployment, involving the development... ...without direct management authority ~ Track...Full timeTemporary workLocal areaFlexible hours$220k - $300k
...delivering a full-stack inference platform for... ...into SambaRack, rack-scale hardware that lets customers deploy state-of-the-art models... ...Machine Learning Engineer, you will be... ...role bridges advanced LLM research and practical... ...without direct management authority ~ Track...Full timeTemporary workLocal areaWorldwideFlexible hours$195k - $285k
...compute. As a Principal Hardware Design Engineer, you will be a cornerstone of our... ...solve the industry's most massive LLM inference challenges.Required Qualifications•... ...Compute Project) standards or large-scale data center deployments.• Knowledge of liquid cooling integration...3 days per week$262.5k - $394k
...Principal Data Engineering Lead - Services Special... ...at a larger scale, coordinating and... ...lineage, metadata management, schema governance... ...development, testing, deployment, and operations.... ...integrate ML model inference - including LLMs and... ...TensorRT/TensorRT-LLM), and serving frameworks...Relocation$272k - $431.25k
...exceptional Principal Perception Engineer to lead the design and... ...using techniques such as large-scale pretraining, distillation, and... ...robustness, and are ready for deployment at scale.Provide technical... ...development and optimizing training or inference pipelines through custom CUDA...Full time$240k - $290k
...fast-growing company with the scale and impact of an established... ...and proven by global deployment, we’re solving some of the most... ...looking for an experienced Senior Engineering Manager to build and lead our AI... ...well as our cloud-only AI and LLM-based products. You will...Work at officeWorldwideFlexible hoursShift work$207k - $300k
...driving complex 1-6 month Engineers and 2-4 week Strike... ...infrastructure (e.g., model deployment, model evaluation, data... ...generative AI tools or LLM interfaces into... ...distributed enterprise scale (Franchises) via structured... .... Software Engineering Managers have not only the technical...Permanent employment$249k
...everywhere.Director, Security Engineering Our Technology Team... ...at enterprise scale. This is fundamentally... ...and maintain secrets management, certificate lifecycle... ...security architecture for LLM-based applications, RAG... ...tuning practices, and deployment architecture through...Full timeShift workDay shift- ...Tuesday.About the RoleFleet Engineering owns the full lifecycle of Lambda... ...— new product introduction, deployment, operation, and reliability... ...provisioning, firmware management, out-of-band access, and power... ...while driving efficiency at scale.We are hiring multiple...Work at officeLocal areaWork from homeFlexible hours
$114.6k - $234.6k
...that operate reliably at cloud scale.Only Oracle brings together... ...measurement, analytics, and inference systems for data center and... ...analysis while maintaining engineering, security, and quality standards... ..., training and evaluation, deployment, versioning, monitoring, and...Temporary workFlexible hours$140k - $215k
...AI-native platform. We work on large scale distributed systems, processing almost... ...CrowdStrike is seeking an experienced SDET Engineering Manager to lead our quality engineering... ...and integrating automated testing into deployment workflowsExperience developing and deploying...Full timeContract workWork experience placementWork at officeLocal areaWorldwide2 days per week3 days per week$180k - $230k
...operating as one integrated managed service. The Company... ...seeking a Principal Engineer to own the end-to-end... ...internal tools (KNOC, Deployment, Service, Calibration,... ...video analytics, AI inference, digital twin engines,... ...experience architecting large-scale, real-time distributed...Full time$141k - $221k
...Country: USA SummaryThe Senior Staff Engineer, Program/Project Management is accountable for planning and... ...customer feedbackCoordinate site-wide deployment efforts.Implement change as directed... ...development - from drawing board to full-scale production and after-market services...Temporary workWork at officeLocal areaRemote workRelocation$169k - $338k
...including data sourcing, feature engineering, model training, deployment, and monitoring.Champion... ...end-to-end ML lifecycle management, including feature... ...Our team develops and scales AI-driven pricing solutions... ...machine learning, causal inference, and optimization techniques...Full timeTemporary workPart timeWork experience placement- ...Machine Learning Engineer, you will drive the... ...ExperimentationDesign, develop, and deploy production-grade... ..., retrieval, LLM-based systems) to... ...and online inference at scaleOversee... ...closely with product managers, designers, and... ...programs of work that scale across the...Local areaRemote work
$206.4k - $379.1k
...Foundry is Adobe's enterprise managed-service offering for... ...media-intelligence layer, and deployed across new and existing... ...Principal Machine Learning Engineer to serve as the technical... ...and served at enterprise scale. You will set the inference architecture and technical...Full timeTemporary workLocal areaWorldwide$224k - $356.5k
...the power and flexibility to develop and deploy breakthrough artificial intelligence... ...platforms — from data center to vehicle — and scale effortlessly from simple deployments to... ...pipeline. We are seeking a seasoned Engineering Manager to lead our rapidly growing NvStreams...Full time$206.4k - $379.1k
...seeking a Principal Service Engineer to serve as the... ...while productizing and scaling a rapidly growing portfolio... ....Design and architect inference infrastructure for... ...with Product Managers, TPMs, and engineering... ...management in large-scale deployments.Hands-on expertise with...Full timeTemporary workLocal areaWorldwide- ...the future of transportation on a global scale. The Data Scaling team owns the Data... ...collaborative, high-impact team of AI/ML engineers, data scientists and engineers who are... ...pivotal in designing, architecting, and deploying data curation and training pipelines while...Full timeLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours
- ..., more than 35 live GPU cluster deployments globally, and backing from a top... ...Silicon Valley VC, the company is scaling rapidly across its product and engineering organization. This... ...partnering closely with Product Management, Architecture, and Customer Success...Full time
$60 - $70 per hour
Engineering Program Manager W2 Contract Pay Rate: $60 - $70 per hour Location: Cupertino, CA - Hybrid... ...Understanding of software development and deployment lifecycles. Familiarity with CI/CD,... ..., Alibaba Cloud, or other large-scale cloud environments. Experience with...Hourly payContract work$300k - $325k
...travel, primarily supporting engineering and manufacturing operations... ...product development, manufacturing scale-up, and commercialization of... ...costs, and accelerating deployment with Tier 1 customers. This... ...guiding architecture decisions, managing a large multidisciplinary...RelocationRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Engineering Manager, LLM Inference & Deployment at Scale. Be the first to apply!


