Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Engineering Manager, LLM Inference & Deployment at Scale

$224k - $356.5k
Full-time

NVIDIA

NVIDIA is looking for an Engineering Manager to lead the team responsible for deploying and serving Large Language Models (LLMs) and Vision-Language Models (VLMs) at scale.Our team builds and operates an AI inference platform that enables customers to deploy and run brand new generative AI models efficiently across NVIDIA GPU platforms. The platform operates at the intersection of model optimization, inference systems, distributed computing, and production infrastructure. In this role, you will manage a team of engineers responsible for taking modern models from research and experimentation into highly optimized, reliable, and scalable production deployments. You will work closely with research scientists, software engineers, and hardware specialists to push the boundaries of AI inference performance and deliver a world-class model deployment platform.What you will be doing:Lead, mentor, and grow a high-performing team building and operating a platform for deploying GenAI models at scale.Drive the deployment and optimization of LLMs and VLMs for low-latency, high-throughput, and cost-efficient inference.Analyze, profile, and optimize end-to-end deep learning workloads across the model, inference stack, distributed systems, and GPU hardware.Work closely with research teams and model developers to bring new architectures and models from prototype to production.Drive technical strategy and roadmap for model deployment, inference optimization, scalability, reliability, and performance.Establish engineering standards for benchmarking, profiling, production deployment, and continuous performance optimization.Collaborate with internal and external partners to enable seamless deployment of rapidly evolving GenAI models..What we want to see:Master’s or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.8+ years overall experience, including 3 years of management/leadership.Strong hands-on experience with LLMs and/or VLMs and a solid understanding of modern deep learning architectures.Deep expertise in inference optimization techniques such as quantization, speculative decoding, continuous batching, prefix caching, and KV-cache optimization.Experience with disaggregated inference/serving, distributed inference, multi-node deployments, and GPU cluster orchestration.Strong technical leadership, communication, and people-management skills.Ways to stand out from the crowd:Experience building or operating AI inference, model serving, or model deployment platforms.Practical experience in working with TensorRT, TensorRT-LLM, vLLM, SGLang, or comparable inference/serving frameworks.Experience building highly available, scalable, and observable production services for AI workloads.NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you!#LI-HybridYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD for Level 3, and 272,000 USD - 431,250 USD for Level 4.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until September 1, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Engineering Manager, LLM Inference & Deployment at Scale in Santa Clara, CA vacancy
  • $224k - $356.5k

    NVIDIA is seeking an Engineering Manager to lead the development of an agentic...  ...optimizing GenAI models deployed at scale. In this role, you will...  ...visibility into model behavior, inference performance, reliability,...  ...signals across large scale LLM and VLM deployments. It will... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $224k - $356.5k

     ...seeking an exceptional Manager, Deep Learning Inference Software, to lead a world-class engineering team advancing the state of AI model deployment. You will shape the...  ...time inference at every scale, from datacenter clusters...  ...-scale models for LLM, multimodal, and generative... 
    Suggested
    Full time
    Remote work
    Worldwide

    Nvidia

    Santa Clara, CA
    17 hours ago
  • $224k - $356.5k

     ...standard for assessing LLM serving performance across various inference frameworks. Hyperscalers...  ...improving efficiency, and scaling. As Technical Lead Manager, you will lead the engineering team within NVIDIA’s...  ..., and Kubernetes-native deployment.Taking ownership for the... 
    Suggested
    Full time
    Local area
    Remote work
    Worldwide

    Nvidia

    Santa Clara, CA
    1 day ago
  • $224k - $356.5k

     ...accelerating it. We are accelerating LLM inference across the stack and across...  ...highly skilled and driven Engineering Manager to take the lead in...  ...experience for LLM deployment. Lead software development...  ...Proven ability to lead and scale high-performing engineering... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    17 hours ago
  •  ...ROLEWe are seeking a Principal GenAI Inference Optimization Engineer to join our Models and...  ...and cost efficiency for real-world deployment of large-scale models, working across the software...  ...throughput, and cost efficiency for LLM and multimodal model serving in production... 
    Suggested

    AMD

    San Jose, CA
    3 days ago
  • $207k - $300k

     ...infrastructure (e.g., model deployment, model evaluation,...  ...in a people management or team leadership role...  ...Master’s degree or PhD in Engineering, Computer Science, or...  ...distributed computing, large-scale system design,...  ...supporting training and inference across Alphabet. You will... 
    Immediate start
    Worldwide

    Google

    Sunnyvale, CA
    1 day ago
  •  ...deliver industry-leading training and inference speeds; over 10 times faster than...  ...-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra...  ...The RoleWe're hiring a Principal Engineer for our Inference Cloud Platform. This... 

    Cerebras Systems

    Sunnyvale, CA
    4 days ago
  • $201k - $402k

     ...OpportunityWe are looking for a Senior Engineering Manager to lead the design, development, and scaling of a next-generation...  ...capabilities in on-prem and hybrid deployments.You will lead a team...  ...GPU scheduling, training, and inference systemsExecution & Operational... 
    Work at office
    Local area
    Remote work
    Relocation package
    3 days per week

    Nutanix

    San Jose, CA
    2 days ago
  • $220k - $300k

     ...Principal Machine Learning Engineer San Jose,...  ...and operations at scale. SambaNova Suite...  ...training and fine-tuning, inference optimization,...  ...bridges advanced LLM research and practical deployment, involving the development...  ...without direct management authority ~ Track... 
    Full time
    Temporary work
    Local area
    Flexible hours

    SambaNova Systems

    San Jose, CA
    17 hours ago
  • $220k - $300k

     ...delivering a full-stack inference platform for...  ...into SambaRack, rack-scale hardware that lets customers deploy state-of-the-art models...  ...Machine Learning Engineer, you will be...  ...role bridges advanced LLM research and practical...  ...without direct management authority ~ Track... 
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours

    SambaNova

    San Jose, CA
    17 hours ago
  • $195k - $285k

     ...compute. As a Principal Hardware Design Engineer, you will be a cornerstone of our...  ...solve the industry's most massive LLM inference challenges.Required Qualifications•...  ...Compute Project) standards or large-scale data center deployments.• Knowledge of liquid cooling integration... 
    3 days per week

    d-Matrix

    Santa Clara, CA
    3 days ago
  • $262.5k - $394k

     ...Principal Data Engineering Lead - Services Special...  ...at a larger scale, coordinating and...  ...lineage, metadata management, schema governance...  ...development, testing, deployment, and operations....  ...integrate ML model inference - including LLMs and...  ...TensorRT/TensorRT-LLM), and serving frameworks... 
    Relocation

    Apple

    Cupertino, CA
    17 hours ago
  • $272k - $431.25k

     ...exceptional Principal Perception Engineer to lead the design and...  ...using techniques such as large-scale pretraining, distillation, and...  ...robustness, and are ready for deployment at scale.Provide technical...  ...development and optimizing training or inference pipelines through custom CUDA... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $240k - $290k

     ...fast-growing company with the scale and impact of an established...  ...and proven by global deployment, we’re solving some of the most...  ...looking for an experienced Senior Engineering Manager to build and lead our AI...  ...well as our cloud-only AI and LLM-based products. You will... 
    Work at office
    Worldwide
    Flexible hours
    Shift work

    Sonatus

    Sunnyvale, CA
    1 day ago
  • $207k - $300k

     ...driving complex 1-6 month Engineers and 2-4 week Strike...  ...infrastructure (e.g., model deployment, model evaluation, data...  ...generative AI tools or LLM interfaces into...  ...distributed enterprise scale (Franchises) via structured...  .... Software Engineering Managers have not only the technical... 
    Permanent employment

    Google

    Mountain View, CA
    1 day ago
  • $249k

     ...everywhere.Director, Security Engineering Our Technology Team...  ...at enterprise scale. This is fundamentally...  ...and maintain secrets management, certificate lifecycle...  ...security architecture for LLM-based applications, RAG...  ...tuning practices, and deployment architecture through... 
    Full time
    Shift work
    Day shift

    Expedia

    San Jose, CA
    2 days ago
  •  ...Tuesday.About the RoleFleet Engineering owns the full lifecycle of Lambda...  ...— new product introduction, deployment, operation, and reliability...  ...provisioning, firmware management, out-of-band access, and power...  ...while driving efficiency at scale.We are hiring multiple... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    17 hours ago
  • $114.6k - $234.6k

     ...that operate reliably at cloud scale.Only Oracle brings together...  ...measurement, analytics, and inference systems for data center and...  ...analysis while maintaining engineering, security, and quality standards...  ..., training and evaluation, deployment, versioning, monitoring, and... 
    Temporary work
    Flexible hours

    Oracle Corporation

    Santa Clara, CA
    17 hours ago
  • $140k - $215k

     ...AI-native platform. We work on large scale distributed systems, processing almost...  ...CrowdStrike is seeking an experienced SDET Engineering Manager to lead our quality engineering...  ...and integrating automated testing into deployment workflowsExperience developing and deploying... 
    Full time
    Contract work
    Work experience placement
    Work at office
    Local area
    Worldwide
    2 days per week
    3 days per week

    CrowdStrike

    Sunnyvale, CA
    17 hours ago
  • $180k - $230k

     ...operating as one integrated managed service. The Company...  ...seeking a Principal Engineer to own the end-to-end...  ...internal tools (KNOC, Deployment, Service, Calibration,...  ...video analytics, AI inference, digital twin engines,...  ...experience architecting large-scale, real-time distributed... 
    Full time

    Knightscope

    Sunnyvale, CA
    17 hours ago
  • $141k - $221k

     ...Country: USA SummaryThe Senior Staff Engineer, Program/Project Management is accountable for planning and...  ...customer feedbackCoordinate site-wide deployment efforts.Implement change as directed...  ...development - from drawing board to full-scale production and after-market services... 
    Temporary work
    Work at office
    Local area
    Remote work
    Relocation

    Celestica

    San Jose, CA
    3 days ago
  • $169k - $338k

     ...including data sourcing, feature engineering, model training, deployment, and monitoring.Champion...  ...end-to-end ML lifecycle management, including feature...  ...Our team develops and scales AI-driven pricing solutions...  ...machine learning, causal inference, and optimization techniques... 
    Full time
    Temporary work
    Part time
    Work experience placement

    Walmart

    Sunnyvale, CA
    17 hours ago
  •  ...Machine Learning Engineer, you will drive the...  ...ExperimentationDesign, develop, and deploy production-grade...  ..., retrieval, LLM-based systems) to...  ...and online inference at scaleOversee...  ...closely with product managers, designers, and...  ...programs of work that scale across the... 
    Local area
    Remote work

    Atlassian

    Mountain View, CA
    17 hours ago
  • $206.4k - $379.1k

     ...Foundry is Adobe's enterprise managed-service offering for...  ...media-intelligence layer, and deployed across new and existing...  ...Principal Machine Learning Engineer to serve as the technical...  ...and served at enterprise scale. You will set the inference architecture and technical... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Jose, CA
    17 hours ago
  • $224k - $356.5k

     ...the power and flexibility to develop and deploy breakthrough artificial intelligence...  ...platforms — from data center to vehicle — and scale effortlessly from simple deployments to...  ...pipeline. We are seeking a seasoned Engineering Manager to lead our rapidly growing NvStreams... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $206.4k - $379.1k

     ...seeking a Principal Service Engineer to serve as the...  ...while productizing and scaling a rapidly growing portfolio...  ....Design and architect inference infrastructure for...  ...with Product Managers, TPMs, and engineering...  ...management in large-scale deployments.Hands-on expertise with... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Jose, CA
    4 days ago
  •  ...the future of transportation on a global scale. The Data Scaling team owns the Data...  ...collaborative, high-impact team of AI/ML engineers, data scientists and engineers who are...  ...pivotal in designing, architecting, and deploying data curation and training pipelines while... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Mountain View, CA
    17 hours ago
  •  ..., more than 35 live GPU cluster deployments globally, and backing from a top...  ...Silicon Valley VC, the company is scaling rapidly across its product and engineering organization. This...  ...partnering closely with Product Management, Architecture, and Customer Success... 
    Full time
    Santa Clara, CA
    a month ago
  • $60 - $70 per hour

    Engineering Program Manager W2 Contract Pay Rate: $60 - $70 per hour Location: Cupertino, CA - Hybrid...  ...Understanding of software development and deployment lifecycles. Familiarity with CI/CD,...  ..., Alibaba Cloud, or other large-scale cloud environments. Experience with... 
    Hourly pay
    Contract work

    Bayside Solutions

    Cupertino, CA
    2 days ago
  • $300k - $325k

     ...travel, primarily supporting engineering and manufacturing operations...  ...product development, manufacturing scale-up, and commercialization of...  ...costs, and accelerating deployment with Tier 1 customers. This...  ...guiding architecture decisions, managing a large multidisciplinary... 
    Relocation
    Relocation package

    The Hiring Method, LLC

    Sunnyvale, CA
    27 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Engineering Manager, LLM Inference & Deployment at Scale. Be the first to apply!