Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Engineering Manager, LLM Inference & Deployment at Scale

$224k - $356.5k

NVIDIA

NVIDIA is looking for an Engineering Manager to lead the team responsible for deploying and serving Large Language Models (LLMs) and Vision-Language Models (VLMs) at scale.Our team builds and operates an AI inference platform that enables customers to deploy and run brand new generative AI models efficiently across NVIDIA GPU platforms. The platform operates at the intersection of model optimization, inference systems, distributed computing, and production infrastructure. In this role, you will manage a team of engineers responsible for taking modern models from research and experimentation into highly optimized, reliable, and scalable production deployments. You will work closely with research scientists, software engineers, and hardware specialists to push the boundaries of AI inference performance and deliver a world-class model deployment platform.What you will be doing:Lead, mentor, and grow a high-performing team building and operating a platform for deploying GenAI models at scale.Drive the deployment and optimization of LLMs and VLMs for low-latency, high-throughput, and cost-efficient inference.Analyze, profile, and optimize end-to-end deep learning workloads across the model, inference stack, distributed systems, and GPU hardware.Work closely with research teams and model developers to bring new architectures and models from prototype to production.Drive technical strategy and roadmap for model deployment, inference optimization, scalability, reliability, and performance.Establish engineering standards for benchmarking, profiling, production deployment, and continuous performance optimization.Collaborate with internal and external partners to enable seamless deployment of rapidly evolving GenAI models..What we want to see:Master’s or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.8+ years overall experience, including 3 years of management/leadership.Strong hands-on experience with LLMs and/or VLMs and a solid understanding of modern deep learning architectures.Deep expertise in inference optimization techniques such as quantization, speculative decoding, continuous batching, prefix caching, and KV-cache optimization.Experience with disaggregated inference/serving, distributed inference, multi-node deployments, and GPU cluster orchestration.Strong technical leadership, communication, and people-management skills.Ways to stand out from the crowd:Experience building or operating AI inference, model serving, or model deployment platforms.Practical experience in working with TensorRT, TensorRT-LLM, vLLM, SGLang, or comparable inference/serving frameworks.Experience building highly available, scalable, and observable production services for AI workloads.NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you!#LI-HybridYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD for Level 3, and 272,000 USD - 431,250 USD for Level 4.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until September 1, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Engineering Manager, LLM Inference & Deployment at Scale in Santa Clara, CA vacancy
  • $224k - $356.5k

     ...standard for assessing LLM serving performance across various inference frameworks. Hyperscalers...  ...improving efficiency, and scaling. As Technical Lead Manager, you will lead the engineering team within NVIDIA’s...  ..., and Kubernetes-native deployment.Taking ownership for the... 
    Suggested
    Full time
    Local area
    Remote work
    Worldwide

    Nvidia

    Santa Clara, CA
    4 days ago
  • $224k - $356.5k

     ...seeking an exceptional Manager, Deep Learning Inference Software, to lead a world-class engineering team advancing the state of AI model deployment. You will shape the...  ...time inference at every scale, from datacenter clusters...  ...-scale models for LLM, multimodal, and generative... 
    Suggested
    Full time
    Remote work
    Worldwide

    Nvidia

    Santa Clara, CA
    3 days ago
  • $224k - $356.5k

     ...accelerating it. We are accelerating LLM inference across the stack and across...  ...highly skilled and driven Engineering Manager to take the lead in...  ...experience for LLM deployment. Lead software development...  ...Proven ability to lead and scale high-performing engineering... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...ROLEWe are seeking a Principal GenAI Inference Optimization Engineer to join our Models and...  ...and cost efficiency for real-world deployment of large-scale models, working across the software...  ...throughput, and cost efficiency for LLM and multimodal model serving in production... 
    Suggested

    AMD

    San Jose, CA
    1 day ago
  • $201.3k - $352.3k

     ...all started when engineer Fred Luddy wrote code...  ..., and product managers with a dual mission...  ...Design, own, and scale automated testing...  ...window efficiency, and inference costs.Lead a High-...  ...deep experience deploying or testing agentic...  ...ROUGE, BLEU, G-Eval, LLM-as-a-judge... 
    Suggested
    Work experience placement
    Work at office
    Immediate start
    Remote work
    Flexible hours
    Shift work

    ServiceNow

    Santa Clara, CA
    3 days ago
  • A leading technology company in California is seeking an Engineering Manager to lead the development of cutting-edge LLM/VLM technologies. In this hands-on leadership role, you will manage a team responsible for optimizing runtime and frameworks, while collaborating with... 

    NVIDIA Corporation

    Santa Clara, CA
    4 days ago
  • $201k - $402k

     ...OpportunityWe are looking for a Senior Engineering Manager to lead the design, development, and scaling of a next-generation...  ...capabilities in on-prem and hybrid deployments.You will lead a team...  ...GPU scheduling, training, and inference systemsExecution & Operational... 
    Work at office
    Local area
    Remote work
    Relocation package
    3 days per week

    Nutanix

    San Jose, CA
    10 hours ago
  • $154k - $312k

     ...platform, its Zero Trust Engine, and the powerful...  ...research and enterprise-scale business impact. Leveraging...  ...the architecture and deployment of large-scale, production...  ...Production-Grade Inference Systems: Design, optimize...  ...leveraging the latest LLM serving technologies such... 

    Netskope

    Santa Clara, CA
    2 days ago
  •  ...deliver industry-leading training and inference speeds; over 10 times faster than...  ...-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra...  ...The RoleWe're hiring a Principal Engineer for our Inference Cloud Platform. This... 

    Cerebras Systems

    Sunnyvale, CA
    2 days ago
  • $224k - $356.5k

     ...systems.We are looking for an Engineering Manager to build and lead the team...  ...spanning agent runtimes, local inference, Windows integration,...  ...software, and local or hybrid deployment, with the judgment to guide...  ...patterns.Experience with TensorRT-LLM, Ollama, llama.cpp, vLLM,... 
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    2 days ago
  • $171.6k - $245k

     ...The team works across product management, engineering, applied AI/ML, data,...  ...customer-ready capabilities at scale. This is a highly multi-functional...  ...continuous integration and deployment.Communicate program status,...  ...at least one generative AI, LLM-based, or agentic capability... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    10 hours ago
  • $270k - $300k

     ...Senior Engineering Manager, MLOpsPalo Alto, California, United...  ...company that is rapidly scaling while maintaining a...  ...to automated deployment and real-time monitoring...  ...infrastructure cost per inference.Champion Operational...  ...Monitor emerging trends in LLM-ops, vector databases... 
    Seasonal work
    Local area

    Quince

    Palo Alto, CA
    3 days ago
  • $240.1k - $420.2k

     ...DescriptionIt all started when engineer Fred Luddy wrote code that...  ...agentic AI and enterprise-scale search systems that power...  ...autonomous systems safe to deploy at scale. Retrieval and grounding...  ...or retrieval. Exposure to LLM fine-tuning or inference optimization in... 
    Work experience placement
    Work at office
    Immediate start
    Remote work
    Flexible hours

    ServiceNow

    Santa Clara, CA
    1 day ago
  • $272k - $431.25k

     ...Machine Learning (ML) Engineer to join the GPU accelerated...  ...for running large scale workloads for ETL, SQL,...  .../DL model training and inference pipelines, spanning many...  ...partners and customers on the deployment of complex machine...  ..., including LLM/GenAI, reinforcement learning... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $272k - $425.5k

    A leading technology firm in Santa Clara seeks a Principal Software Engineer for Large-Scale LLM Memory and Storage Systems. The role involves designing a unified memory layer for large-scale inference, integrating with LLM serving engines, and mentoring engineers. Candidates... 
    Remote job

    NVIDIA Corporation

    Santa Clara, CA
    1 day ago
  • $195k - $285k

     ...compute. As a Principal Hardware Design Engineer, you will be a cornerstone of our...  ...solve the industry's most massive LLM inference challenges.Required Qualifications•...  ...Compute Project) standards or large-scale data center deployments.• Knowledge of liquid cooling integration... 
    3 days per week

    d-Matrix

    Santa Clara, CA
    1 day ago
  • $262.5k - $394k

     ...Principal Data Engineering Lead - Services Special...  ...at a larger scale, coordinating and...  ...lineage, metadata management, schema governance...  ...development, testing, deployment, and operations.Align...  ...ML model inference - including LLMs and...  ...TensorRT/TensorRT-LLM), and serving frameworks... 
    Relocation

    Apple

    Cupertino, CA
    2 days ago
  •  ...Tuesday.About the RoleFleet Engineering owns the full lifecycle of Lambda...  ...— new product introduction, deployment, operation, and reliability...  ...provisioning, firmware management, out-of-band access, and power...  ...while driving efficiency at scale.We are hiring multiple... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  • $185.1k - $284.1k

     ...everyday drivers at unprecedented scale. Join us to help deliver the next generation...  ...results into clear feedback for engineering and leadership, and help accelerate validated AV deployment at scale. We are looking for an Engineering Manager to lead a team building the... 
    Full time
    Local area
    Remote work
    Work from home
    Flexible hours

    General Motors

    Sunnyvale, CA
    1 day ago
  • $240k - $290k

     ...fast-growing company with the scale and impact of an established...  ...and proven by global deployment, we’re solving some of the most...  ...looking for an experienced Senior Engineering Manager to build and lead our AI...  ...well as our cloud-only AI and LLM-based products. You will... 
    Work at office
    Worldwide
    Flexible hours
    Shift work

    Sonatus

    Sunnyvale, CA
    4 days ago
  • $272k - $431.25k

     ...exceptional Principal Perception Engineer to lead the design and...  ...using techniques such as large-scale pretraining, distillation, and...  ...robustness, and are ready for deployment at scale.Provide technical...  ...development and optimizing training or inference pipelines through custom CUDA... 
    Full time

    Nvidia

    Santa Clara, CA
    10 hours ago
  • $225k - $325k

     ...impact on a global scale. At Goodwin, we...  ...peers. Director of Engineering The Director of...  ..., and production deployment of AI-driven...  ...system latency, model inference performance, and infrastructure...  ..., product management, and business...  ...system architectures, LLM orchestration... 

    Goodwin Procter

    Palo Alto, CA
    5 days ago
  • $141k - $221k

     ...Country: USA SummaryThe Senior Staff Engineer, Program/Project Management is accountable for planning and...  ...customer feedbackCoordinate site-wide deployment efforts.Implement change as directed...  ...development - from drawing board to full-scale production and after-market services... 
    Temporary work
    Work at office
    Local area
    Remote work
    Relocation

    Celestica

    San Jose, CA
    7 hours ago
  •  ...THIS FEATURED OPPORTUNITY The Engineering Program Manager will join the Channel Strategy &...  ...teams and ensure on-time delivery and deployment of key initiatives Define milestones...  ...managing projects involving LLM-based products or AI-driven systems... 
    Local area
    Flexible hours

    INSPYR Solutions

    Cupertino, CA
    1 day ago
  • $250k - $300k

     .... We're hiring a strong engineering leader to own the Cloud &...  ...Tech Lead ready to step into management OR a Manager who still codes...  ...for cloud infrastructure, deployments, and platform capabilities...  ...environments and production systems at scale Drive improvements in... 
    Work at office
    Shift work

    Orkes

    Cupertino, CA
    1 day ago
  • $140k - $215k

     ...AI-native platform. We work on large scale distributed systems, processing almost...  ...CrowdStrike is seeking an experienced SDET Engineering Manager to lead our quality engineering...  ...and integrating automated testing into deployment workflowsExperience developing and deploying... 
    Full time
    Contract work
    Work experience placement
    Work at office
    Local area
    Worldwide
    2 days per week
    3 days per week

    CrowdStrike

    Sunnyvale, CA
    2 days ago
  • $160k - $240k

     ...Senior Engineering Manager (Backend & Platform)Fiserv is a global leader in payments and financial...  ...processing financial transactions at scale.Establish operational excellence,...  ...guarantee platform reliability, zero-downtime deployments, and fast incident recovery.Direct... 
    Contract work
    Work at office
    Worldwide
    Monday to Friday

    Fiserv

    Sunnyvale, CA
    3 days ago
  • $237.6k - $356.4k

     ...Engineering Manager, FileMaker ServerAt Apple, great ideas have a way of becoming great products...  ...platform technologies that power enterprise-scale solutions and directly impact customer...  ...(APIs), and cloud or on-premises deployments.Review designs and code to ensure... 
    Relocation

    Apple

    Sunnyvale, CA
    3 days ago
  • $204k - $343k

     ...flexibility and trust our employees to manage their schedules responsibly....  .... About the role As an Engineering Manager on the Data...  ...cutting‑edge AI techniques to scale these capabilities, setting technical...  ...development, safety, and deployment milestones. At Applied Intuition... 
    Full time
    For contractors
    For subcontractor
    Casual work
    Work at office
    Remote work
    Day shift

    Applied Intuition

    Sunnyvale, CA
    2 days ago
  •  ...CrowdStrike SDET Engineering ManagerAs a global leader in cybersecurity...  ...platform. We work on large scale distributed systems, processing...  ...experienced SDET Engineering Manager to lead our quality...  ...integrating automated testing into deployment workflowsExperience developing... 
    Contract work
    Work at office
    Worldwide
    2 days per week
    3 days per week

    CrowdStrike

    Sunnyvale, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Engineering Manager, LLM Inference & Deployment at Scale. Be the first to apply!