Manager, Large Language Model Inference
$184k - $287.5kNVIDIA
At NVIDIA, we aren't just powering the AI revolution-we're accelerating it. The TensorRT inference platform is the backbone of modern AI, delivering the industry's fastest and most efficient deployment of cutting-edge deep learning models on every NVIDIA GPU. With demand for AI exploding, particularly in the realm of large language models (LLMs) and vision language models (VLMs, VLAs), we are significantly expanding our team. We're seeking a highly skilled and driven Engineering Manager to take the lead in developing the next generation of LLM/VLM/VLA inference software technologies that will define the future of AI. This is a high-impact, hands-on leadership role at the intersection of deep technical expertise and world-class management. You won't just manage; you'll architect and guide a brilliant team of engineers who are building the core LLM inference runtime. Your work will be highly collaborative, interfacing directly with NVIDIA Researchers, GPU Architects, and other teams across the company to ensure we ship production-grade, lightning-fast software that sets the global standard for AI performance. What You'll Be Doing: Lead and grow a team responsible for specialized kernel development, runtime optimizations, and frameworks for LLM inference. Drive the design, development, and delivery of production inference software, targeting NVIDIA's next-generation enterprise and edge hardware platforms. Integrating cutting-edge technologies developed at NVIDIA and offering an intuitive developer experience for LLM deployment. Lead software development execution, with responsibility for project planning, milestone delivery, and cross-functional coordination. What We Need to See: MS, PhD, or equivalent experience in Computer Science, Computer Engineering, AI, or a related technical field. 7+ overall years of overall software engineering experience, including 3+ years of technical leadership experience. Proven ability to lead and scale high-performing engineering teams, especially across distributed and cross-functional groups. Strong background in C++ or Python, with expertise in software design and delivering production-quality software libraries. Demonstrated expertise in large language models (LLM) and/or vision language models (VLM). Ways to Stand Out from the Crowd: Deep understanding of GPU architecture, CUDA programming, and system-level performance tuning. Background in LLM inference or working with frameworks such as TensorRT-LLM, vLLM, or SGLang. Passion for building scalable, user-friendly APIs and enabling developers in the AI ecosystem. Have a proven track record of growing and managing a team that encourages idea sharing, empowers team members, and provides opportunities for professional growth. We are widely considered to be one of the technology world's most desirable employers, and we have some of the most forward-thinking and hardworking people in the world working with us. Due to outstanding growth, our best-in-class teams are rapidly growing. If you're a creative self-starter with a real passion for technology, then come join us. #LI-Hybrid Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 2, and 224,000 USD - 356,500 USD for Level 3. You will also be eligible for equity and benefits . Applications for this job will be accepted at least until November 4, 2025. NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law. #J-18808-Ljbffr NVIDIA Corporation
$184k - $287.5k
...forefront of the generative AI revolution! The Algorithmic Model Optimization Team specifically focuses on optimizing generative AI models such as large language models (LLM) and diffusion models for maximal inference efficiency using techniques ranging from neural...LanguageFull time$193.3k - $261.5k
...to optimize the latest models to run really fast on... ...Development Engineer on the Inference Model Enablement team,... ..., and product managers to architect and deliver... ...a modern programming language such as Java, C++, or... ...Machine Learning and Large Language Model fundamentals...LanguageInternshipLocal areaFlexible hours$165.2k - $223.6k
...enabling unparalleled ML inference and training... ...running a wide range of models and supporting novel architecture... ...spaces that are very large, yet our teams remain... ...massive scale large language models like the Llama... ..., and product managers to deliver state-of-the...LanguageWork experience placementInternshipLocal areaFlexible hours- ...to deliver industry-leading training and inference speeds; over 10 times faster than GPU-... ...computation. Cerebras works with the leading model labs, global enterprises, and cutting-... ..., hardware architecture, product management, and AI research teams while helping shape...Suggested
- ...industry-leading training and inference speeds and allows users to run large-scale ML applications with less hardware management. Cerebras’ customers include top model labs, global enterprises, and... ...environments. Familiarity with large language models, foundation model...Language
$192.2k - $260k
...'s Delivery Foundation Model team, where you'll work... ...amounts of Amazon data and infer at Amazon scale, taking... ...Python, C++ or other languages- Strong publication... ...deployments- Experience with large-scale distributed... ...judgment, effectively manage stress and work safely...LanguageLocal areaWorldwideFlexible hours- ...Institute of Foundation Models We are a dedicated... ...understanding, using, and risk-managing foundation models. Our... ...in the Vision Language Model (VLM) team, your... ...research and development of large‑scale VLM systems,... ...model modularity, and inference optimization. Build...Language
$160.5k - $240.7k
...developers optimize and deploy machine learning models on edge and mobile hardware. AIMET is... .... Applications range from quantizing large language models (LLMs) and generative AI models... ..., Falcon, or similar families) for inference optimization Familiarity with AIMET, GPTQ...LanguageWork experience placementImmediate startWork from home$272k - $431.25k
...for a new generation of interactive world-model systems. With this release, NVIDIA has set the standard for fidelity, real-time inference performance in world models. And, we've... ...experience, CI/CD, and release management.Partner with research, simulation, rendering...Full time$197.3k - $225.1k
...Lead AI Engineer (Vision model customization, VLM) Overview At Capital One... ...research scientists, technical program managers, and product managers to deliver AI-... ...including foundation model training, large language model inference, similarity search, guardrails, model...LanguageFull timePart timeLocal area$163k - $236k
...product strategy for improving our models and agentic capabilities.... ...of experience in product management or a related technical role.Experience... ...AI agents for medium-to-large businesses and enterprises.... ...building evaluations for Large Language Models (LLMs) and agents....Language$174.72k - $295.68k
...Machine Learning Engineers with strong expertise in generative modeling and large-scale deep learning systems, along with solid software... ...actions, and apply predictive pre-training to improve Vision-Language-Action (VLA) driving performance. Extend prediction beyond...LanguageFull time$174.72k - $295.68k
...Learning Engineer / Research Scientist to drive the modeling and algorithmic development of XPENG’s next-generation Vision-Language-Action (VLA) Foundation Model — the core brain... ...experts to design, train, and deploy large-scale multi-modal models that unify vision, language...LanguageFull time$117.7k - $221.4k
...curate high-value data from large-scale real-world sensor streams... ...depends not only on stronger models, but also on better... ...processing, featurization, and inference foundations that power scalable... ...other frequency dictated by your manager}. This job may be eligible for...Full timeLocal areaRemote workWork from homeRelocation packageFlexible hours$211.4k - $286k
...foundational artificial intelligence and large language models to stay at the forefront of the... ...seeking an experienced Applied Science Manager to lead and grow a team of applied scientists... ..., scikit-learn, numpy, scipy or edge inference frameworksAmazon is an equal...LanguageLocal areaWorldwideFlexible hours$224k - $356.5k
...building cutting-edge infrastructure for large-scale foundation model training in the Generalist Embodied... ..., CUDA programming, and cluster management tools like Kubernetes.Strong programming... ...in Python and a high-performance language such as C++ for efficient system...LanguageFull time$184k - $287.5k
...-powered application is built. We are seeking a senior vision language model engineer to design and build agentic data and training workflows... ...architecture skills demonstrated through contributions to large internal or open-source projects.Experience in robotic systems...LanguageFull time$156k - $255.6k
...team is looking for Business Development Managers based in Sunnyvale. Key Responsibilities... ...of orchestrating resources within a large-scale global organization like Alibaba Group... ...’s Degree in STEM is highly preferred Languages: Native or professional fluency in English...LanguageLocal area$175k - $350k
...pioneering this future with human-centered AI models that unite emotional intelligence (EQ)... ...and perspectives. Platform — large-language models (LLMs) and APIs that enable builders... ...quality targets. Collaborate with inference, safety, and product teams to land improvements...Language- ...founding member. MEGA is building foundation models for general-purpose robots beyond not... ...foundation models, embodied intelligence and large-scale machine learning. You will have... ...capabilities. Your day-to-day work may span vision-language-action models, world and action models,...LanguageFull timeWork from home
$126k - $423k
...flexibility and trust our employees to manage their schedules responsibly.... ...of miles of data from large fleets, and deploy methods they... ...pretraining world‑action foundation model with various world modalities... ..., human data incorporation, language modality, and spatial...LanguageFull timeFor contractorsFor subcontractorCasual workWork at officeImmediate startRemote workDay shift$75.5k - $104k
...seeking a highly motivated Technical Program Manager to join the Knowledge Management (KM)... ...programming to cleanse, integrate, and analyze large datasets Develop automated data‑... ...or native speaker in one or more of these languages: Japanese, Korean, Simplified Chinese, Traditional...LanguageFull timeRelocation$184k - $287.5k
...looking for a technical product marketing manager who is passionate about AI frameworks... ...that convey the value of training and inference frameworks, such as PyTorch, JAX,... ...expertise - Familiarity with popular large language models like DeepSeek, GPT-OSS, Gemma and Phi...LanguageFull timeWork experience placement$150k
...You will join the Grok Voice Model team to help build the world'... ...processing, frontier speech-language pre-training, and intensive post... ...: Design and execute large-scale speech data curation and... ...scale distributed training and inference systems on Kubernetes. Proactive...LanguageTemporary work$212.7k - $287.7k
...run really fast on the Trainium hardware.As an SDM for the LLM Inference Model Enablement team, you will lead a team of expert AI/ML... ...using distributed inference libraries. You should be capable of managing demanding, fast-changing priorities. You should have a strong...Local areaFlexible hours$152k - $241.5k
...innovators to roll out and enhance AI inference solutions at scale, demonstrating NVIDIA... ...Inference Server, or TensorRT-LLM for model optimization and serving.GPU orchestration... ..., UCX).Demonstrated success in tuning large language models for low-latency inference in...LanguageFull time$195.2k - $262.2k
...developers and enterprises from data and model training through to production deployment... ...without the cost and complexity of building large in-house AI/ML infrastructure. Built by... .... From large-scale GPU orchestration to inference optimization, we own the hard problems...Temporary workImmediate startRemote work$152k - $241.5k
...-driven Developer Relations Manager focused on Foundational AI Research... ...the next generation of AI models, systems, and methods. In... ...AI systems, including large language models, multimodal models, reasoning... ...systems, training methods, inference systems, model serving, and...LanguageFull time$174k - $252k
...software development in one or more programming languages (e.g., Python, C, C++, Java, JavaScript).... ...of experience building and developing large-scale infrastructure or distributed... ...as GOLang, Rust, or Java.Experience in ML model coding languages (e.g., Python).Google's...Language- ...Systems is seeking an engineering leader to build and scale the Inference Model Scaling organization. You will define the technical vision... .... Work across compiler, runtime, cloud, hardware, product management, and AI research to shape the future of AI inference at Cerebras...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Manager, Large Language Model Inference. Be the first to apply!
- phlebotomy manager Santa Clara, CA
- pharma manager Santa Clara, CA
- apparel manager Santa Clara, CA
- sox manager Santa Clara, CA
- certification manager Santa Clara, CA
- full time manager Santa Clara, CA
- help desk manager Santa Clara, CA
- transaction manager Santa Clara, CA
- onsite manager Santa Clara, CA
- manager sodexo Santa Clara, CA


