Senior Product Manager - AI Inference Performance
$208k - $327.75kNVIDIA Gruppe
What You'll Be Doing: Own the inference performance roadmap. Set direction across the stack: how models are represented, how memory and state are managed, how requests are scheduled and served, and how tokens get generated. The techniques change fast. Judge which ones matter, then decide what we build, what we adopt, and what we retire. Build platforms, not one-offs. Deliver capabilities that generalize across model families, deployment topologies, and customer sizes. Build for easy adoption, sane defaults, and extensibility. Agentic and Multi-Turn Workloads: Define the performance strategy for agentic applications, where long-running sessions, tool-call stalls, and unpredictable output lengths break the assumptions built into single-turn serving. Drive capabilities around cross-turn cache reuse, request prioritization, and efficient handling of idle time in agent loops. Framework & Ecosystem Strategy: Define how our optimizations land across TensorRT-LLM, vLLM, SGLang, and NVIDIA Dynamo. Partner with open-source communities and internal engineering teams so customers get great performance on NVIDIA hardware. Benchmarking & Performance Claims: Own how performance is measured, published, and reproduced. Define the benchmark methodology, the metrics that matter (TTFT, ITL, throughput per GPU, cost per million tokens), and the guardrails that keep our numbers credible. Run the product day to day. Own release readiness, quality bars, regression tracking, customer blocking issues, and the feedback loop from production deployments back into the roadmap. What We Need to See: 12+ years in product management at a technology company, or comparable time as a founder, engineering lead, or technical product owner. Depth in AI inference optimization: KV caching and reuse, quantization, speculative decoding, disaggregated serving. Know how each one moves accuracy, latency, and cost. Familiarity with the inference and orchestration frameworks customers use: TensorRT-LLM, vLLM, SGLang, NVIDIA Dynamo, and the surrounding serving and scheduling ecosystem. Proven track record of working independently — you can take an ambiguous problem space, define the strategy, and drive it to a shipped outcome without waiting to be told what to do next. Operational experience running a live product: release management, quality and regression rigor, customer issues, and support processes. Skill at translating low-level capability into business value — lower TCO, faster response, better GPU utilization — for engineers and executives alike. BS, MS, or PhD in Computer Science, Computer Engineering, or another relevant area of study (or equivalent experience). Ways to Stand Out From the Crowd: Engineering experience with LLM inference performance: profiling, kernel-level or serving-level optimization, or building a serving stack! Open-source contributions or product leadership in vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or Dynamo. Production experience at scale counts too: capacity planning, autoscaling, SLA management, or stateful multi-turn applications. A habit of reading the relevant research and translating it into roadmap decisions — you have intuition for where model architectures and serving techniques are heading next! Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 208,000 USD - 327,750 USD. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until August 17, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law. #J-18808-Ljbffr NVIDIA Gruppe
$168k - $258.75k
Inference is the fastest growing and most competitive area in Generative AI today. It is where AI models impact our daily life,... ...ever bit of accuracy and performance matters for quality, safety... ...deployment techniques. As a Senior Product Manager for AI Platform Inference...SeniorPerformanceFull time$208k - $327.75k
...the unlimited potential of AI to define the next era of computing... ...for a highly technical Product Manager to own the products that... ...customers extract the best possible performance from AI models and... ...running on NVIDIA hardware. Every inference deployment — from a single-...SeniorPerformanceFull time- NVIDIA is seeking a Senior Product Manager for AI Inference Performance (Finance) to own optimization strategies across the inference stack and drive platform capabilities for diverse deployments. You will translate deep optimization into broadly adoptable products, collaborate...SeniorPerformance
- NVIDIA in Santa Clara, CA, seeks a senior product leader to own the inference performance roadmap, shaping how models are represented, memory/state is managed, and tokens are generated. You will build scalable platforms across model families and deployment topologies for...SeniorPerformance
- ...mission is to build great products that accelerate next-... ...experiences—from AI and data centers, to... ...ROLE:AMD is seeking a Senior Product Manager to drive strategy and... ...on large-scale model inference on AMD Instinct™ and... ...source community, high-performance computing, and...SeniorPerformanceRemote work
- Cerebras Systems is seeking a Product Manager for AI Models to lead the strategic model portfolio that defines our product roadmap, deciding which models to ship and how they perform. You will partner with leading AI labs, drive launches, and ensure production-grade performance...SeniorPerformance
$134.47k - $220.92k
...Directs a comprehensive product strategy from product... ...OpportunityAs the Senior Product Manager for Artificial Intelligence... ...If you want to build AI that doesn't just... ...optimization, and inference speed.Guardrails & Safety... ...and model performance metrics into clear, business...SeniorPerformanceTemporary workWork at officeWork from homeWorldwideHome officeRelocation packageFlexible hours- ...builds the world's largest AI chip, 56 times larger... ...-leading training and inference speeds; over 10 times... ...AI inference. As the Product Manager for AI Models, you'll... ...models ship, how they perform, and how the world... ...or above the level of Senior PM.5+ years of total technical...SeniorPerformanceWork experience placementWork at officeRemote workShift work
$168k - $258.75k
...the unlimited potential of AI to define the next era of computing... ...for a highly technical Product Manager to own the strategy, roadmap... ..., multimodal AI, and edge inference into the vehicle cabin experience... ...optimize product performance based on real-world usage.What...SeniorPerformanceFull time$150k - $200k
...hooked. Other standout products include Chapters, where... ...ReelShort is seeking a senior product manager to define and launch a new AI companion and chat experience... ...per user, including inference cost Run competitive... ...eligible for an annual performance bonus, competitive...SeniorPerformanceWork experience placementLocal areaAfternoon shift- AMD in Santa Clara is seeking a Senior Product Manager to drive strategy and execution for ROCm, AMD’s open-source GPU software stack. This role involves managing the product roadmap for inference capabilities and engaging with the open-source community to enhance AMD'...SeniorRemote job
- ...Sr. Technical Product Manager, AI Data PlatformsThe AI infrastructure market is being redefined... ...governance and rollback features, high-performance storage protocols, enterprise data connectors... ...vector database partners, LLM and inference platform vendors, and enterprise data...SeniorPerformanceShift work
- Google is seeking a Senior Product Manager for Host Networking Software in Sunnyvale, CA. In this role you will... ...and align roadmaps with customer needs, with a focus on high-performance host networking infrastructure for AI training and inference. #J-18808-Ljbffr GoogleSeniorPerformance
$192k - $278k
...libraries for training and inference (e.g., NCCL, NIXL and other... ...to define requirements and performance of system level host networking... ...experience.8 years of experience in product management or a related technical role.... ...the next generation of AI training and inference. The...SeniorPerformanceWorldwide$179k - $254.3k
...are received.Splunk is looking for a Senior Product Manager to join our AI Foundations Team! We are building... ...pretraining, post-training, evaluation, and inference stack around them that power our... ...every model and every inference is performant, scalable, and built on a...SeniorPerformanceFull timeTemporary workLocal areaFlexible hours- Advanced Micro Devices is looking for a Senior Product Manager in Santa Clara, CA, to drive the strategy for ROCm, AMD's GPU software stack... ...strong background in product management, GPU computing, and AI inference technologies. This role provides an opportunity to work in...Senior
$148k - $235.75k
We are looking for a Senior Technical Product Marketing Manager. This role will be located in... ...business and pivotal in our inference marketing. You will be... ...our leadership position in AI inference.Want to join a... ...intelligence and high performance computing. Come grow your...SeniorPerformanceFull time- NVIDIA is seeking a Senior Product Manager for AI Platform Inference to build tools, SDKs, and libraries that enable developers to deploy inference on NVIDIA... ...strong knowledge of inference deployment and performance optimization, a BS/MS in CS/CE (or equivalent), and...Performance
$124.5k - $272k
...for the cloud and AI era. We secure and... ...and control without performance trade-offs.At... ...are available at Senior Staff and above. Candidates... ..., you own the inference and optimization... ...fast, efficient, and production-grade. You fine-... ...-cache and memory management, sparsity, fine-tuning...SeniorPerformance- CoreWeave is hiring a Senior Engineer for its Benchmarking & Performance team to write, profile, and optimize GPU kernels on the LLM inference path. You will improve latency and throughput and collaborate with product, orchestration, and hardware teams to achieve strict...SeniorPerformance
- ...world’s largest AI chip with wafer-scale... ...training and inference speeds and allows... ...with less hardware management. Cerebras’... ...iteration and higher productivity for AI applications... .... About The Role Senior Director of Technical... .../inference performance, scaling architectures...SeniorPerformance
$184k - $287.5k
...that wants to change how the AI inference industry works. With agent... ...support to push the boundaries of performance, squeezing value and... ...of our stack and bring back product insights that benefit the whole... ...best-practices for such as managing kubernetes clusters, configuring...SeniorPerformanceFull time- ...seeks a Principal System Software Engineer for AI Inference Execution. You will join the software team to productize the AI compute engine's SW stack, developing deployment... ...on Linux, and experience with distributed, high-performance software. #J-18808-Ljbffr Jobleads-USSeniorPerformance
- NVIDIA is seeking a Senior Agentic AI Software Engineer (Finance) to build agentic systems and advance AI inference workloads in real-world finance contexts. You will design and... ...drives experimental agents, optimize performance, and contribute to cutting-edge AI research...SeniorPerformance
$150k - $275k
A leading AI infrastructure company based in San Jose is seeking a highly skilled Supercomputing... .... This role involves developing high-performance networking solutions and optimizing software communication across inference nodes. Candidates should have strong C/C++ skills...SeniorPerformanceRelocation package- NVIDIA is seeking a highly technical Product Manager to own AI inference optimization on NVIDIA hardware, from single-GPU workstations to large data... ..., and cross-functional collaboration to deliver measurable performance gains and compelling #J-18808-Ljbffr NVIDIASeniorPerformance
- NVIDIA seeks a Product Manager for Inference to enable developers to deploy high-performance AI on NVIDIA GPUs. You will shape product strategy, roadmaps, and go-to-market plans, collaborating with internal teams and external developers to build model-optimization software...SeniorPerformance
- NVIDIA Corporation in Santa Clara, CA is seeking a Senior Product Manager for AI Platform Inference to lead the development of tools, SDKs, and libraries... ...collaborating with developers to optimize model deployment performance, with strong emphasis on GenAI concepts and GPU-...SeniorPerformance
- ...Airwallex, Fireblocks, and Bridge. Our AI Agents help compliance teams do more... ...shouldn't be. The Role As a Senior Product Manager on AiPrise's AI team , you'll drive... ...expertise in customer behavior, product performance, and data trends to shape strategy....SeniorPerformanceLive inHome officeFlexible hours
- Adobe is seeking a Senior Product Manager to lead AI/LLM work for adobe.com visitors and internal teams. You will define the vision, set the roadmap... ..., and business teams to ship features that balance performance and cost. You will build and evaluate LLM systems, including...SeniorPerformance
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Product Manager - AI Inference Performance. Be the first to apply!
- services product manager Santa Clara, CA
- product manager data analytics Santa Clara, CA
- iot product manager Santa Clara, CA
- sr technical product manager Santa Clara, CA
- product line manager Santa Clara, CA
- salesforce product manager Santa Clara, CA
- product offering manager Santa Clara, CA
- product strategy manager Santa Clara, CA
- associate product manager web Santa Clara, CA
- product manager lighting Santa Clara, CA


