Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Product Manager - AI Inference Performance

$208k - $327.75k

NVIDIA Gruppe

What You'll Be Doing: Own the inference performance roadmap. Set direction across the stack: how models are represented, how memory and state are managed, how requests are scheduled and served, and how tokens get generated. The techniques change fast. Judge which ones matter, then decide what we build, what we adopt, and what we retire. Build platforms, not one-offs. Deliver capabilities that generalize across model families, deployment topologies, and customer sizes. Build for easy adoption, sane defaults, and extensibility. Agentic and Multi-Turn Workloads: Define the performance strategy for agentic applications, where long-running sessions, tool-call stalls, and unpredictable output lengths break the assumptions built into single-turn serving. Drive capabilities around cross-turn cache reuse, request prioritization, and efficient handling of idle time in agent loops. Framework & Ecosystem Strategy: Define how our optimizations land across TensorRT-LLM, vLLM, SGLang, and NVIDIA Dynamo. Partner with open-source communities and internal engineering teams so customers get great performance on NVIDIA hardware. Benchmarking & Performance Claims: Own how performance is measured, published, and reproduced. Define the benchmark methodology, the metrics that matter (TTFT, ITL, throughput per GPU, cost per million tokens), and the guardrails that keep our numbers credible. Run the product day to day. Own release readiness, quality bars, regression tracking, customer blocking issues, and the feedback loop from production deployments back into the roadmap. What We Need to See: 12+ years in product management at a technology company, or comparable time as a founder, engineering lead, or technical product owner. Depth in AI inference optimization: KV caching and reuse, quantization, speculative decoding, disaggregated serving. Know how each one moves accuracy, latency, and cost. Familiarity with the inference and orchestration frameworks customers use: TensorRT-LLM, vLLM, SGLang, NVIDIA Dynamo, and the surrounding serving and scheduling ecosystem. Proven track record of working independently — you can take an ambiguous problem space, define the strategy, and drive it to a shipped outcome without waiting to be told what to do next. Operational experience running a live product: release management, quality and regression rigor, customer issues, and support processes. Skill at translating low-level capability into business value — lower TCO, faster response, better GPU utilization — for engineers and executives alike. BS, MS, or PhD in Computer Science, Computer Engineering, or another relevant area of study (or equivalent experience). Ways to Stand Out From the Crowd: Engineering experience with LLM inference performance: profiling, kernel-level or serving-level optimization, or building a serving stack! Open-source contributions or product leadership in vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or Dynamo. Production experience at scale counts too: capacity planning, autoscaling, SLA management, or stateful multi-turn applications. A habit of reading the relevant research and translating it into roadmap decisions — you have intuition for where model architectures and serving techniques are heading next! Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 208,000 USD - 327,750 USD. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until August 17, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law. #J-18808-Ljbffr NVIDIA Gruppe

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Senior Product Manager - AI Inference Performance in Santa Clara, CA vacancy
  • $168k - $258.75k

    Inference is the fastest growing and most competitive area in Generative AI today. It is where AI models impact our daily life,...  ...ever bit of accuracy and performance matters for quality, safety...  ...deployment techniques. As a Senior Product Manager for AI Platform Inference... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    15 hours ago
  • $208k - $327.75k

     ...the unlimited potential of AI to define the next era of computing...  ...for a highly technical Product Manager to own the products that...  ...customers extract the best possible performance from AI models and...  ...running on NVIDIA hardware. Every inference deployment — from a single-... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • NVIDIA is seeking a Senior Product Manager for AI Inference Performance (Finance) to own optimization strategies across the inference stack and drive platform capabilities for diverse deployments. You will translate deep optimization into broadly adoptable products, collaborate... 
    Senior
    Performance

    Nvidia Corporation in

    Santa Clara, CA
    4 days ago
  • NVIDIA in Santa Clara, CA, seeks a senior product leader to own the inference performance roadmap, shaping how models are represented, memory/state is managed, and tokens are generated. You will build scalable platforms across model families and deployment topologies for... 
    Senior
    Performance

    NVIDIA Gruppe

    Santa Clara, CA
    4 days ago
  •  ...mission is to build great products that accelerate next-...  ...experiences—from AI and data centers, to...  ...ROLE:AMD is seeking a Senior Product Manager to drive strategy and...  ...on large-scale model inference on AMD Instinct™ and...  ...source community, high-performance computing, and... 
    Senior
    Performance
    Remote work

    AMD

    Santa Clara, CA
    15 hours ago
  • Cerebras Systems is seeking a Product Manager for AI Models to lead the strategic model portfolio that defines our product roadmap, deciding which models to ship and how they perform. You will partner with leading AI labs, drive launches, and ensure production-grade performance... 
    Senior
    Performance

    Cerebras

    Sunnyvale, CA
    3 days ago
  • $134.47k - $220.92k

     ...Directs a comprehensive product strategy from product...  ...OpportunityAs the Senior Product Manager for Artificial Intelligence...  ...If you want to build AI that doesn't just...  ...optimization, and inference speed.Guardrails & Safety...  ...and model performance metrics into clear, business... 
    Senior
    Performance
    Temporary work
    Work at office
    Work from home
    Worldwide
    Home office
    Relocation package
    Flexible hours

    McAfee

    San Jose, CA
    2 days ago
  •  ...builds the world's largest AI chip, 56 times larger...  ...-leading training and inference speeds; over 10 times...  ...AI inference. As the Product Manager for AI Models, you'll...  ...models ship, how they perform, and how the world...  ...or above the level of Senior PM.5+ years of total technical... 
    Senior
    Performance
    Work experience placement
    Work at office
    Remote work
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    2 days ago
  • $168k - $258.75k

     ...the unlimited potential of AI to define the next era of computing...  ...for a highly technical Product Manager to own the strategy, roadmap...  ..., multimodal AI, and edge inference into the vehicle cabin experience...  ...optimize product performance based on real-world usage.What... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $150k - $200k

     ...hooked. Other standout products include Chapters, where...  ...ReelShort is seeking a senior product manager to define and launch a new AI companion and chat experience...  ...per user, including inference cost Run competitive...  ...eligible for an annual performance bonus, competitive... 
    Senior
    Performance
    Work experience placement
    Local area
    Afternoon shift

    Crazy Maple Studio

    Sunnyvale, CA
    5 days ago
  • AMD in Santa Clara is seeking a Senior Product Manager to drive strategy and execution for ROCm, AMD’s open-source GPU software stack. This role involves managing the product roadmap for inference capabilities and engaging with the open-source community to enhance AMD'... 
    Senior
    Remote job

    AMD

    Santa Clara, CA
    2 days ago
  •  ...Sr. Technical Product Manager, AI Data PlatformsThe AI infrastructure market is being redefined...  ...governance and rollback features, high-performance storage protocols, enterprise data connectors...  ...vector database partners, LLM and inference platform vendors, and enterprise data... 
    Senior
    Performance
    Shift work

    DDN Storage

    Santa Clara, CA
    3 days ago
  • Google is seeking a Senior Product Manager for Host Networking Software in Sunnyvale, CA. In this role you will...  ...and align roadmaps with customer needs, with a focus on high-performance host networking infrastructure for AI training and inference. #J-18808-Ljbffr Google
    Senior
    Performance

    Google

    Sunnyvale, CA
    5 days ago
  • $192k - $278k

     ...libraries for training and inference (e.g., NCCL, NIXL and other...  ...to define requirements and performance of system level host networking...  ...experience.8 years of experience in product management or a related technical role....  ...the next generation of AI training and inference. The... 
    Senior
    Performance
    Worldwide

    Google

    Sunnyvale, CA
    3 days ago
  • $179k - $254.3k

     ...are received.Splunk is looking for a Senior Product Manager to join our AI Foundations Team! We are building...  ...pretraining, post-training, evaluation, and inference stack around them that power our...  ...every model and every inference is performant, scalable, and built on a... 
    Senior
    Performance
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    2 days ago
  • Advanced Micro Devices is looking for a Senior Product Manager in Santa Clara, CA, to drive the strategy for ROCm, AMD's GPU software stack...  ...strong background in product management, GPU computing, and AI inference technologies. This role provides an opportunity to work in... 
    Senior

    Advanced Micro Devices

    Santa Clara, CA
    3 days ago
  • $148k - $235.75k

    We are looking for a Senior Technical Product Marketing Manager. This role will be located in...  ...business and pivotal in our inference marketing. You will be...  ...our leadership position in AI inference.Want to join a...  ...intelligence and high performance computing. Come grow your... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • NVIDIA is seeking a Senior Product Manager for AI Platform Inference to build tools, SDKs, and libraries that enable developers to deploy inference on NVIDIA...  ...strong knowledge of inference deployment and performance optimization, a BS/MS in CS/CE (or equivalent), and... 
    Performance

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $124.5k - $272k

     ...for the cloud and AI era. We secure and...  ...and control without performance trade-offs.At...  ...are available at Senior Staff and above. Candidates...  ..., you own the inference and optimization...  ...fast, efficient, and production-grade. You fine-...  ...-cache and memory management, sparsity, fine-tuning... 
    Senior
    Performance

    Netskope

    Santa Clara, CA
    3 days ago
  • CoreWeave is hiring a Senior Engineer for its Benchmarking & Performance team to write, profile, and optimize GPU kernels on the LLM inference path. You will improve latency and throughput and collaborate with product, orchestration, and hardware teams to achieve strict... 
    Senior
    Performance

    CoreWeave

    Sunnyvale, CA
    1 day ago
  •  ...world’s largest AI chip with wafer-scale...  ...training and inference speeds and allows...  ...with less hardware management. Cerebras’...  ...iteration and higher productivity for AI applications...  .... About The Role Senior Director of Technical...  .../inference performance, scaling architectures... 
    Senior
    Performance

    Cerebras

    Sunnyvale, CA
    3 days ago
  • $184k - $287.5k

     ...that wants to change how the AI inference industry works. With agent...  ...support to push the boundaries of performance, squeezing value and...  ...of our stack and bring back product insights that benefit the whole...  ...best-practices for such as managing kubernetes clusters, configuring... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...seeks a Principal System Software Engineer for AI Inference Execution. You will join the software team to productize the AI compute engine's SW stack, developing deployment...  ...on Linux, and experience with distributed, high-performance software. #J-18808-Ljbffr Jobleads-US
    Senior
    Performance

    Jobleads-US

    Santa Clara, CA
    4 days ago
  • NVIDIA is seeking a Senior Agentic AI Software Engineer (Finance) to build agentic systems and advance AI inference workloads in real-world finance contexts. You will design and...  ...drives experimental agents, optimize performance, and contribute to cutting-edge AI research... 
    Senior
    Performance

    Nvidia Corporation in

    Santa Clara, CA
    3 days ago
  • $150k - $275k

    A leading AI infrastructure company based in San Jose is seeking a highly skilled Supercomputing...  .... This role involves developing high-performance networking solutions and optimizing software communication across inference nodes. Candidates should have strong C/C++ skills... 
    Senior
    Performance
    Relocation package

    Jobleads-US

    San Jose, CA
    4 days ago
  • NVIDIA is seeking a highly technical Product Manager to own AI inference optimization on NVIDIA hardware, from single-GPU workstations to large data...  ..., and cross-functional collaboration to deliver measurable performance gains and compelling #J-18808-Ljbffr NVIDIA
    Senior
    Performance

    NVIDIA

    Santa Clara, CA
    4 days ago
  • NVIDIA seeks a Product Manager for Inference to enable developers to deploy high-performance AI on NVIDIA GPUs. You will shape product strategy, roadmaps, and go-to-market plans, collaborating with internal teams and external developers to build model-optimization software... 
    Senior
    Performance

    NVIDIA Gruppe

    Santa Clara, CA
    1 day ago
  • NVIDIA Corporation in Santa Clara, CA is seeking a Senior Product Manager for AI Platform Inference to lead the development of tools, SDKs, and libraries...  ...collaborating with developers to optimize model deployment performance, with strong emphasis on GenAI concepts and GPU-... 
    Senior
    Performance

    NVIDIA Corporation

    Santa Clara, CA
    4 days ago
  •  ...Airwallex, Fireblocks, and Bridge. Our AI Agents help compliance teams do more...  ...shouldn't be. The Role As a Senior Product Manager on AiPrise's AI team , you'll drive...  ...expertise in customer behavior, product performance, and data trends to shape strategy.... 
    Senior
    Performance
    Live in
    Home office
    Flexible hours

    AiPrise

    San Jose, CA
    3 days ago
  • Adobe is seeking a Senior Product Manager to lead AI/LLM work for adobe.com visitors and internal teams. You will define the vision, set the roadmap...  ..., and business teams to ship features that balance performance and cost. You will build and evaluate LLM systems, including... 
    Senior
    Performance

    Adobe

    San Jose, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Product Manager - AI Inference Performance. Be the first to apply!