Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Director, Product Management - AI Inference Platform (San Jose)

Full-time

Arm

Location: San Jose, California

We are seeking a Director / Principal Product Manager to lead the strategy and execution of a next-generation AI Inference Platform. This role sits at the intersection of hardware and software, defining how modern AI models are executed, served, and optimized at scale across distributed compute environments.

You will own the core platform stack—driving innovation in performance, efficiency, and scalability while partnering closely with engineering, research, and infrastructure teams to deliver world-class AI systems.

Responsibilities

  • Define and drive the product vision, strategy, and roadmap for large-scale AI inference systems
  • Own product direction across compute execution, inference serving, and control plane systems
  • Drive inference architecture and orchestration strategy, including distributed serving, routing, batching, and scheduling
  • Define capabilities for KV cache management, memory optimization, and token lifecycle efficiency (prefill vs decode)
  • Partner with engineering to enable hardware–software co-design, improving performance across accelerators and interconnects
  • Shape the development of inference platform capabilities that deliver measurable gains in latency, throughput, and cost efficiency
  • Own and evolve control plane services (APIs, policy engines, context/state management, usage/accounting)
  • Translate complex technical systems into clear product value and customer impact
  • Align multi-functional collaborators across infrastructure, research, and product organizations
  • Define and track key performance metrics (P99 latency, TTFT, throughput, cost per inference/token, availability)
  • Influence platform direction across constantly evolving AI workloads (LLMs, multimodal, agentic systems)

Required Skills And Experience

  • Proven experience in product management within AI/ML infrastructure, cloud platforms, or distributed systems
  • Deep understanding of AI inference systems and large-scale serving architectures
  • Solid understanding of LLM inference concepts (prefill vs decode, KV cache, token streaming)
  • Demonstrated ability to deliver scalable, high-performance platform products
  • Strong technical expertise and ability to collaborate directly with engineering teams

Nice To Have Skills And Experience

  • Experience with accelerator-based systems and performance optimization
  • Familiarity with modern inference frameworks (e.g., TensorRT-LLM, vLLM)
  • Experience working on production-scale AI systems

Compensation

$260,000-$351,800 per year

Equal Opportunities

Arm is an equal opportunity employer, committed to providing an environment of mutual respect where equal opportunities are available to all applicants and colleagues. We are a diverse organization of dedicated and innovative individuals, and don’t discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.

#J-18808-Ljbffr
Vacancy posted more than 2 months ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Director, Product Management - AI Inference Platform (San Jose). Be the first to apply!