Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Infra Platform PM: Scale Compute, Reliability & Cost

Cssmerge

Atoms is seeking an experienced Infrastructure Product Manager to lead strategy for internal infrastructure teams, overseeing a complex ecosystem from Core Compute (Kubernetes) to in‑house storage primitives and security platforms. You will drive roadmaps, ensure reliability, and optimize costs at scale. You’ll collaborate with engineering leadership to align product development with long‑term architecture while championing developer velocity and efficient operations. #J-18808-Ljbffr Cssmerge

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Infra Platform PM: Scale Compute, Reliability & Cost in San Francisco, CA vacancy
  • $190k - $253.75k

     ...data and AI infrastructure platform so our customers can use deep...  ...for interfacing with data to scaling our services and...  ...started.At Databricks, the Compute Infrastructure organization...  ...deliver extreme elasticity, reliability and cost efficiency.We are looking for... 
    Platform
    Local area
    Immediate start
    Worldwide

    DataBricks

    San Francisco, CA
    21 hours ago
  • Mercor in San Francisco is seeking a Platform Product Manager to own the LLM gateway and...  ...high-leverage role partnering with the infra team and leadership to translate...  ...with internal PMs and engineers to ensure reliability, cost transparency, and scalable usage policies... 
    Platform

    Mercor

    San Francisco, CA
    3 days ago
  •  ...internal infrastructure teams, enabling developers to move faster and reduce costs. You will oversee a complex ecosystem from Core Compute (Kubernetes) to in-house storage primitives, driving reliability and efficiency while partnering with engineering leadership to align... 
    Platform

    ATOMS Careers page

    San Francisco, CA
    2 days ago
  • Harvey is building the AI platform trusted by the world’s leading law firms and enterprises...  ...Engineer to design, operate, and scale Harvey’s compute, networking, Kubernetes, and workflow platforms. You’ll improve reliability, security, and efficiency while partnering... 
    Platform

    Neura Market

    San Francisco, CA
    21 hours ago
  •  ...Product Manager to lead strategy for our internal services and platforms, including Core Compute (Kubernetes) and Networking (Service Mesh/Istio). You’ll help developers move faster, ship reliably, and reduce costs across cloud and on‑prem deployments. As a force multiplier... 
    Platform

    Atoms

    San Francisco, CA
    4 days ago
  • Harvey is building the AI platform trusted by the world’s leading law firms and enterprises...  ...Engineer to design, operate, and scale Harvey’s global compute and networking infrastructure,...  ...foundations. You will improve reliability, capacity planning, automation, and... 
    Platform

    Harvey

    San Francisco, CA
    2 days ago
  • $152.5k - $205k

     ...s leading internet financial platform companies, building the foundation...  ...to power trusted, internet-scale financial innovation. Learn...  ...for:As a Senior Site Reliability Engineer on Circle’s platform...  ..., performance, security, and cost-effectiveness of the systems... 
    Platform
    Flexible hours

    Circle

    San Francisco, CA
    4 days ago
  •  ...role you will help scale and optimize our...  ...managing GPU/TPU compute and job orchestration...  ...‑scale training reliable, reproducible, and...  ...research, data, and platform engineers to...  ...while controlling cost. Partner with researchers...  ...needs into infra capabilities and guide... 
    Platform
    Full time

    Monograph

    San Francisco, CA
    4 days ago
  • $280k - $350k

     ...best data and AI platform. The Databricks AI...  ...to all.About the Scaling Research TeamThe Databricks...  ...efficiency, and compute efficiency through...  ...into performant, reliable infrastructure....  ..., throughput, and cost.Work hands‑on with...  ...systems and infra teams to push the... 
    Platform
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    2 days ago
  • $150k - $300k

    Platform Engineer — Infra / Reliability Specialist Join to apply for the Platform Engineer — Infra / Reliability Specialist role at Poly Platform Engineer...  ...service to store files at breathtaking speed and scale, and improving our AI architecture to become the "search... 
    Platform
    Full time
    Local area

    Poly

    San Francisco, CA
    4 days ago
  • $202k

     ...Software Engineer - ML Infra About the Role &...  ...accurately and on time to scaling the simulation platforms that drive billions...  ...where performance, reliability, and scale cannot be...  ...'s degree in Computer Science, Computer Engineering...  ..., and operational cost. Exceptional... 
    Platform
    Full time
    Work at office
    Remote work

    Uber

    San Francisco, CA
    4 days ago
  • Scale AI, Inc. in San Francisco is seeking an experienced Platform Product Manager who will own the core infrastructure that underpins all FD team work, define the foundation, and ensure reliable delivery. You will prioritize platform work, identify patterns, maintain a... 
    Platform

    Scale AI, Inc.

    San Francisco, CA
    2 days ago
  • Vapi is hiring a Product Manager to scale our Agent Platform, focusing on agent building and model management for voice AI across customer lifecycles...  ...in a growing SF startup environment. You bring 5+ years of PM experience, strong technical reasoning, and a proven ability... 
    Platform

    Doist

    San Francisco, CA
    2 days ago
  • DoorDash is hiring a Senior Software Engineer to lead the Spark Platform, setting the long-term direction for our in-house Spark...  ...spanning runtime, shuffle service, and scheduler, ensuring reliability at scale across data, analytics, and ML workloads. You will partner with... 
    Platform

    Doordashusa

    San Francisco, CA
    4 days ago
  •  .... is seeking a engineer for the Supercomputing Platform & Infrastructure to design, build, and operate large-scale GPU infrastructure powering model training and...  ...Kubernetes clusters, and ensure reproducibility and reliability of thousands of GPUs. This role offers visa... 
    Platform
    Visa sponsorship
    Relocation package

    Magic AI, Inc

    San Francisco, CA
    12 hours ago
  • Reflection's Compute Platform team leads a Kubernetes-based, multi-cloud compute layer across neo-clouds. As Compute Platform Lead you will...  ...on to contribute and ensure the compute fleet supports large-scale training runs. This role also manages vendor relationships, drives... 
    Platform

    B Capital

    San Francisco, CA
    2 days ago
  • $300k - $360k

    Director, Infrastructure Compute Procurement Anthropic...  ...mission is to create reliable, interpretable, and...  ...assets at scale. This is a function‑building...  ...own supplier strategy, cost optimization, and category...  ...lease administration platform, partnering with Infra/Workplace, Finance,... 
    Platform
    Permanent employment
    Contract work
    Interim role
    Work at office
    Local area
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    21 hours ago
  •  ...evaluate, and serve agent telemetry at scale. What You'll Do Design and build...  ...petabyte scale. Partner with the platform and infra teams on scalability, reliability, and the systems that back our...  ...managing the rate-limit, retry, and cost dynamics that come with them. Familiarity... 
    Platform

    Judgment Labs Inc.

    San Francisco, CA
    2 days ago
  •  ...of: distributed data platforms real-time decision systems...  ...and orchestration reliability for long-horizon AI...  ...systems like: streaming compute platforms distributed...  ...engines large-scale data infrastructure many...  ...Foundations Engineer (Deep Infra) to design and operate... 
    Platform

    Rox Data Corp

    San Francisco, CA
    4 days ago
  •  ...for a world-class Site Reliability Engineer to ensure the...  ...our AI infrastructure platform. You'll be building...  ...that power agentic AI at scale. Your mission: keep...  ...stateful, serverless compute engine rock-solid as we...  ...with the founders, the infra team, and the dev team... 
    Platform

    Blaxel, Inc

    San Francisco, CA
    2 days ago
  •  ...cause and effect. We believe that scaling on physics will enable an understanding...  ...it all, delivering performant, reliable, and cost‑efficient compute to ensure research is able to iterate...  ...‑as‑code Knowledge of cloud platforms (GCP, AWS, or Azure) and their ML/AI... 
    Platform

    Causal Labs

    San Francisco, CA
    2 days ago
  • $230k - $385k

     ...also powering production at scale. We own the platform end to end: backend systems...  ...durable, available, and cost-effectiveEvolve the federation...  ...performance, reliability, and operational excellence...  ...offer of employment: protect computer hardware entrusted to you from... 
    Platform
    Work at office
    Local area
    Flexible hours

    OpenAI

    San Francisco, CA
    21 hours ago
  • $232k - $319k

     ...too, let's talk.The Infrastructure Platform and Shared Services TeamOkta...  ...technical leader to help us continue to scale the service with great people and reliable, cost-effective, and efficient...  ...tooling. What you’ll be doing Lead the Infra platform and shared services org... 
    Platform
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    3 days ago
  •  ...evaluation possible at scale. This is a critical...  ...users that scales, is reliable, and makes the complexities...  ..., usage metering, cost attribution, audit logging...  ...with our core evaluation platform, Arena data, and...  ...Familiarity with the modern AI infra stack (vLLM, LiteLLM,... 
    Platform
    Permanent employment
    Shift work

    Arena AI

    San Francisco, CA
    3 days ago
  • $157k - $239k

     ...the adventure?As a Site Reliability Engineer with strong...  ...observable, and secure as we scale. You coordinate with...  ...like a regular SRE: platform reliability and...  ...environmentDegree in Computer Science or a related field...  ...experience in FinOps and cost-optimized... 
    Platform
    Full time
    Temporary work

    Loft Orbital

    San Francisco, CA
    3 days ago
  • $290k - $365k

     ...s mission is to create reliable, interpretable, and steerable...  ...Program Manager on the Compute team, you will help...  ...running efficiently at scale. Our compute fleet is...  ...providers and hardware platforms. Partner with engineering...  ...between utilization, cost, latency, and reliability... 
    Platform
    Visa sponsorship

    Anthropic

    San Francisco, CA
    4 days ago
  • $126k - $204.5k

     ...developer velocity and reliability for the entire engineering...  ...an infrastructure platform that enables the engineering...  ...and maintain high-scale developer tooling and backend...  ...for performance, cost-efficiency, and near real...  ...logic.Operating Systems & Compute: Deep knowledge of... 
    Platform
    Full time
    Local area
    Remote work

    Palo Alto Networks

    San Francisco, CA
    2 days ago
  • $213k - $263k

     ...to a range of vehicle platforms and product use cases....  ...autolabels at a massive scale, serving as the foundation...  ...state-of-the-art computer vision, deep learning,...  ...Collaborate closely with the ML Infra, Perception, Behavior,...  ...designing scalable and reliable systems. ~ Experience... 
    Platform
    Full time
    Remote work

    Waymo

    San Francisco, CA
    2 days ago
  • $230k - $390k

     ...Sierra, we’re creating a platform to help businesses...  ...Engineer on our Site Reliability team at Sierra, you will...  ...robust, performant, and cost-effective operation....  ...engineering teams.Degree in Computer Science or a related...  ...models, or large-scale model deployment.Past... 
    Platform
    Full time
    Flexible hours

    Sierra

    San Francisco, CA
    21 hours ago
  •  ...way. The RoleThe Director of Platform & Reliability Engineering will lead a...  ...Forge's platform capabilities scale with the business.Location:...  ...experience.Lead cloud capacity, cost, and architecture planning...  ...roadmaps.Bachelor's degree in Computer Science or a closely related... 
    Platform
    Work at office
    Local area
    2 days per week
    3 days per week

    Forge Global

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Infra Platform PM: Scale Compute, Reliability & Cost. Be the first to apply!