Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

HPC Engineer - AI Infrastructure

$275k
Full-time

Ready to take the next step in your career?

Join a VC-backed GPU cloud company at the founding stage, building the next generation of managed GPU compute and inference infrastructure. Backed by leading venture investors, the organization delivers fully managed AI infrastructure and offers the opportunity to help shape the technical foundation of a rapidly growing GPU cloud platform.

This company is seeking a Founding Engineer to take ownership of core GPU cloud infrastructure across Kubernetes, Slurm, GPU orchestration, distributed storage, high-bandwidth networking, and inference platforms. This hands-on role offers founding-level ownership, significant influence over platform architecture, and the opportunity to work closely with the founders to build and scale production AI infrastructure from the ground up.

Don’t miss out on this exciting opportunity and apply today!

Responsibilities:

  • Design, build, and operate large-scale Kubernetes and Slurm based GPU clusters in production
  • Build and manage GPU orchestration layers on top of core scheduling infrastructure
  • Architect and operate distributed object storage, NVMe storage clusters, and storage systems for large-scale AI workloads
  • Design and manage high-bandwidth networking supporting distributed training and inference at scale
  • Design, deploy, and optimise a scalable inference stack for serving large AI models with low-latency, high-throughput GPU acceleration
  • Build telemetry, observability, and automated remediation across the GPU fleet
  • Implement power-aware infrastructure mechanisms including GPU power capping, power-aware scheduling, and telemetry loops
  • Own operational health, reliability, and performance of the platform end to end
  • Work directly with founders on architecture, roadmap, and technical strategy
  • Help define engineering culture, standards, and hiring as one of the first technical team members

Skills/Must Have:

  • 3 to 6 years of hands-on experience operating large-scale AI or HPC infrastructure in production
  • Proven experience operating large-scale Kubernetes and Slurm clusters
  • Experience building and managing GPU orchestration layers on top of core schedulers
  • Strong understanding of distributed object storage, NVMe storage clusters, and high-bandwidth networking
  • Deep knowledge of storage architectures for large-scale AI infrastructure
  • Comfort operating with founding-level ownership across the full infrastructure stack
  • Based in or willing to relocate to San Francisco

Benefits:

  • Founding engineer equity
  • Full benefits package

Salary:

  • $275,000 Base
Vacancy posted more than 2 months ago
Similar jobs that could be interesting for youBased on the HPC Engineer - AI Infrastructure in San Francisco, CA vacancy
  • $230k

     ...safety, reliability, and responsible AI deployment over unchecked growth.About the roleAs a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for...  ...efficiency of our supercomputing infrastructure.Our team empowers strong engineers... 
    Suggested
    Work at office
    Local area
    Flexible hours

    OpenAI

    San Francisco, CA
    2 days ago
  •  ...Job Description Job Description We’re looking for an experienced HPC infrastructure engineer to lead bringup, administration, and operations on is probably the largest anime AI training cluster in the world . You’ll serve as the bridge between our researchers and... 
    Suggested
    Work at office
    Visa sponsorship

    Spellbrush

    San Francisco, CA
    28 days ago
  • $200k

     ...opportunities? Join a rapidly scaling AI cloud infrastructure provider building next-generation GPU...  ...opportunity for a Senior Storage Engineer to take ownership of the high-performance...  ...performance storage platforms supporting AI and HPC workloads Manage and optimize... 
    Suggested
    Remote work

    Jobleads-US

    San Francisco, CA
    1 day ago
  • $225k

     ...your career? Join a rapidly scaling AI cloud infrastructure provider building next-generation GPU...  ...opportunity has arisen for a Senior Network Engineer to design, deploy, and operate ultra-...  ...environments optimized for AI and HPC traffic patterns. Ready to make a move... 
    Suggested
    Remote work

    Jobleads-US

    San Francisco, CA
    2 days ago
  • $190k - $270k

    Staff Software Engineer - AI Research InfrastructureP-1215At Databricks, we are obsessed with...  ...a Staff Software Engineer, AI Research Infrastructure, you will be developing and running the...  ...processing, and model training (e.g., HPC clusters, GPU fleets, or cloud‑based systems... 
    Suggested
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    15 hours ago
  • $238k - $300k

     ...in a single system. Combined with industry leading AI, the Motive platform gives you complete visibility and...  ...more.About the Role: We are looking for a seasoned engineering leader who will lead the Core Infrastructure and Operations Engineering. You will be responsible... 
    Temporary work
    Work experience placement

    Motive Technologies

    San Francisco, CA
    2 days ago
  • $168k - $252k

    About Vercel:Vercel is the agentic infrastructure company. We free people and agents to ship what...  .... As the team behind Next.js, v0, and AI SDK, we create products that help builders...  ....About the role:We're hiring a DevRel Engineer to work inside the product teams building... 
    Work at office
    Remote work
    Work from home
    Worldwide
    Monday to Friday
    Flexible hours

    Vercel

    San Francisco, CA
    4 days ago
  • $237.6k - $297k

    We are seeking a highly skilled Infrastructure Security Engineer to join our team. This role is integral to ensuring the security and integrity of our...  ....About Us:At Scale, our mission is to develop reliable AI systems for the world's most important decisions. Our products... 
    Full time

    Scale AI

    San Francisco, CA
    1 day ago
  • $208.45k - $364.8k

     ...career you love? It’s Possible.At Pinterest, AI isn't just a feature, it's a powerful...  ...process here.Pinterest is seeking an Engineering Manager to lead the Ads Serving Platform...  ...the end-to-end ad-request lifecycle.Drive infrastructure performance, cost efficiency, and... 
    Work at office
    Local area
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    1 day ago
  •  ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers...  ...and on-call rotation for Network Engineering teamYouHave 10+ years of experience in...  ...like Terraform/Ansible/SaltHands-on with HPC/AI networking: RoCEv2 and/or InfiniBand... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda

    San Francisco, CA
    2 days ago
  •  ...Gimlet is building the next generation of AI infrastructure: large-scale AI datacenters and the...  ...Gimlet Labs is seeking a Network Engineer to design, build, and scale the network...  ...policies. Understand high‑performance AI/HPC networking concepts such as RoCEv2, InfiniBand... 

    Gimlet Labs

    San Francisco, CA
    2 days ago
  • $266k

     ...security culture.About the RoleOpenAI is seeking a Security Engineer to join our Infrastructure Security (InfraSec) team. InfraSec protects the...  ...storage, and the critical services that power our frontier AI models. Our charter includes securing everything from bare... 
    Work at office
    Local area
    Remote work
    Flexible hours

    OpenAI

    San Francisco, CA
    2 days ago
  • $190k - $280k

     ...Senior Network Engineer San Francisco About the Role Together AI is looking for a Senior Network Engineer to design...  ...and operate the global network infrastructure supporting our production services...  ...supporting GPU clusters, HPC environments, distributed storage... 
    Full time

    Together AI

    San Francisco, CA
    2 days ago
  • $225k - $275k

     ...intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we...  ...a Senior Staff Network Deployment Engineer to serve as the technical owner of how...  ...footprint of high-performance compute (HPC) and GPU-based AI infrastructure, you will... 
    Temporary work
    Remote work

    Crusoe

    San Francisco, CA
    3 days ago
  • $188k - $275k

     ...CoreWeave is The Essential Cloud for AI™. Built for pioneers by...  ...CoreWeave combines superior infrastructure performance with deep...  ...What You'll Do: The Field Engineering organization at CoreWeave is...  ...InfiniBand/RoCE fabric validation and HPC performance benchmarking,... 
    Permanent employment
    Full time
    Contract work
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    San Francisco, CA
    1 day ago
  • $250k

     ...Ready to architect AI infrastructure that powers next-generation research and cloud platforms...  ...to join as a Senior Inference Platform Engineer at an early stage and help define the architecture...  ...distributed systems (ML inference, HPC, or similar). ~ Proficiency in Python... 
    Permanent employment
    San Francisco, CA
    more than 2 months ago
  • $235k - $260k

     ...fit. About the Company Our client builds the data layer that AI agents run on. One API call turns any URL into clean, structured...  ...to power real products. The Opportunity As a Search Engineer, you will own ranking quality and relevance for an LLM-driven search... 
    Temporary work
    Immediate start

    Lavendo

    San Francisco, CA
    16 days ago
  • $115k - $158k

    Who We AreHP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s...  ...assembling a diverse, world-class team—engineers, designers, researchers, and product minds...  ...Software Engineer, Tooling and Development Infrastructure, you will play a critical role in... 
    Full time
    Temporary work
    Local area
    Flexible hours

    HP IQ

    San Francisco, CA
    4 days ago
  • $150k - $200k

    Who We AreNotion is the collaborative AI workspace where teams and agents think together...  ...life’s work.About the Role:The Product Infrastructure team works on creating abstractions and...  ...of problems up-front for product engineers.Solve hard technical challenges such as... 
    Local area

    Notion Labs

    San Francisco, CA
    15 hours ago
  • $299k - $334k

    Who We AreNotion is the collaborative AI workspace where teams and agents think together...  ...a reliable system often means finding an engineer, provisioning a database, and building an application.We’re building database infrastructure for everyone else. We want everyday... 
    Local area
    Weekend work

    Notion Labs

    San Francisco, CA
    3 days ago
  • $175k - $300k

     ...About the Role Join an early-stage B2B integration infrastructure team as its first infrastructure engineer. You will own the platform that runs integrations and...  .... Strong engineering judgment when using AI coding agents, and experience with self-hosted, single... 
    Visa sponsorship

    Clera

    San Francisco, CA
    2 days ago
  • $166k - $225k

     ...breakthroughs. We do this by building and running the world's best data and AI infrastructure platform so our customers can use deep data insights to improve their business. Founded by engineers — and customer obsessed — we leap at every opportunity to solve technical... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    2 days ago
  • $250k - $300k

     ...9Salary$250,000-$300,000 / yearSoftware Engineer - Growth InfrastructureJob Number: 26-01...  ...looking for a Software Engineer - Growth Infrastructure for our client in San Francisco, CA....  ...onboarding, and matching at scale for a top AI startup.This is not a traditional growth... 
    Full time
    Internship
    Visa sponsorship

    Eclaro International

    San Francisco, CA
    2 days ago
  • $196k - $294k

    About Vercel:Vercel is the agentic infrastructure company. We free people and agents to ship what...  .... As the team behind Next.js, v0, and AI SDK, we create products that help builders...  ...starts in v0, in an agent, or in an engineer’s editor, every customer builds — and millions... 
    Work from home
    Worldwide
    Flexible hours

    Vercel

    San Francisco, CA
    2 days ago
  • $122.4k - $158.4k

    The backbone of any technology org is its infrastructure, come join the team giving that skeleton...  ...team has built.We are looking for an engineer to join our team to help build out the “...  ...agentic infrastructure: systems that let AI agents safely interact with our infrastructure... 

    Mercury

    San Francisco, CA
    3 days ago
  •  ...seamlessly in the face of incredible growth.As a Senior Software Engineer (Infrastructure), you will be a core technical contributor on the IT...  ...logging and metrics enabled by defaultTooling, Scripting & AI : Build internal CLI tools,AI plugins and automation scripts... 
    Worldwide

    DataBricks

    San Francisco, CA
    1 day ago
  • $230k - $390k

    About usSierra is the leading platform for customer-facing AI agents, working with many of the world's biggest brands — including...  ...teams for Google Workspace.What you’ll doAs a Software Engineer, Infrastructure at Sierra, you will be responsible for designing, building,... 
    Full time
    Flexible hours

    Sierra

    San Francisco, CA
    1 day ago
  • $164k

    About the roleThe Infrastructure Engineering organization comprises three sub-teams: Core Infrastructure (which manages the infrastructure systems...  ...and technical expertise.Be a constant advocate and adopter of AI tools and evangelize agentic AI workflows in a way that is... 
    Full time
    Work at office
    Local area
    Remote work

    Chime

    San Francisco, CA
    2 days ago
  • $293k - $385k

    About the TeamThe GPT Infrastructure team builds systems that turn advances in model inference...  ...execution, compilers and runtimes, performance engineering, secure partner integrations, evaluation...  ...teams.Preferred SkillsExperience with AI infrastructure, model inference,... 
    Work at office
    Local area
    Remote work
    Flexible hours

    OpenAI

    San Francisco, CA
    15 hours ago
  • $230k - $405k

    About the Team:Compute Infrastructure builds the platform that turns enormous amounts of compute into a reliable engine for frontier AI. We design, provision, schedule, operate, and optimize the systems that connect accelerators, CPUs, networks, storage, data centers,... 
    Work at office
    Local area
    Flexible hours

    OpenAI

    San Francisco, CA
    15 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to HPC Engineer - AI Infrastructure. Be the first to apply!