HPC Engineer - AI Infrastructure
$275kReady to take the next step in your career?
Join a VC-backed GPU cloud company at the founding stage, building the next generation of managed GPU compute and inference infrastructure. Backed by leading venture investors, the organization delivers fully managed AI infrastructure and offers the opportunity to help shape the technical foundation of a rapidly growing GPU cloud platform.
This company is seeking a Founding Engineer to take ownership of core GPU cloud infrastructure across Kubernetes, Slurm, GPU orchestration, distributed storage, high-bandwidth networking, and inference platforms. This hands-on role offers founding-level ownership, significant influence over platform architecture, and the opportunity to work closely with the founders to build and scale production AI infrastructure from the ground up.
Don’t miss out on this exciting opportunity and apply today!
Responsibilities:
- Design, build, and operate large-scale Kubernetes and Slurm based GPU clusters in production
- Build and manage GPU orchestration layers on top of core scheduling infrastructure
- Architect and operate distributed object storage, NVMe storage clusters, and storage systems for large-scale AI workloads
- Design and manage high-bandwidth networking supporting distributed training and inference at scale
- Design, deploy, and optimise a scalable inference stack for serving large AI models with low-latency, high-throughput GPU acceleration
- Build telemetry, observability, and automated remediation across the GPU fleet
- Implement power-aware infrastructure mechanisms including GPU power capping, power-aware scheduling, and telemetry loops
- Own operational health, reliability, and performance of the platform end to end
- Work directly with founders on architecture, roadmap, and technical strategy
- Help define engineering culture, standards, and hiring as one of the first technical team members
Skills/Must Have:
- 3 to 6 years of hands-on experience operating large-scale AI or HPC infrastructure in production
- Proven experience operating large-scale Kubernetes and Slurm clusters
- Experience building and managing GPU orchestration layers on top of core schedulers
- Strong understanding of distributed object storage, NVMe storage clusters, and high-bandwidth networking
- Deep knowledge of storage architectures for large-scale AI infrastructure
- Comfort operating with founding-level ownership across the full infrastructure stack
- Based in or willing to relocate to San Francisco
Benefits:
- Founding engineer equity
- Full benefits package
Salary:
- $275,000 Base
- ...safety, reliability, and responsible AI deployment over unchecked growth.... ...About the role As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible... ...efficiency of our supercomputing infrastructure. Our team empowers strong engineers...SuggestedFull time
- ...Job Description Job Description We’re looking for an experienced HPC infrastructure engineer to lead bringup, administration, and operations on is probably the largest anime AI training cluster in the world . You’ll serve as the bridge between our researchers and...SuggestedWork at officeVisa sponsorship
$220k - $300k
...000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role Sentry's Infrastructure Engineering team is what makes operating Sentry simple, safe, and seamless for every other engineering team in...SuggestedHourly payFull time$298k - $310k
...cutting-edge machine learning, data science, and anti-fraud AI, we have served over 18 million customers as of 2025... ...profitability for sustainable growth. This role The Senior Engineering Manager, Cloud Infrastructure leads the engineering team responsible for the company’s...SuggestedFull time$180k - $277k
...Seattle About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers... .... ~ Background in hyperscale, cloud, or HPC data centre environments (preferred). Technical...SuggestedContract workFor contractorsFlexible hours- ...Gimlet is building the next generation of AI infrastructure: large-scale AI datacenters and the... ...Role Gimlet Labs is seeking a Network Engineer to design, build, and scale the network... ...policies. Understand high‑performance AI/HPC networking concepts such as RoCEv2, InfiniBand...
$100k - $140k
...; Seattle About Nscale Nscale is the GPU cloud engineered for AI. We provide cost‑effective, high‑performance infrastructure for AI start‑ups and large enterprise customers... ...and storage or networking add‑ons. Deeper GPU/HPC concepts such as RDMA/InfiniBand, performant distributed...Flexible hours$250k
...Ready to architect AI infrastructure that powers next-generation research and cloud platforms... ...to join as a Senior Inference Platform Engineer at an early stage and help define the architecture... ...distributed systems (ML inference, HPC, or similar). ~ Proficiency in Python...Permanent employment$250k
...your career? Join a rapidly scaling AI cloud infrastructure provider building next-generation GPU... ...company is looking for a Senior Storage Engineer with experience supporting high-... ...performance storage platforms supporting AI and HPC workloads Manage and optimize...Permanent employmentRemote work- ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers... ...and on-call rotation for Network Engineering team You Have 10+ years of... ...Terraform/Ansible/Salt Hands-on with HPC/AI networking: RoCEv2 and/or InfiniBand...Work at officeLocal areaWork from homeFlexible hours
$100k - $300k
About Cogent Cogent is an Applied AI Lab building the next generation of AI agents for cybersecurity. AI has fundamentally... ...Deepmind and SAIL About the Role We're hiring a Senior+ Infrastructure Engineer to help build a secure, reliable, cost-effective, and agent-...Full time- ...About the Role We are seeking a Cloud Infrastructure Engineer to help design and evolve the platforms that power OpenAI’s products. In this... ...operate the systems that power some of the most widely used AI products in the world. Your work will directly impact the...Full time
- ...Who We Are Serval is the AI platform for IT teams — replacing legacy systems like... ...expanded into a horizontal automation engine adopted by HR, Finance, Legal, Security,... ...modern enterprises. As a Software Engineer, Infrastructure, you'll build and scale the foundational...Full time
- ...platform will ultimately become the perception engine for a company's physical footprint,... ...the fast-approaching world of physical AI and robotics. We are a small, fast-... ...Responsibilities Specter is hiring an ML infrastructure engineer to build and scale the machine...Full time
- ...better, more human customer experiences with AI. We are primarily an in-person company... .... What you'll do The Payments Infrastructure team builds the trust boundary between a... ...data. Make payments something other engineers can use without becoming compliance experts...Full timeFlexible hours
$128.5k - $200k
...Semgrep gets smarter as you build, with AI that learns your context to cut false positives... ...dev . About the role Semgrep’s Infrastructure team is responsible for the cloud... ...-security space, collaborate with other engineers on the Infrastructure team to create a robust...Full timeCurrently hiringLocal areaRemote workWeekend work3 days per week- ...inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence... ...uniting applied AI research, flexible infrastructure, and seamless developer tooling, we... ...Conviction. Join us and help build the platform engineers turn to to ship AI products. THE...Full timeFlexible hours
$131k - $176k
...Wanna join the adventure? As a Space Infrastructure Software Engineer, you are responsible for scaling our ability to operate a heterogeneous... ...service across Earth observation, IoT connectivity, on-orbit AI, national security missions, and more. Leveraging our...Full timeTemporary workWork at officeRelocation packageFlexible hours- ...Combinator. We're solving the hard part of AI-driven software creation: correctness,... ...You'll Be Responsible For Platform & Infrastructure Maintain stability of our platform... ...~4+ years of software/platform engineering experience with production systems ~ Strong...Full timeFlexible hours
- ...About the Team The ChatGPT team works across research, engineering, product, and design to bring OpenAI’s technology to the world. We... ...to learn from deployment and broadly distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and...Full timeWork at officeRelocation package
- ...About the Role This role broadly owns infrastructure across the stack. If it’s running in the... ...The company embraces both large-scale AI and robotics as core to its DNA. Our team... ...GPT-4 to hundreds of millions of users, engineered the foundations of autonomous driving, built...Full time
$209k - $240k
...notes, projects, calendar, and email—with AI built in to find answers and automate... ...office workdays. About the Product Infrastructure Team: The Product Infrastructure team... ...classes of problems up-front for product engineers. Solve hard technical challenges such...Full timeWork at officeLocal area- ...About the Role We are looking for engineers to operate the next generation of compute... ...distributed systems engineering with hands-on infrastructure work on our largest datacenters. You... ...computing About OpenAI OpenAI is an AI research and deployment company...Full time
- ...About Lightfield Lightfield is an AI-native CRM that assembles itself from your email... ...an experienced, creative, and versatile engineer who is eager to tackle the challenge of... ..., building, and scaling the core infrastructure and systems powering Lightfield's AI-driven...Full timeWork from home
- ...hardware caching Much more! About the Role As an engineer within Fleet infrastructure, you will design, write, deploy, and operate infrastructure... ...system in support of OpenAI's mission to advance AI capabilities responsibly. This role is based in San Francisco...Full timeWork at officeRelocation package
- ...About the Team We’re hiring software engineers to join our broader Infrastructure organization, which supports multiple high-impact teams. Depending on your... ...that powers some of the most widely used AI systems in the world. You’ll help ensure our systems are...Full time
- ...About Runloop Runloop.ai is building the foundational infrastructure for the next generation of AI development. We provide AI engineers and data scientists with lightning-fast, secure, and reproducible code sandboxes for agents. Our platform enables teams to experiment...Full timeWork at officeWork from home1 day per week
- ...and cost-efficient requires world-class infrastructure. The Caching Infrastructure team is responsible... .... We’re looking for an experienced engineer to help design and scale this critical... .... About OpenAI OpenAI is an AI research and deployment company...Full time
$215k - $265k
...reduce the risks from scheming frontier AI systems. We work with and are trusted by... ...mitigations. We're looking for a Software Engineer to build the platform that the rest of... .... Build and maintain Apollo's cloud infrastructure . This means IaC, networking,...Full timeWork experience placementWork at officeImmediate startVisa sponsorshipFlexible hours- ...built world At Bedrock, we’re moving AI out of the lab and into the real world.... ...project schedules of billion-dollar infrastructure projects and improving safety on job sites... ...with construction veterans and world-class engineers to solve physical-world problems that...Full timeWork at officeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to HPC Engineer - AI Infrastructure. Be the first to apply!
- senior infrastructure engineer San Francisco, CA
- infrastructure engineering manager San Francisco, CA
- infrastructure engineer San Francisco, CA
- security infrastructure engineer San Francisco, CA
- lead infrastructure engineer San Francisco, CA
- infrastructure developer San Francisco, CA
- principal infrastructure engineer San Francisco, CA
- remote infrastructure engineer San Francisco, CA
- data infrastructure engineer San Francisco, CA
- senior infrastructure engineer



