Software Engineer, GPU Infrastructure - HPC
OpenAI
About the team
The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth.
About the role
As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise.
Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure.
Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale.
In this role, you will:
Build and maintain automation systems for provisioning and managing server fleets.
Develop tools to monitor server health, performance, and lifecycle events.
Collaborate with clusters, networking, and infrastructure teams.
Partner with external operators to ensure a high level of quality.
Identify and fix performance bottlenecks and inefficiencies.
Continuously improve automation to reduce manual work.
You might thrive in this role if you have:
Experience managing large-scale server environments.
A balance of strengths in building and operationalizing.
Proficiency in Python, Go, or similar languages.
Strong Linux, networking, and server hardware knowledge.
Comfort digging into noisy data with SQL, PromQL, and Pandas or any other tool.
Prior hardware expertise is not required for this role.
Bonus Skills:
Experience with low level details of hardware components, protocols, and associated Linux tooling (e.g., PCIe, Infiniband, networking, power management, kernel perf tuning)
Knowledge of hardware management protocols (e.g., IPMI, Redfish).
High-performance computing (HPC) or distributed systems experience.
Prior experience developing, managing, or designing hardware.
Familiarity with monitoring tools (e.g., Prometheus, Grafana).
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity.
We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.
For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement .
Qualified applicants with arrest or conviction records will be considered for employment in accordance with applicable law, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.
We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link .
At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.
- ...This role will support the fleet infrastructure team at OpenAI. The fleet team focuses on running... ..., most reliable, and frictionless GPU fleet to support OpenAI’s general purpose... ...Much more! About the Role As an engineer within Fleet infrastructure, you will design...SuggestedFull timeWork at officeRelocation package
$300k
...Join a stealth-mode hyperscale infrastructure startup building a 300MW+ AI... .... The Principal Software Engineer will take ownership of the software... ..., workload scheduling, GPU resource management, infrastructure... ...powering large-scale HPC and AI infrastructure environments...SuggestedFull timeRemote workFlexible hours- ...applied AI research, flexible infrastructure, and seamless developer... ...and help build the platform engineers turn to to ship AI products.... ...foundational engineers to lead our GPU Networking efforts, making RDMA... ...to architect the software fabric that unifies thousands...SuggestedFull timeFlexible hours
$275k
...your career? Join a VC-backed GPU cloud company at the founding... ...GPU compute and inference infrastructure. Backed by leading venture investors... ...is seeking a Founding Engineer to take ownership of core GPU... ...operating large-scale AI or HPC infrastructure in production...SuggestedFull timeRelocation- ...About the Role This role broadly owns infrastructure across the stack. If it’s running in the... ...responsible for: The compute (both CPU and GPU compute) that’s required to cost... ...GPT-4 to hundreds of millions of users, engineered the foundations of autonomous driving, built...SuggestedFull time
- ...working systems and build any software needed for running large-... ...the Role We are looking for engineers to operate the next generation... ...systems engineering with hands-on infrastructure work on our largest... ...bare-metal Linux environments, GPU hardware, and large-scale networking...Full time
- ...Exa is building a search engine from scratch to serve every AI agent. We build massive-scale infrastructure to crawl the web, train state-of... ...compute, we also own a $5M H200 GPU cluster (and soon 5x'ing that... ...Design GPU scheduling software so we max out our cluster utilization...Full timeH1b
$215k - $265k
...scheming mitigations. We're looking for a Software Engineer to build the platform that the rest of... ...Build and maintain Apollo's cloud infrastructure . This means IaC, networking,... ...signals for their service. Multi-cloud GPU orchestration: The platform that finds...Full timeWork experience placementWork at officeImmediate startVisa sponsorshipFlexible hours- ...Background Specter is creating a software-defined "control plane" for... ...become the perception engine for a company's physical footprint... ...Specter is hiring an ML infrastructure engineer to build and scale the... ...DDP, DeepSpeed, Ray) and GPU cluster management. Strong...Full time
$230k - $405k
About the Team:Compute Infrastructure builds the platform that turns enormous... ...of compute into a reliable engine for frontier AI. We design,... ...data centers, orchestration software, agent infrastructure, developer... ...protocols, RDMA, NCCL, GPU hardware behavior, benchmarking...Work at officeLocal areaFlexible hours- ...Platform and builds the shared infrastructure that helps DoorDash, Wolt,... ...DeepSeek) ourselves — real-time GPU serving, high-throughput... ...model serving and inference engines, fine-tuning and training pipelines... ...of industry experience in software engineeringDeep backend engineering...Hourly payWork at officeLocal areaRemote workFlexible hours
- ...performance, flexibility, and resiliency across our infrastructure. We are forming a team to generalize our... ...AMD. About the Role We’re hiring engineers to scale and optimize OpenAI’s inference infrastructure across emerging GPU platforms. You’ll work across the stack -...Full time
- ...applied AI research, flexible infrastructure, and seamless developer... ...and help build the platform engineers turn to to ship AI products.... ...Baseten is building its own GPU infrastructure for large-scale... ...not. We are hiring a Lead Software Engineer to build a first-class...Full timeFlexible hours
- ...Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable... .... Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re seeking a GPU Kernel Engineer to join our team at the cutting...Full timeFlexible hours
$190k - $270k
Staff Software Engineer - AI Research InfrastructureP-1215At Databricks, we are obsessed with... ...Software Engineer, AI Research Infrastructure, you will be developing and running... ...processing, and model training (e.g., HPC clusters, GPU fleets, or cloud‑based systems)Enable...Local areaWorldwide- ...will de-risk the largest infrastructure build-out in history. When people finance GPU clusters, the... ...looking for a high agency engineer to help build the compute... ...with the orchestration software managing virtual machines... ...running on cutting-edge HPC hardware worth billions...Long term contractFull timeContract workFixed term contractWork at officeLocal areaVisa sponsorshipShift work
$230k - $385k
...security culture.About the RoleOpenAI is seeking a Security Software Engineer to join the Infrastructure Security (InfraSec) team.InfraSec safeguards the core of OpenAI’s research and production environments—GPU supercomputing clusters, multi-cloud infrastructure,...Work at officeLocal areaFlexible hours- ...Software Engineer Voxel's perception system is the technical core of everything we ship. Our... ...software engineer to own the ML Infrastructure that powers how Voxel trains and ships... ...Prefect, or similar) Familiarity with GPU performance profiling and optimization...Work at officeFlexible hours
- ...Exa Infrastructure Engineer Exa is an applied AI lab building a search engine unlike the world has... ...org. That could mean building GPU cluster orchestration in Kubernetes, map... ...scales for our agent fleet Automate software maintenance and improvements for the whole...H1b
$176k - $220k
...institutions Work together with engineers, scientists, operators, and... ...AI Human data is the core infrastructure to AI advancement. Frontier... ...We’re looking for a Senior Software Engineer to join our ML Infrastructure... ...fine-tuned models, including GPU serving, batching, and...Full timeWork at officeRemote workFlexible hours- About the TeamThe GPT Infrastructure team builds systems that turn advances... ...and runtimes, performance engineering, secure partner integrations... ...the RoleWe are seeking a software engineer to help build the platform... ...platforms.Familiarity with GPU or accelerator architecture,...Remote work
$200k - $250k
...Fluidstack, a leading cloud provider, is looking for a Software Engineer, Infrastructure Platform to build the foundational platforms that enable our... ...data Create DCIM platforms for rack operations, server/GPU deployment, OS installation, quality assurance, and white-...Full timeLocal area$180k - $250k
...next generation of AI products. We build the infrastructure, tools, and model access that teams need to... ...ambitious teams build on. You are a hands-on engineer who builds the software and processes that keep a large fleet of GPU servers healthy and productive. You write...Local areaRelocation package$148.7k - $201.2k
...disciplinary team of scientists, engineers, and technicians, on a... ...computer.We are looking to hire an HPC Platform Engineer to develop,... ...-performance computing (HPC) infrastructure on AWS that CQC scientists... ...maintaining clusters, managing OS and software stacks, building containers,...Local areaFlexible hours- ...Whoever deploys frontier compute infrastructure fastest will decide whether... ...teams spanning hardware and software. Speed and scale are our key... ...forward. The Production Engineering Team Examples of key... ...0s of GWs: at our scale, a GPU failure isn't a ticket. It's...Local area
$200k - $250k
...office. The opportunity We're hiring a Senior Cloud Infrastructure Engineer to join the Cloud Infrastructure & Platform (CIN) team at Altruist... .... Architect infrastructure for AI/ML workloads: GPU-enabled compute, SageMaker endpoints, Bedrock integration, vector...Work at officeImmediate start3 days per week- ...Whoever deploys frontier compute infrastructure fastest will decide whether... ...teams spanning hardware and software. Speed and scale are our key... ...forward. The Production Engineering Team Examples of key... ...platforms, so every new site and GPU generation lands cleanly...Local area
- ...Sciforium is an AI infrastructure company developing next-... ...hands-on support from AMD engineers the team is scaling... ...runtime, service, and GPU layers, working closely... ...) ~3+ years of software engineering experience... ...to open-source ML or HPC infrastructure Benefits...Full timeWork at officeFlexible hours
$209k - $240k
...office workdays. About the Product Infrastructure Team: The Product Infrastructure... ...classes of problems up-front for product engineers. Solve hard technical challenges such... ...values, and enthusiastic about making software toolmaking ubiquitous, we want to hear...Full timeWork at officeLocal area$170k - $216k
...across 15+ U.S. states. The Simulation Infrastructure team creates reliable, scalable, and... ...products that evaluate the Waymo Driver's software stack at a massive scale. We solve... ...for a broad range of customers Software Engineers, Product, Data Science, System Engineering...Full timeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer, GPU Infrastructure - HPC. Be the first to apply!
- agile software developer San Francisco, CA
- software developer internship no experience San Francisco, CA
- intermediate software engineer San Francisco, CA
- software engineer staff San Francisco, CA
- experienced software developer San Francisco, CA
- work from home software developer San Francisco, CA
- software developer no experience San Francisco, CA
- software developer fintech San Francisco, CA
- software data engineer San Francisco, CA
- financial software developer San Francisco, CA



