Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer, GPU Infrastructure - HPC

Full-time

OpenAI

About the team

The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth.

About the role

As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise.

Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure.

Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale.

In this role, you will:

  • Build and maintain automation systems for provisioning and managing server fleets.

  • Develop tools to monitor server health, performance, and lifecycle events.

  • Collaborate with clusters, networking, and infrastructure teams.

  • Partner with external operators to ensure a high level of quality.

  • Identify and fix performance bottlenecks and inefficiencies.

  • Continuously improve automation to reduce manual work.

You might thrive in this role if you have:

  • Experience managing large-scale server environments.

  • A balance of strengths in building and operationalizing.

  • Proficiency in Python, Go, or similar languages.

  • Strong Linux, networking, and server hardware knowledge.

  • Comfort digging into noisy data with SQL, PromQL, and Pandas or any other tool.

Prior hardware expertise is not required for this role.

Bonus Skills:

  • Experience with low level details of hardware components, protocols, and associated Linux tooling (e.g., PCIe, Infiniband, networking, power management, kernel perf tuning)

  • Knowledge of hardware management protocols (e.g., IPMI, Redfish).

  • High-performance computing (HPC) or distributed systems experience.

  • Prior experience developing, managing, or designing hardware.

  • Familiarity with monitoring tools (e.g., Prometheus, Grafana).

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. 

We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.

For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement .

Qualified applicants with arrest or conviction records will be considered for employment in accordance with applicable law, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.

We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link .

At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Software Engineer, GPU Infrastructure - HPC in San Francisco, CA vacancy
  •  ...This role will support the fleet infrastructure team at OpenAI. The fleet team focuses on running...  ..., most reliable, and frictionless GPU fleet to support OpenAI’s general purpose...  ...Much more! About the Role As an engineer within Fleet infrastructure, you will design... 
    Suggested
    Full time
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    1 day ago
  • $300k

     ...Join a stealth-mode hyperscale infrastructure startup building a 300MW+ AI...  .... The Principal Software Engineer will take ownership of the software...  ..., workload scheduling, GPU resource management, infrastructure...  ...powering large-scale HPC and AI infrastructure environments... 
    Suggested
    Full time
    Remote work
    Flexible hours
    San Francisco, CA
    18 days ago
  •  ...applied AI research, flexible infrastructure, and seamless developer...  ...and help build the platform engineers turn to to ship AI products....  ...foundational engineers to lead our GPU Networking efforts, making RDMA...  ...to architect the software fabric that unifies thousands... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  • $275k

     ...your career? Join a VC-backed GPU cloud company at the founding...  ...GPU compute and inference infrastructure. Backed by leading venture investors...  ...is seeking a Founding Engineer to take ownership of core GPU...  ...operating large-scale AI or HPC infrastructure in production... 
    Suggested
    Full time
    Relocation
    San Francisco, CA
    16 days ago
  •  ...About the Role This role broadly owns infrastructure across the stack. If it’s running in the...  ...responsible for: The compute (both CPU and GPU compute) that’s required to cost...  ...GPT-4 to hundreds of millions of users, engineered the foundations of autonomous driving, built... 
    Suggested
    Full time

    The Generalist

    San Francisco, CA
    1 day ago
  •  ...working systems and build any software needed for running large-...  ...the Role We are looking for engineers to operate the next generation...  ...systems engineering with hands-on infrastructure work on our largest...  ...bare-metal Linux environments, GPU hardware, and large-scale networking... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...Exa is building a search engine from scratch to serve every AI agent. We build massive-scale infrastructure to crawl the web, train state-of...  ...compute, we also own a $5M H200 GPU cluster (and soon 5x'ing that...  ...Design GPU scheduling software so we max out our cluster utilization... 
    Full time
    H1b

    Exa

    San Francisco, CA
    1 day ago
  • $215k - $265k

     ...scheming mitigations. We're looking for a Software Engineer to build the platform that the rest of...  ...Build and maintain Apollo's cloud infrastructure . This means IaC, networking,...  ...signals for their service. Multi-cloud GPU orchestration: The platform that finds... 
    Full time
    Work experience placement
    Work at office
    Immediate start
    Visa sponsorship
    Flexible hours

    Apollo Research

    San Francisco, CA
    1 day ago
  •  ...Background Specter is creating a software-defined "control plane" for...  ...become the perception engine for a company's physical footprint...  ...Specter is hiring an ML infrastructure engineer to build and scale the...  ...DDP, DeepSpeed, Ray) and GPU cluster management. Strong... 
    Full time

    S.e. Specter

    San Francisco, CA
    1 day ago
  • $230k - $405k

    About the Team:Compute Infrastructure builds the platform that turns enormous...  ...of compute into a reliable engine for frontier AI. We design,...  ...data centers, orchestration software, agent infrastructure, developer...  ...protocols, RDMA, NCCL, GPU hardware behavior, benchmarking... 
    Work at office
    Local area
    Flexible hours

    OpenAI

    San Francisco, CA
    3 days ago
  •  ...Platform and builds the shared infrastructure that helps DoorDash, Wolt,...  ...DeepSeek) ourselves — real-time GPU serving, high-throughput...  ...model serving and inference engines, fine-tuning and training pipelines...  ...of industry experience in software engineeringDeep backend engineering... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    2 days ago
  •  ...performance, flexibility, and resiliency across our infrastructure. We are forming a team to generalize our...  ...AMD. About the Role We’re hiring engineers to scale and optimize OpenAI’s inference infrastructure across emerging GPU platforms. You’ll work across the stack -... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...applied AI research, flexible infrastructure, and seamless developer...  ...and help build the platform engineers turn to to ship AI products....  ...Baseten is building its own GPU infrastructure for large-scale...  ...not. We are hiring a Lead Software Engineer to build a first-class... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  •  ...Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable...  .... Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re seeking a GPU Kernel Engineer to join our team at the cutting... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  • $190k - $270k

    Staff Software Engineer - AI Research InfrastructureP-1215At Databricks, we are obsessed with...  ...Software Engineer, AI Research Infrastructure, you will be developing and running...  ...processing, and model training (e.g., HPC clusters, GPU fleets, or cloud‑based systems)Enable... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    3 days ago
  •  ...will de-risk the largest infrastructure build-out in history. When people finance GPU clusters, the...  ...looking for a high agency engineer to help build the compute...  ...with the orchestration software managing virtual machines...  ...running on cutting-edge HPC hardware worth billions... 
    Long term contract
    Full time
    Contract work
    Fixed term contract
    Work at office
    Local area
    Visa sponsorship
    Shift work

    The San Francisco Compute Company

    San Francisco, CA
    1 day ago
  • $230k - $385k

     ...security culture.About the RoleOpenAI is seeking a Security Software Engineer to join the Infrastructure Security (InfraSec) team.InfraSec safeguards the core of OpenAI’s research and production environments—GPU supercomputing clusters, multi-cloud infrastructure,... 
    Work at office
    Local area
    Flexible hours

    OpenAI

    San Francisco, CA
    5 days ago
  •  ...Software Engineer Voxel's perception system is the technical core of everything we ship. Our...  ...software engineer to own the ML Infrastructure that powers how Voxel trains and ships...  ...Prefect, or similar) Familiarity with GPU performance profiling and optimization... 
    Work at office
    Flexible hours

    Voxel

    San Francisco, CA
    3 days ago
  •  ...Exa Infrastructure Engineer Exa is an applied AI lab building a search engine unlike the world has...  ...org. That could mean building GPU cluster orchestration in Kubernetes, map...  ...scales for our agent fleet Automate software maintenance and improvements for the whole... 
    H1b

    Exa Labs

    San Francisco, CA
    4 days ago
  • $176k - $220k

     ...institutions Work together with engineers, scientists, operators, and...  ...AI Human data is the core infrastructure to AI advancement. Frontier...  ...We’re looking for a Senior Software Engineer to join our ML Infrastructure...  ...fine-tuned models, including GPU serving, batching, and... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Handshake

    San Francisco, CA
    a month ago
  • About the TeamThe GPT Infrastructure team builds systems that turn advances...  ...and runtimes, performance engineering, secure partner integrations...  ...the RoleWe are seeking a software engineer to help build the platform...  ...platforms.Familiarity with GPU or accelerator architecture,... 
    Remote work

    OpenAI

    San Francisco, CA
    2 days ago
  • $200k - $250k

     ...Fluidstack, a leading cloud provider, is looking for a Software Engineer, Infrastructure Platform to build the foundational platforms that enable our...  ...data Create DCIM platforms for rack operations, server/GPU deployment, OS installation, quality assurance, and white-... 
    Full time
    Local area

    Fluidstack

    San Francisco, CA
    1 day ago
  • $180k - $250k

     ...next generation of AI products. We build the infrastructure, tools, and model access that teams need to...  ...ambitious teams build on. You are a hands-on engineer who builds the software and processes that keep a large fleet of GPU servers healthy and productive. You write... 
    Local area
    Relocation package

    features and labels

    San Francisco, CA
    1 day ago
  • $148.7k - $201.2k

     ...disciplinary team of scientists, engineers, and technicians, on a...  ...computer.We are looking to hire an HPC Platform Engineer to develop,...  ...-performance computing (HPC) infrastructure on AWS that CQC scientists...  ...maintaining clusters, managing OS and software stacks, building containers,... 
    Local area
    Flexible hours

    Amazon

    San Francisco, CA
    3 days ago
  •  ...Whoever deploys frontier compute infrastructure fastest will decide whether...  ...teams spanning hardware and software. Speed and scale are our key...  ...forward. The Production Engineering Team Examples of key...  ...0s of GWs: at our scale, a GPU failure isn't a ticket. It's... 
    Local area

    Fluidstack

    San Francisco, CA
    4 days ago
  • $200k - $250k

     ...office. The opportunity We're hiring a Senior Cloud Infrastructure Engineer to join the Cloud Infrastructure & Platform (CIN) team at Altruist...  .... Architect infrastructure for AI/ML workloads: GPU-enabled compute, SageMaker endpoints, Bedrock integration, vector... 
    Work at office
    Immediate start
    3 days per week

    Altruist

    San Francisco, CA
    1 day ago
  •  ...Whoever deploys frontier compute infrastructure fastest will decide whether...  ...teams spanning hardware and software. Speed and scale are our key...  ...forward. The Production Engineering Team Examples of key...  ...platforms, so every new site and GPU generation lands cleanly... 
    Local area

    Fluidstack

    San Francisco, CA
    4 days ago
  •  ...Sciforium is an AI infrastructure company developing next-...  ...hands-on support from AMD engineers the team is scaling...  ...runtime, service, and GPU layers, working closely...  ...) ~3+ years of software engineering experience...  ...to open-source ML or HPC infrastructure Benefits... 
    Full time
    Work at office
    Flexible hours

    Sciforium

    San Francisco, CA
    1 day ago
  • $209k - $240k

     ...office workdays. About the Product Infrastructure Team: The Product Infrastructure...  ...classes of problems up-front for product engineers. Solve hard technical challenges such...  ...values, and enthusiastic about making software toolmaking ubiquitous, we want to hear... 
    Full time
    Work at office
    Local area

    Notion

    San Francisco, CA
    1 day ago
  • $170k - $216k

     ...across 15+ U.S. states. The Simulation Infrastructure team creates reliable, scalable, and...  ...products that evaluate the Waymo Driver's software stack at a massive scale. We solve...  ...for a broad range of customers Software Engineers, Product, Data Science, System Engineering... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer, GPU Infrastructure - HPC. Be the first to apply!