Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer, GPU Infrastructure - HPC

Full-time

OpenAI

About the team

The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth.

About the role

As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise.

Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure.

Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale.

In this role, you will:

  • Build and maintain automation systems for provisioning and managing server fleets.

  • Develop tools to monitor server health, performance, and lifecycle events.

  • Collaborate with clusters, networking, and infrastructure teams.

  • Partner with external operators to ensure a high level of quality.

  • Identify and fix performance bottlenecks and inefficiencies.

  • Continuously improve automation to reduce manual work.

You might thrive in this role if you have:

  • Experience managing large-scale server environments.

  • A balance of strengths in building and operationalizing.

  • Proficiency in Python, Go, or similar languages.

  • Strong Linux, networking, and server hardware knowledge.

  • Comfort digging into noisy data with SQL, PromQL, and Pandas or any other tool.

Prior hardware expertise is not required for this role.

Bonus Skills:

  • Experience with low level details of hardware components, protocols, and associated Linux tooling (e.g., PCIe, Infiniband, networking, power management, kernel perf tuning)

  • Knowledge of hardware management protocols (e.g., IPMI, Redfish).

  • High-performance computing (HPC) or distributed systems experience.

  • Prior experience developing, managing, or designing hardware.

  • Familiarity with monitoring tools (e.g., Prometheus, Grafana).

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. 

We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.

For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement .

Qualified applicants with arrest or conviction records will be considered for employment in accordance with applicable law, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.

We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link .

At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.

Vacancy posted 8 hours ago
Similar jobs that could be interesting for youBased on the Software Engineer, GPU Infrastructure - HPC in San Francisco, CA vacancy
  •  ...This role will support the fleet infrastructure team at OpenAI. The fleet team focuses on running...  ..., most reliable, and frictionless GPU fleet to support OpenAI’s general purpose...  ...Much more! About the Role As an engineer within Fleet infrastructure, you will design... 
    Suggested
    Full time
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    8 hours ago
  •  ...applied AI research, flexible infrastructure, and seamless developer...  ...and help build the platform engineers turn to to ship AI products....  ...foundational engineers to lead our GPU Networking efforts, making RDMA...  ...to architect the software fabric that unifies thousands... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    8 hours ago
  • $180k - $250k

     ...You are a hands-on engineer who builds the software and processes that keep a large fleet of GPU servers healthy and productive. You write systems and tooling for managing...  ...with configuration management and infrastructure-as-code: Ansible, Terraform, cloud-init ~... 
    Suggested
    Remote job
    Full time
    Local area
    Relocation package

    Falò

    San Francisco, CA
    8 hours ago
  •  ...Exa is building a search engine from scratch to serve every AI agent. We build massive-scale infrastructure to crawl the web, train state-of...  ...compute, we also own a $5M H200 GPU cluster (and soon 5x'ing that...  ...Design GPU scheduling software so we max out our cluster utilization... 
    Suggested
    Full time
    H1b

    Exa

    San Francisco, CA
    8 hours ago
  • $215k - $265k

     ...scheming mitigations. We're looking for a Software Engineer to build the platform that the rest of...  ...Build and maintain Apollo's cloud infrastructure . This means IaC, networking,...  ...signals for their service. Multi-cloud GPU orchestration: The platform that finds... 
    Suggested
    Full time
    Work experience placement
    Work at office
    Immediate start
    Visa sponsorship
    Flexible hours

    Apollo Research

    San Francisco, CA
    8 hours ago
  •  ...About the Role This role broadly owns infrastructure across the stack. If it’s running in the...  ...responsible for: The compute (both CPU and GPU compute) that’s required to cost...  ...GPT-4 to hundreds of millions of users, engineered the foundations of autonomous driving, built... 
    Full time

    The Generalist

    San Francisco, CA
    8 hours ago
  •  ...working systems and build any software needed for running large-...  ...the Role We are looking for engineers to operate the next generation...  ...systems engineering with hands-on infrastructure work on our largest...  ...bare-metal Linux environments, GPU hardware, and large-scale networking... 
    Full time

    OpenAI

    San Francisco, CA
    8 hours ago
  •  ...Background Specter is creating a software-defined "control plane" for...  ...become the perception engine for a company's physical footprint...  ...Specter is hiring an ML infrastructure engineer to build and scale the...  ...DDP, DeepSpeed, Ray) and GPU cluster management. Strong... 
    Full time

    S.e. Specter

    San Francisco, CA
    8 hours ago
  •  ...performance, flexibility, and resiliency across our infrastructure. We are forming a team to generalize our...  ...AMD. About the Role We’re hiring engineers to scale and optimize OpenAI’s inference infrastructure across emerging GPU platforms. You’ll work across the stack -... 
    Full time

    OpenAI

    San Francisco, CA
    8 hours ago
  •  ...applied AI research, flexible infrastructure, and seamless developer...  ...and help build the platform engineers turn to to ship AI products....  ...Baseten is building its own GPU infrastructure for large-scale...  ...not. We are hiring a Lead Software Engineer to build a first-class... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    8 hours ago
  •  ...Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable...  .... Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re seeking a GPU Kernel Engineer to join our team at the cutting... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    8 hours ago
  •  ...truly understand each other. The Role We're hiring a Senior Software Engineer (Infrastructure) to be a technical driver across teams. You're the...  ...team. Your Mission Work across teams to own and extend our GPU infra as well as our traditional cloud infra (AWS) Work closely... 
    Flexible hours

    Tavus

    San Francisco, CA
    2 days ago
  •  ...will de-risk the largest infrastructure build-out in history. When people finance GPU clusters, the...  ...looking for a high agency engineer to help build the compute...  ...with the orchestration software managing virtual machines...  ...running on cutting-edge HPC hardware worth billions... 
    Long term contract
    Full time
    Contract work
    Fixed term contract
    Work at office
    Local area
    Visa sponsorship
    Shift work

    The San Francisco Compute Company

    San Francisco, CA
    8 hours ago
  • $160k - $230k

    Senior Software Engineer - Together Cloud Infrastructure Together AI is building the AI Acceleration Cloud, an end-to-end platform for the full generative AI...  ...at scale a plus Experience with DPUs/SmartNICs a plus GPU programming, NCCL, CUDA knowledge a plus Responsibilities... 
    Full time
    Work at office
    Remote work

    Together AI

    San Francisco, CA
    3 hours ago
  • $200k - $250k

     ...Fluidstack, a leading cloud provider, is looking for a Software Engineer, Infrastructure Platform to build the foundational platforms that enable our...  ...data Create DCIM platforms for rack operations, server/GPU deployment, OS installation, quality assurance, and white-... 
    Full time
    Local area

    Fluidstack

    San Francisco, CA
    8 hours ago
  •  ..., and Google Lens. Before that, Clay led the product and design teams for Google Workspace.  What you’ll do As a Software Engineer, Infrastructure at Sierra, you will be responsible for designing, building, and maintaining the core systems that make our AI platform... 
    Full time
    Flexible hours

    Sierra

    San Francisco, CA
    24 days ago
  •  ...teams for Google Workspace.  What you'll do The Payments Infrastructure team builds the trust boundary between a live conversation...  ...plaintext cardholder data. Make payments something other engineers can use without becoming compliance experts: drive the platform... 
    Full time
    Flexible hours

    Sierra

    San Francisco, CA
    24 days ago
  • $200k - $400k

     ...and grow as a team. About the Team The Infrastructure team builds and operates the foundations...  ...managed and on‑prem environments. ML Infra: GPU and model‑serving platforms for LLM...  ...’re hiring a Senior Data Infrastructure Engineer to design, build, and operate the data systems... 
    Full time
    Work at office
    Local area

    Decagon

    San Francisco, CA
    2 days ago
  •  ...Sciforium is an AI infrastructure company developing next-...  ...hands-on support from AMD engineers the team is scaling...  ...runtime, service, and GPU layers, working closely...  ...) ~3+ years of software engineering experience...  ...to open-source ML or HPC infrastructure Benefits... 
    Full time
    Work at office
    Flexible hours

    Sciforium

    San Francisco, CA
    8 hours ago
  •  ...Orb: Orb is redefining what billing software can be, turning one of the most complex...  ...the role: As a senior member of the infrastructure team, you’ll play a key role in maintaining...  ...customer growth Partner with other engineering teams to ensure we build reliable... 
    Full time
    Work at office
    Remote work
    3 days per week

    Orb

    San Francisco, CA
    8 hours ago
  • $130k - $175k

     ...goal is to enable and empower Kiddom’s engineering by building a scalable and sustainable...  ...new and existing services. Practicing Infrastructure as Code (IaC) wherever possible, giving...  ...related field ~5+ years professional software engineering experience ~ Experience... 
    Permanent employment
    Full time
    Work at office
    Local area
    Remote work

    Kiddom

    San Francisco, CA
    8 hours ago
  • $145.4k - $188.1k

     ...together. About the Machine Learning Infrastructure Team At Thumbtack, our challenges...  ...these experiences. We empower product engineering teams by providing scalable, high-performance...  ...blog . The challenge As a Software Engineer on the ML Infrastructure team,... 
    Full time
    Seasonal work
    Local area

    Thumbtack

    San Francisco, CA
    8 hours ago
  •  ...About the Team The ChatGPT team works across research, engineering, product, and design to bring OpenAI’s technology to the world. We seek to learn from deployment and broadly distribute the benefits of AI, while ensuring that this powerful tool is used responsibly... 
    Full time
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    8 hours ago
  • $170k - $216k

     ...across 15+ U.S. states. The Simulation Infrastructure team creates reliable, scalable, and...  ...products that evaluate the Waymo Driver's software stack at a massive scale. We solve...  ...for a broad range of customers Software Engineers, Product, Data Science, System Engineering... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    8 hours ago
  •  ...economy has grown 6x over the last decade, software businesses have gone from not worrying...  .... We're building the agentic infrastructure layer at Anrok—the systems that let AI...  ...product builds on. We're looking for an engineer who pairs a strong distributed-systems... 
    Full time
    Work at office
    Remote work
    Home office
    Flexible hours
    3 days per week

    Anrok, Inc.

    San Francisco, CA
    8 hours ago
  •  ...the data collection systems, operational capability, exabyte-scale data warehouse, and software toolchain – to help our partners drive the field forward. Infrastructure engineers build the platform that support our ever-increasing scale of data collection. Sample... 
    Full time

    Xdof

    San Francisco, CA
    8 hours ago
  •  ...accelerating project schedules of billion-dollar infrastructure projects and improving safety on job...  ...construction veterans and world-class engineers to solve physical-world problems that...  ...Onboard Infrastructure team builds the software foundation Bedrock's autonomous... 
    Full time
    Work at office
    Flexible hours

    Bedrock Robotics

    San Francisco, CA
    8 hours ago
  •  ...including stablecoins. You’ll help design, deploy and operate the infrastructure that makes this possible. This is a hands-on devops...  ...role sits at the intersection of infrastructure, platform engineering and security. You’ll define how Modern Treasury scales and operates... 
    Full time
    Work at office
    Local area
    Immediate start

    Modern Treasury

    San Francisco, CA
    8 hours ago
  • $120k - $170k

     ...Horowitz to Blackrock and Fidelity, and employs a team of 450 engineers and entrepreneurs. Astranis designs, builds, and...  ...,000 sq. ft. headquarters in Northern California, USA. Software Engineer - Infrastructure As a Software Engineer on the Infrastructure team,... 
    Permanent employment
    Full time
    Remote work
    Flexible hours

    Astranis

    San Francisco, CA
    8 hours ago
  •  ...systems fast, reliable, and cost-efficient requires world-class infrastructure. The Caching Infrastructure team is responsible for...  ...diverse range of use cases. We’re looking for an experienced engineer to help design and scale this critical infrastructure. The ideal... 
    Full time

    OpenAI

    San Francisco, CA
    8 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer, GPU Infrastructure - HPC. Be the first to apply!