Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer, GPU Infrastructure - HPC

$230k

OpenAI

About the teamThe Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth.About the roleAs a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise.Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure.Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale.In this role, you will:Build and maintain automation systems for provisioning and managing server fleets.Develop tools to monitor server health, performance, and lifecycle events.Collaborate with clusters, networking, and infrastructure teams.Partner with external operators to ensure a high level of quality.Identify and fix performance bottlenecks and inefficiencies.Continuously improve automation to reduce manual work.You might thrive in this role if you have:Experience managing large-scale server environments.A balance of strengths in building and operationalizing.Proficiency in Python, Go, or similar languages.Strong Linux, networking, and server hardware knowledge.Comfort digging into noisy data with SQL, PromQL, and Pandas or any other tool.Prior hardware expertise is not required for this role.Bonus Skills:Experience with low level details of hardware components, protocols, and associated Linux tooling (e.g., PCIe, Infiniband, networking, power management, kernel perf tuning)Knowledge of hardware management protocols (e.g., IPMI, Redfish).High-performance computing (HPC) or distributed systems experience.Prior experience developing, managing, or designing hardware.Familiarity with monitoring tools (e.g., Prometheus, Grafana).About OpenAIOpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement.Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance.We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link.OpenAI Global Applicant Privacy PolicyAt OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.Compensation Range: $230K - $490KLocationSan Francisco; New York CityEmployment TypeFull timeDepartmentScalingCompensation$230K – $490K • Offers EquityThe base pay offered may vary depending on multiple individualized factors, including market location, job-related knowledge, skills, and experience. If the role is non-exempt, overtime pay will be provided consistent with applicable laws. In addition to the salary range listed above, total compensation also includes generous equity, performance-related bonus(es) for eligible employees, and the following benefits.Medical, dental, and vision insurance for you and your family, with employer contributions to Health Savings AccountsPre-tax accounts for Health FSA, Dependent Care FSA, and commuter expenses (parking and transit)401(k) retirement plan with employer matchPaid parental leave (up to 24 weeks for birth parents and 20 weeks for non-birthing parents), plus paid medical and caregiver leave (up to 8 weeks)Paid time off: flexible PTO for exempt employees and up to 15 days annually for non-exempt employees13+ paid company holidays, and multiple paid coordinated company office closures throughout the year for focus and recharge, plus paid sick or safe time (1 hour per 30 hours worked, or more, as required by applicable state or local law)Mental health and wellness supportEmployer-paid basic life and disability coverageAnnual learning and development stipend to fuel your professional growthDaily meals in our offices, and meal delivery credits as eligibleRelocation support for eligible employeesAdditional taxable fringe benefits, such as charitable donation matching and wellness stipends, may also be provided.More details about our benefits are available to candidates during the hiring process.This role is at-will and OpenAI reserves the right to modify base pay and other compensation components at any time based on individual performance, team or company results, or market conditions.

Vacancy posted 6 days ago
Similar jobs that could be interesting for youBased on the Software Engineer, GPU Infrastructure - HPC in San Francisco, CA vacancy
  •  ...Whoever deploys frontier compute infrastructure fastest will decide whether...  ...teams spanning hardware and software. Speed and scale are our key...  ...world. The Production Engineering Team Examples of key...  ...0s of GWs: at our scale, a GPU failure isn't a ticket. It's... 
    Suggested
    Full time
    Local area

    Fluidstack

    San Francisco, CA
    15 hours ago
  • $240k - $280k

     ...About the Role We're looking for a Software Engineer to build the systems that treat infrastructure as software. This role owns the software state machines that...  ...physical host from discovery, inference bring-up to GPU driver/CUDA stack, health validation, and decommission... 
    Suggested
    Full time

    Together Ai

    San Francisco, CA
    8 hours ago
  •  ...applied AI research, flexible infrastructure, and seamless developer...  ...and help build the platform engineers turn to to ship AI products....  ...foundational engineers to lead our GPU Networking efforts, making RDMA...  ...to architect the software fabric that unifies thousands... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    15 hours ago
  • $300k

     ...Join a stealth-mode hyperscale infrastructure startup building a 300MW+ AI...  .... The Principal Software Engineer will take ownership of the software...  ..., workload scheduling, GPU resource management, infrastructure...  ...powering large-scale HPC and AI infrastructure environments... 
    Suggested
    Full time
    Remote work
    Flexible hours
    San Francisco, CA
    a month ago
  •  ...Background Specter is creating a software-defined "control plane" for...  ...become the perception engine for a company's physical footprint...  ...Specter is hiring an ML infrastructure engineer to build and scale the...  ...DDP, DeepSpeed, Ray) and GPU cluster management. Strong... 
    Suggested
    Full time

    S.e. Specter

    San Francisco, CA
    15 hours ago
  • $180k - $250k

     ...next generation of AI products. We build the infrastructure, tools, and model access that teams need to...  ...ambitious teams build on. You are a hands-on engineer who builds the software and processes that keep a large fleet of GPU servers healthy and productive. You write... 
    Full time
    Local area
    Relocation package

    Falò

    San Francisco, CA
    15 hours ago
  • $137.1k - $201.6k

     ...Platform and builds the shared infrastructure that helps DoorDash, Wolt,...  ...DeepSeek) ourselves — real-time GPU serving, high-throughput...  ...model serving and inference engines, fine-tuning and training pipelines...  ...of industry experience in software engineering ~ Deep backend... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    8 hours ago
  •  ...You’ll Do Design and operate secure, multi‑tenant container infrastructure with fast startup and smart autoscaling. Ship cloud...  ...systems. Nice to Have gVisor/Kata/Firecracker; Cilium/eBPF; GPU scheduling; serverless autoscaling (KEDA/Knative/Karpenter).... 
    Full time
    Remote work

    Julius Ai

    San Francisco, CA
    15 hours ago
  •  ...Exa is building a search engine from scratch to serve every AI agent. We build massive-scale infrastructure to crawl the web, train state-of...  ...compute, we also own a $5M H200 GPU cluster (and soon 5x'ing that...  ...Design GPU scheduling software so we max out our cluster utilization... 
    Full time
    H1b

    Exa

    San Francisco, CA
    15 hours ago
  • $215k - $265k

     ...scheming mitigations. We're looking for a Software Engineer to build the platform that the rest of...  ...Build and maintain Apollo's cloud infrastructure . This means IaC, networking,...  ...signals for their service. Multi-cloud GPU orchestration: The platform that finds... 
    Full time
    Work experience placement
    Work at office
    Immediate start
    Visa sponsorship
    Flexible hours

    Apollo Research

    San Francisco, CA
    15 hours ago
  • $250k

     ...already in place.   Role Overview We’re looking for an ML infrastructure engineer to design and build the core systems that enable scalable,...  ...parallelism) and the systems concerns of keeping large GPU jobs efficient ~ Prior experience as a technical lead or mentor... 
    Full time

    Epsilon Labs, Inc.

    San Francisco, CA
    15 hours ago
  •  ...working systems and build any software needed for running large-...  ...the Role We are looking for engineers to operate the next generation...  ...systems engineering with hands-on infrastructure work on our largest...  ...bare-metal Linux environments, GPU hardware, and large-scale networking... 
    Full time

    OpenAI

    San Francisco, CA
    15 hours ago
  •  ...About the Role This role broadly owns infrastructure across the stack. If it’s running in the...  ...responsible for: The compute (both CPU and GPU compute) that’s required to cost...  ...GPT-4 to hundreds of millions of users, engineered the foundations of autonomous driving, built... 
    Full time

    The Generalist

    San Francisco, CA
    15 hours ago
  • $275k

     ...your career? Join a VC-backed GPU cloud company at the founding...  ...GPU compute and inference infrastructure. Backed by leading venture investors...  ...is seeking a Founding Engineer to take ownership of core GPU...  ...operating large-scale AI or HPC infrastructure in production... 
    Full time
    Relocation
    San Francisco, CA
    a month ago
  •  ...Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable...  .... Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re seeking a GPU Kernel Engineer to join our team at the cutting... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    15 hours ago
  •  ...applied AI research, flexible infrastructure, and seamless developer...  ...and help build the platform engineers turn to to ship AI products....  ...Baseten is building its own GPU infrastructure for large-scale...  ...not. We are hiring a Lead Software Engineer to build a first-class... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    15 hours ago
  •  ...well as accelerating research progression via model inference. About the Role We’re hiring engineers to scale and optimize OpenAI’s inference infrastructure across emerging GPU platforms. You’ll work across the stack - from low-level kernel performance to high-level... 
    Full time

    OpenAI

    San Francisco, CA
    15 hours ago
  •  ...About the job FriendliAI is looking for a Cloud Infrastructure Engineer to own the architecture and evolution of the cluster platform behind our GPU-accelerated AI inference cloud. As a Software Engineer, Cloud Infrastructure, you will design how our clusters are built... 
    Permanent employment
    Full time
    Flexible hours

    FriendliAI

    San Francisco, CA
    15 hours ago
  • $130.6k - $192k

     ...next 10B even better. About the Role We’re hiring an Infrastructure Software Engineer. In this role, you’ll work with multiple stakeholders...  ...processing, and organization of petabyte-scale datasets GPU-accelerated distributed computing for data preparation and... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    15 hours ago
  •  ...Software Engineer Voxel's perception system is the technical core of everything we ship. Our...  ...software engineer to own the ML Infrastructure that powers how Voxel trains and ships...  ...Prefect, or similar) Familiarity with GPU performance profiling and optimization... 
    Work at office
    Flexible hours

    Voxel

    San Francisco, CA
    15 hours ago
  •  ...Cloud, is a leader in AI cloud infrastructure serving tens of thousands of...  .... One person, one GPU. If you'd like to build the...  ...About the Role As a Senior Software Engineer on Lambda’s Core Cloud Platform...  ...metal, GPU infrastructure, HPC, Kubernetes, Slurm, or large... 
    Full time
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda

    San Francisco, CA
    15 hours ago
  • $176k - $220k

     ...institutions Work together with engineers, scientists, operators, and...  ...AI Human data is the core infrastructure to AI advancement. Frontier...  ...We’re looking for a Senior Software Engineer to join our ML Infrastructure...  ...fine-tuned models, including GPU serving, batching, and... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Handshake

    San Francisco, CA
    23 days ago
  •  ...Whoever deploys frontier compute infrastructure fastest will decide whether...  ...teams spanning hardware and software. Speed and scale are our key...  ...world. The Production Engineering Team Examples of key...  ...platforms, so every new site and GPU generation lands cleanly... 
    Local area

    Fluidstack

    San Francisco, CA
    1 day ago
  • $230k

    This role will support the fleet infrastructure team at OpenAI. The fleet team focuses on running...  ..., most reliable, and frictionless GPU fleet to support OpenAI’s general purpose...  ...hardware caching Much more!About the RoleAs an engineer within Fleet infrastructure, you will... 
    Work at office
    Local area
    Relocation package
    Flexible hours

    OpenAI

    San Francisco, CA
    6 days ago
  • $230k - $405k

     ...About the Team:Compute Infrastructure builds the platform that turns enormous...  ...of compute into a reliable engine for frontier AI. We design,...  ...data centers, orchestration software, agent infrastructure, developer...  ...protocols, RDMA, NCCL, GPU hardware behavior, benchmarking... 
    Work at office
    Local area
    Flexible hours

    OpenAI

    San Francisco, CA
    a month ago
  •  ...applied AI lab building a search engine unlike the world has ever...  ...this is the place for you. Our Infrastructure Team builds the underlying...  ...org. That could mean building GPU cluster orchestration in...  ...for our agent fleet Automate software maintenance and improvements... 
    H1b

    Exa Corporation

    San Francisco, CA
    5 days ago
  • $220k - $292k

    Software Engineer, Infrastructure Platform owns the foundation the company runs on: compute and deployment, reliability and on-call, CI/CD and monorepo...  ...'ll Own Own our platform: GCP, Kubernetes, Temporal, the GPU fleet behind cloud export, and the deploy and rollback machinery... 
    Work at office
    Remote work
    Flexible hours

    Descript

    San Francisco, CA
    5 days ago
  • $175k - $215k

     ...redefines the future of global mobility. The Waymo Logs Infrastructure powers critical decision making and model development...  ...and clients   You have: ~4+ years of professional software engineering experience ~ Experience working on large-scale distributed... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    15 hours ago
  •  ...business with billions in revenue The Role Handshake is building the infrastructure layer that powers the next generation of AI agents across our platform. As a Senior Software Engineer on our Agentic Infrastructure team, you'll be at forefront of AI at... 
    Full time
    Freelance
    Internship
    Work at office
    Remote work
    Flexible hours

    Handshake

    San Francisco, CA
    15 hours ago
  •  ...of the great threats AI presents: mass-manufactured social engineering. Countless scams, deepfakes, and other social engineering attacks...  ...for an experienced backend engineer to build out the infrastructure required to rapidly scale up our engineering organization. A... 
    Full time
    Work at office
    Flexible hours

    Doppel

    San Francisco, CA
    15 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer, GPU Infrastructure - HPC. Be the first to apply!