Software Engineer, Workload Enablement
OpenAI
About the Team
The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization.
About the Role
We’re hiring an SW Engineer to enable production workloads and end-to-end testing on new platforms. This role will include creating new test harnesses and platform stress benchmarks, porting existing inference and training workloads to new, sometimes early-access, systems/hardware, analyzing performance and bottlenecks, and characterizing the end-to-end behavior of new systems (compute, comms, storage, control plane, and failure modes).
Key Responsibilities
Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar.
Build a suite of benchmarks and stress tests that capture real E2E behavior of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts.
Deep-dive performance on distributed training/inference:
Collective performance and tuning (across NCCL/RCCL and internal libraries)
Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects
Create repeatable test harnesses that run in CI / lab environments and produce actionable outputs (pass/fail, performance score, regression detection).
Partner with systems + fleet bring-up engineers to ensure the platform is not only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops).
Work cross-functionally with vendors and internal stakeholders by producing clear bug reports, minimal repros, and prioritized issue lists.
Qualifications
BS in CS/EE (or equivalent practical experience).
5+ years in one or more of: ML systems, performance engineering, distributed systems, or HPC.
Strong hands-on experience with:
PyTorch and modern LLM training/inference stacks
Large-scale distributed training concepts (data/model/pipeline parallel, collective comms)
Experience with RDMA and debugging/optimizing comms libraries (NCCL or RCCL) and their interaction with hardware/network
Proficiency in Python plus comfort reading/writing performance-critical code (C++/CUDA/HIP is a plus).
Strong profiling/debugging skills (e.g., Nsight, rocprof, perf, flamegraphs; ability to reason from traces/counters).
Preferred Skills:
Experience building workload-shaped benchmarks and stress/fault tests that correlate to production behavior (not just synthetic loops or microbenchmarks).
Familiarity with RDMA networking and transport tuning; understanding of how network topology and congestion impact collectives.
Experience running and validating workloads in Kubernetes, and bridging “research code” into robust, repeatable infrastructure.
Hands-on lab experience with early hardware (new NICs, new GPUs/accelerators, early racks).
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity.
We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.
For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement .
Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.
To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form . No response will be provided to inquiries unrelated to job posting compliance.
We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link .
At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.
- ...About the Team The Agent Enablement team works across engineering, product, design, and research to bring our technology to the world. We seek... ...product improvements. Are familiar with enterprise software concepts such as identity, permissions, governance, compliance...SuggestedFull timeTemporary work
- ...hardware architectures like AMD. About the Role We’re hiring engineers to scale and optimize OpenAI’s inference infrastructure across... ...-backed systems. Debug and optimize distributed inference workloads across memory, network, and compute layers. Validate...SuggestedFull time
- ...research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge... ...Greylock, and Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As Baseten's AI...SuggestedFull timeFlexible hours
- ...infrastructure, and seamless developer tooling, we enable companies operating at the frontier of... .... Join us and help build the platform engineers turn to to ship AI products. THE... ...-tenant environments, and cross-cloud workload orchestration. Lead technical...SuggestedFull timeFlexible hours
$300k - $320k
...group of committed researchers, engineers, policy experts, and business... ...role We are looking for software engineers to join our... ...and shared libraries. Our work enables product teams to build and operate... ...Claude to specific customer workloads. The core problem is task-...SuggestedFull timeCurrently hiringWork at officeVisa sponsorshipFlexible hours- ...operates critical infrastructure that enables research at OpenAI. Our mission is simple... ...increasing complexity and size of our workloads, while remaining reliable and easy to... ...We’re looking for a staff-level software engineer to own production-critical infrastructure...Full timeWork at officeRelocation package
- ...computing and make it accessible to software developers of all skill... ...is looking for a Software Engineer to join the Infrastructure team... ...secure, and robust backbone that enables this vision. Our team is... ...execution of distributed workloads. We are seeking a talented...Full time
- ...a small, fast-moving team of engineers focused on delivering a world... ...other non-text modalities. These workloads are inherently more... ...Role We’re looking for a software engineer to help us serve OpenAI... ...audio inputs and outputs. Enable experimental research workflows...Full time
- ...threats AI presents: mass-manufactured social engineering. Countless scams, deepfakes, and other social... ...sophisticated LLM engine to reduce verification workloads by 4x for the highest volume product line, which will enable faster takedowns at scale Built a QA & data...Full timeWork at officeFlexible hours
- ...Kernels team at OpenAI builds the low-level software that accelerates our most ambitious AI... ...inference more efficient. Our work enables OpenAI to push the limits by ensuring models... ...performance optimizations across our AI workloads. You’ll work across the stack,...Full time
- ...seamless developer tooling, we enable companies operating at the... ...and help build the platform engineers turn to to ship AI products.... ...reliability, and ease of use. As a Software Engineer on the Inference... ..., networking, and GPU workloads Make thoughtful engineering...Full timeFlexible hours
$120k - $290k
...horizontal scaling of MySQL, enables businesses to efficiently handle large-scale data workloads — without sacrificing developer... ...role As a PostgreSQL Core Engineer, you will spend most of your... ...consumer tech, and B2B SaaS. As a Software Engineer, you'll be at the...Full timeWorldwide$170k - $235k
...highly optimized SQL queries, enabling seamless exploratory... ...Team, you will join a group of engineers dedicated to building the core... ...across a wide range of query workloads and data architectures Contribute... ...engineering high-quality software systems ~ Demonstrated success...Full timeWork at officeFlexible hours- ...: We’re looking for a foundational Software Engineer (Infrastructure) who thrives on solving... ...platform, and data engineering teams to enable rapid feature development without... ...~ Experience managing containerized workloads using tools like Kubernetes, ECS, or similar...Full timeWork at officeRemote work
- ...infrastructure, and seamless developer tooling, we enable companies operating at the frontier of... .... Join us and help build the platform engineers turn to to ship AI products. THE... ...the foundation that powers modern AI workloads, optimizing every microsecond of...Full timeFlexible hours
- ...We’re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams... ..., ChatGPT, API, and internal research workloads. We’re hiring a Developer Productivity Engineer... ..., and developer workflows that enable our teams to move quickly without compromising...Full time
$120k - $290k
...clustering system for horizontal scaling of MySQL, enables businesses to efficiently handle large-scale data workloads — without sacrificing developer experience.... ...PostgreSQL product, and we're looking for Software Engineers to come help build it from scratch. Our...Full timeWorldwide- ...seamless developer tooling, we enable companies operating at the... ...and help build the platform engineers turn to to ship AI products.... ...stalls, GPU driver problems, and workload symptoms that look like... ...not. We are hiring a Lead Software Engineer to build a first-class...Full timeFlexible hours
- ...to accelerate progress toward AGI by enabling the fastest iteration cycles and highest... ...at scale. About the Role As a software engineer on the Scaling team, you’ll help build... ...validate and optimize distributed training workloads. You will work at the intersection...Full timeWork at officeLocal areaRelocation package3 days per week
- ...vectorSearch aggregation, which enables approximate nearest neighbor... ...customer base, providing engineers an opportunity to make a highly... ...and disrupt industries with software. MongoDB’s unified database platform... ...modernize legacy workloads, embrace innovation, and unleash...Full timeLocal areaWorldwide
$300k
...group of committed researchers, engineers, policy experts, and business... ...customer growth, while enabling breakthrough research by giving... ...you: Have significant software engineering experience,... ..., research, and experimental workloads Building production-grade...Full timeWork at officeWorldwideVisa sponsorshipFlexible hours- ...distribute OpenAI’s API broadly and safely by enabling key API technologies in cloud-native... ...runtime environments for agentic workloads. This work sits at the intersection of production... ...the Role We’re looking for a backend engineer who can quickly understand OpenAI’s...Full timeInternship
- ...Join the engineering teams that bring OpenAI’s ideas safely to the... ...functional teams, including software engineers, product managers,... ...handle our growing user base and workload. This role requires a blend... ...all feel welcome while enabling radical candor and the challenging...Full timeWork experience placementRelocation package
$120k - $290k
...horizontal scaling of MySQL, enables businesses to efficiently handle large-scale data workloads — without sacrificing developer... ...PlanetScale is seeking an engineer to join our Insights team to help... ...Required: ~5+ years of software engineering experience ~ Experience...Full timeWorldwide$310k
...distributed systems. We build the engineering and research infrastructure... ...core model training software and works deep in the stack... ...OpenAI's research velocity, enabling reliable, efficient training... ...visibility into large-scale training workloads and help operate them...- ...The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced... ...operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and...Full time
- ...Position Overview: We are seeking a Software Engineer – Data Center Simulation with 15+ years... ...intuitive front-end visualization tools to enable data-driven decision-making.... ...computing (HPC) environments for simulation workloads. ~ Strong knowledge of software engineering...Full time
- ...seamless developer tooling, we enable companies operating at the... ...and help build the platform engineers turn to to ship AI products.... ...that as LLM and multi-modal workloads scale, the network is the computer... ...to architect the software fabric that unifies thousands...Full timeFlexible hours
- ...cloud-based distributed systems software responsible for the lifecycle... ...for even the largest workloads Work with a collaborative... ...and that we are proud of as engineers Have the opportunity to lead... ...the database for the AI era, enabling innovators to create, transform...Full timeWork at officeLocal areaWorldwide
- ...industries by unleashing the power of software and data. We enable organizations of all sizes to easily... ...by helping them modernize legacy workloads, embrace innovation, and unleash AI.... ...observable system for customers and engineers. The Atlas Search product is quickly...Full timeWork at officeLocal areaWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer, Workload Enablement. Be the first to apply!
- software engineer full time San Francisco, CA
- software system engineer San Francisco, CA
- consulting software engineer San Francisco, CA
- software engineer travel San Francisco, CA
- real time software engineer San Francisco, CA
- network software engineer San Francisco, CA
- senior software engineer remote San Francisco, CA
- entry level software engineer remote San Francisco, CA
- software engineer intern San Francisco, CA
- new grad software engineer San Francisco, CA


