Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer, Reliability

Full-time

OpenAI

Join the engineering teams that bring OpenAI’s ideas safely to the world!!

The Applied Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth.

About the Role

As OpenAI continues to grow, we are looking for experienced, problem-solving engineers to ensure our systems scale. Our success depends on our ability to quickly iterate on products while also ensuring that they are performant and reliable. You will work in a deeply iterative, collaborative, fast-paced environment to bring our technology to millions of users around the world, and ensure it’s delivered with safety and reliability in mind. Successful candidates will play a crucial role in ensuring the reliability, scalability, and performance of our systems as we continue to expand. As a reliability expert, you will be at the forefront of maintaining and enhancing the stability, scalability, and performance of our rapidly evolving infrastructure. You will work closely with cross-functional teams, including software engineers, product managers, and data scientists, to build and maintain resilient systems that can handle our growing user base and workload.

In this role, you will:

  • Design and implement solutions to ensure the scalability of our infrastructure to meet rapidly increasing demands.

  • Build and maintain the load, chaos and synthetic testing software leveraged by development teams to make the systems they design and operate more reliable.

  • Build and maintain automation tools to streamline repetitive tasks and improve system reliability.

  • Build and maintain the platform for CPU/storage, GPU, and network lifecycle management to drive efficiency, accountability and support dynamic optimization of our resources.

  • Implement fault-tolerant and resilient design patterns to minimize service disruptions.

  • Develop and maintain service level objectives (SLOs) and service level indicators (SLIs) to measure and ensure system reliability.

  • Partner with researchers, engineers, product managers, and designers to bring new features and research capabilities to the world.

  • Participate in an on-call rotation to respond to critical incidents and ensure 24/7 system availability.

You might thrive in this role if you:

  • Have a track record of accelerating engineering reliability by empowering your fellow engineers with excellent tooling and systems.

  • Have a humble attitude, an eagerness to help your colleagues, and a desire to do whatever it takes to make the team succeed.

  • Own problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done.

  • Enjoy seeking out and addressing bottlenecks and areas for performance improvement in our systems.

  • Utilize Infrastructure as Code (IaC) principles to automate infrastructure provisioning and configuration management.

  • Are experienced in collaborating with cross-functional teams to ensure that reliability and scalability are considered in the design and development of new features and services.

Qualifications:

  • Bachelor's degree in Computer Science, Information Technology, or a related field (or equivalent work experience).

  • Proven experience as an SWE focused on reliability or a similar role in a fast-paced, rapidly scaling company.

  • Strong proficiency in cloud infrastructure.

  • Proficiency in programming languages.

  • Experience with containerization technologies and container orchestration platforms like Kubernetes.

  • Knowledge of IaC tools such as Terraform or CloudFormation.

  • Excellent problem-solving and troubleshooting skills.

  • Strong communication and collaboration skills.

  • Experience with observability tools such as DataDog, Prometheus, Grafana and Splunk.

  • Experience with microservices architecture and service mesh technologies.

  • Knowledge of security best practices in cloud environments.

This role is exclusively based in our San Francisco HQ. We offer relocation assistance to new employees.

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. 

We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.

For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement .

Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.

To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form . No response will be provided to inquiries unrelated to job posting compliance.

We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link .

At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Software Engineer, Reliability in Remote vacancy
  • $125k - $145k

     ...developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SOFTWARE ENGINEER (FLIGHT RELIAIBLITY) The Flight Reliability software team creates mission critical applications that are used throughout SpaceX to accelerate... 
    Suggested
    Permanent employment
    Full time
    Temporary work
    Remote work
    Worldwide
    Weekend work

    Spacex

    Remote
    1 day ago
  • $150k - $176k

     ...Checkr is recognized on Forbes Cloud 100 2025 List and is a Y Combinator 2024 Breakthrough Company . As a Software Engineer II on the Site Reliability Engineering team within the Platform Engineering group at Checkr, you will identify reliability challenges impacting... 
    Suggested
    Full time
    Work at office
    Local area
    Remote work
    Relocation
    Flexible hours
    3 days per week

    Checkr

    San Francisco, CA
    1 day ago
  • $170k - $216k

     ...U.S. states. The Planner/Perception Reliability team builds out architectures, tools, and...  ...reliability and is accountable for onboard software health while ensuring high development...  ...you will report to a Staff Software Engineer / Tech Lead Manager. You will: Architect... 
    Suggested
    Full time
    Immediate start
    Remote work

    Waymo

    Remote
    1 day ago
  • $200k - $300k

     ...abundance for all. About the Team Our team owns the reliability and testing plan for every software system that runs on the robot or talks to it. That...  ...customers can trust their robots, and whether our engineers can move fast, comes down to this layer. Key... 
    Suggested
    Full time
    Temporary work
    Local area
    Work from home
    Flexible hours

    1x

    Remote
    1 day ago
  • $200k - $300k

     ...Hudson River Trading (HRT) is seeking a Software Engineer focused on GPU reliability to join our Systems Development team. The Systems Development team builds and maintains the platform that is shared by all Systems teams to provision, monitor, and manage HRT’s server... 
    Suggested
    Full time
    Work at office
    Local area
    Immediate start

    Hudson River Trading

    Remote
    1 day ago
  • The roleWe’re looking to hire our first data software engineer at Gridmatic! Looking for a startup-minded eng who works closely with our ML...  ...area to both make sure the data is ingested and transformed reliably, and also be able to build the tooling/abstractions to make... 
    Work at office
    Home office
    Flexible hours
    3 days per week

    Gridmatic

    Cupertino, CA
    3 days ago
  • Role Description We are looking for a Site Reliability Engineer with a strong software development background to join our Scrum team maintaining and improving an Enterprise Generative AI platform. This role focuses on platform stability, performance optimization, uptime... 
    Full time
    Remote work
    Flexible hours

    Corning

    Remote
    4 days ago
  • $152k - $241.5k

     ...of artificial intelligence.We are looking for highly motivated Senior Software Engineers to join our Fabric Networking team with a targeted focus on NVLink Rack-Scale Systems Stability & Reliability. In this role, you will partner closely with architects and developers... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...online without adding hardware, installing software, or changing a line of code. Internet...  ...aspect of Cloudflare's systems: enable more reliable network connectivity for Cloudflare’s...  .... You will work closely with various Engineering teams to translate their requirements... 
    Full time
    Local area

    Cloudflare

    Oklahoma
    1 day ago
  • $140k - $230k

     ...Zoox is seeking a Site Reliability Engineer to help ensure the availability, performance, and resilience of the services that power the development...  ...across engineering: You will partner closely with software engineering teams to elevate our system architecture, streamline... 
    Full time

    Zoox

    Remote
    1 day ago
  • $175k - $250k

     ...customized and developed by our expert team of lawyers, engineers and research scientists. We’ve found product market...  ...Competitive compensation. Role Overview As a Software Engineer on the Site Reliability team at Harvey, you will ensure the reliability, scalability... 
    Full time
    Relocation package

    Harvey

    Remote
    1 day ago
  • The Software Reliability Engineer (SRE) will play a critical role in ensuring that our Warehouse Management Software (WMS) runs seamlessly across both automated and manual facilities. This role focuses on investigating, diagnosing, and resolving operational software issues... 
    Full time
    Local area
    Remote work
    Rotating shift

    Lineage Logistics

    Novi, MI
    4 days ago
  •  ...cloud infrastructure (AWS/GCP/Azure) for performance, cost, and reliability. Improve observability across the platform through...  ...and integrations across systems and tools. Collaborate with engineering, operations, and data teams to understand and support their infrastructure... 
    Full time

    Qdrant

    Remote
    22 days ago
  •  ...for enterprise-grade AI. Founded by the engineers behind Milvus, the world’s most popular...  ..., and far higher expectations for reliability. You'll join a small, fast-moving Cloud...  ...Bachelor's degree in Computer Science, Software Engineering, or a related field, or equivalent... 
    Night shift
    Afternoon shift
    Early shift

    Zilliz

    Redwood City, CA
    a month ago
  • $120k - $190k

     ...Job Description Job Description Senior Software Engineer, Reliability Remote, US About Nametag Nametag is building the future of secure digital identity. Our mission is to make it easy for people and organizations to prove who they are online, safely and seamlessly... 
    Full time
    Remote work
    Visa sponsorship
    Flexible hours

    Nametag

    San Francisco, CA
    17 days ago
  • $142.7k - $158.3k

     ...Responsibilities for this Position Senior Software / Site Reliability Lead Engineer ID: 2026-74192 US-Telework-Telework Required Clearance: Secret, obtainable within reasonable time based on requirements Posted Date: 8/10/2026 Category: Engineering-... 
    Full time
    Remote work
    Flexible hours

    GD Mission Systems

    Sudbury, MA
    2 days ago
  • $174k - $253k

     ...Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical...  ...experience. 5 years of experience with software development in one or more programming...  ...or Engineering. ABOUT THE JOB: Site Reliability Engineering (SRE) is what you get when... 

    Socket

    Sunnyvale, CA
    4 days ago
  • $149.4k - $202k

     ...Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing... 
    Remote work

    Noctua Technology

    Washington DC
    1 day ago
  • $142.7k - $158.3k

     ...Description This position involves owning the reliability of AI services and ensuring that they...  ...error budgets and use them to drive engineering decisions. ~Monitoring and...  ...maintaining consistent standards. ~Your software engineering background allows you to engage... 
    Full time
    Remote work

    General Dynamics Mission Systems, Inc

    Remote
    3 days ago
  • $130k - $165k

     ...Job Title: Senior Software Engineer Company: Snapsheet Job Location: USA, Remote Job Type: Full-time, direct hire Job Department: Technology  Team : Site Reliability Engineering   About Snapsheet: Snapsheet exists to simplify claims. We leverage... 
    Full time
    Temporary work
    Local area
    Remote work
    Visa sponsorship
    Work visa
    Flexible hours

    Snapsheet

    Chicago, IL
    more than 2 months ago
  •  ...eat. We provide tools, resources and support to enable users to reach their health goals. We are looking for a Software Engineer III - Site Reliability to join the MyFitnessPal PEAS team. As a member of the PEAS team, you'll have the opportunity to positively impact... 
    Full time

    MyFitnessPal

    Remote
    18 days ago
  • $193.8k - $285k

    About the TeamThe Reliability Platform role is a key pillar of DoorDash’s Production Lifecycle...  ...toil and repetitive tasks. We use software and agents to “keep the lights on” and focus...  ...itself amazing!About the RoleAs a Software Engineer on the Reliability Platform team, you’ll... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    2 days ago
  •  ...at scale. As a global leader in enterprise recruitment software, SmartRecruiters offers a cloud-based global Hiring...  ...Description We are looking for a passionate Software Engineer with focus on Database Reliability to join our team of Software and DevOps Engineers and... 
    Full time
    Work at office
    Remote work
    Worldwide

    Dev

    Remote
    1 day ago
  •  ...to creating category-leading enterprise software that unleashes that power.To make that...  ...that be you?Your missionAt UiPath’s Site Reliability team, we build the platforms and...  ...repair item tracking, and assertion of engineering best practices across UiPath. We are scaling... 
    Work at office
    Immediate start
    Remote work

    UiPath

    Bellevue, WA
    1 day ago
  •  ...realize their greatest potential. Title and Summary Lead Site Reliability Engineer Job Description Summary Overview: Who is Mastercard...  ...our developers during the application build phase in software run principals that includes operational design, automation,... 
    Full time
    Part time
    Worldwide
    Flexible hours

    Mastercard

    O Fallon, MO
    7 hours ago
  • $120k - $193.7k

    onsemi is looking for a self-driven reliability leader to work in a fast-paced environment and be a key team member to accelerate adoption...  ...for assistance. Bachelor’s or Master’s degree in Electrical Engineering, Semiconductor Device Physics, Materials Science, or relevant... 
    Local area

    Onsemi

    San Jose, CA
    7 hours ago
  • $122k - $207k

     ...their greatest potential. Title and Summary Lead Site Reliability Engineer Overview The Mastercard Business Operations (BizOps)...  ...standards across teams, enabling consistent, high‑quality software delivery at scale. Observability & Self‑Healing Systems... 
    Full time
    Part time
    Worldwide
    Flexible hours
    Shift work

    Mastercard

    O Fallon, MO
    7 hours ago
  •  ...that we serve.The Information Technology group delivers secure, reliable technology solutions that enable DTCC to be the trusted...  ...Impact You Will Have in This RoleAs a Senior Application Support Engineer, you will help power DTCC's global financial markets infrastructure... 
    Remote work
    Flexible hours

    DTCC- The Depository Trust & Clearing Corporation

    Boston, MA
    2 days ago
  •  ...want to be part of an inclusive, adaptable, and forward-thinking organization, apply now.We are currently seeking a Platform Engineer (Reliability & Security) - Oscar Hernandez to join our team in Plano, Texas (US-TX), United States (US).Role 4: Platform Engineer (... 
    Work at office
    Remote work
    Flexible hours

    NTT DATA

    Plano, TX
    3 days ago
  • $136k - $224.25k

    NVIDIA is looking for a Senior Network Reliability Engineer to support and maintain our cloud and datacenter network infrastructures. This network serves the needs across the whole software stack for NVIDIA, from Graphics Drivers to Autonomous Vehicles and Artificial Intelligence... 
    Full time
    Remote work
    Shift work

    Nvidia

    Santa Clara, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer, Reliability. Be the first to apply!