Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Tech Lead, Deployment & Operations - Custom Infrastructure

$342k

OpenAI

About the Team OpenAI’s Hardware organization develops silicon and system‑level solutions designed for the unique demands of advanced AI workloads. The team is responsible for building the next generation of AI‑native silicon while working closely with software and research partners to co‑design hardware tightly integrated with AI models. In addition to delivering production‑grade silicon for OpenAI’s supercomputing infrastructure, the team also creates custom design tools and methodologies that accelerate innovation and enable hardware optimized specifically for AI. About the Role We are seeking a Technical Lead to lead deployment and operations for OpenAI’s Silicon & Systems team. This person will become the Directly‑Responsible Individual responsible for bringing OpenAI’s custom silicon and associated systems into data center environments, ensuring successful deployment, bring‑up, validation, operational readiness, and ongoing reliability at scale. This role sits at the intersection of silicon, systems, infrastructure, data center operations, and software. You will lead a team focused on taking new hardware platforms from lab validation into production data center deployment. You will be responsible for building the operational processes, technical workflows, tooling, and cross‑functional alignment required to deploy and operate custom AI hardware reliably in OpenAI’s supercomputing infrastructure. The ideal candidate is both a strong leader and a deeply technical operator. You should be comfortable staying close to the technical details of hardware bring‑up, fleet deployment, debugging, system validation, data center integration, and production operations. This role requires strong execution, excellent cross‑functional judgment, and the ability to drive clarity in ambiguous, fast‑moving environments. In this role, you will: Lead a team responsible for deployment and operations of OpenAI’s custom silicon and systems in data center environments Own the path from hardware bring‑up and validation through production deployment, operational readiness, and sustained fleet support Partner closely with silicon, systems, software, infrastructure, networking, data center, supply chain, and external partner teams to ensure successful deployment at scale Define deployment processes, operational playbooks, technical readiness criteria, escalation paths, and reliability practices for new hardware platforms Drive cross‑functional execution across lab bring‑up, rack/system integration, data center deployment, fleet monitoring, debugging, and issue resolution Stay hands‑on technically through architecture reviews, deployment planning, failure analysis, operational debugging, and critical system‑level decision‑making Identify gaps in tooling, observability, automation, validation coverage, and operational processes, and build plans to close them Establish clear metrics for deployment readiness, reliability, performance, maintainability, and operational health Build a strong engineering culture grounded in ownership, technical rigor, operational excellence, and high‑velocity execution Ensure OpenAI’s custom hardware platforms can be deployed and operated reliably, repeatably, and safely at scale Be a contributor and technical driver for the architecture and design of future ML systems You might thrive in this role if you: Enjoy mentoring and developing engineers while staying deeply engaged in technical execution Are excited by the challenge of bringing new custom hardware platforms into real‑world production data center environments Can operate across silicon, systems, software, infrastructure, and data center operations Are comfortable leading through ambiguity, especially when the hardware, tooling, and operational model are still being built Have strong judgment around deployment sequencing, technical risk, operational readiness, and when to elevate issues Communicate clearly across technical and operational teams, and can align stakeholders through complex deployment and production issues Care deeply about building practical systems, tools, and processes that work reliably at scale Have a bias toward ownership and are comfortable jumping into urgent technical issues when needed Qualifications 8+ years of engineering experience in hardware systems, infrastructure, data center deployment, production operations, systems engineering, silicon bring‑up, or related technical domains Strong technical depth in one or more of: hardware deployment, data center operations, rack‑scale systems, silicon bring‑up, systems validation, fleet operations, reliability engineering, infrastructure automation, or hardware/software integration Experience bringing complex hardware systems from development or validation into production environments Experience working closely with silicon, systems, software, infrastructure, networking, or data center teams Experience with deployment planning, operational readiness, incident response, debugging, and root‑cause analysis for production systems Experience building tooling, automation, observability, or operational processes that improve deployment quality and fleet reliability Demonstrated ability to hire, develop, and lead senior technical talent Ability to move fluidly between people leadership, technical strategy, and hands‑on operational problem solving Strong written and verbal communication skills, especially in high‑urgency, cross‑functional technical environments Experience working in fast‑moving environments Compensation Range: $342K - $445K USD About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general‑purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US‑based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non‑public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations. We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link. OpenAI Global Applicant Privacy Policy #J-18808-Ljbffr OpenAI

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Tech Lead, Deployment & Operations - Custom Infrastructure in San Francisco, CA vacancy
  •  ...-in-class technology infrastructure to power a global private...  ..., demand from customers, and a need to grow our...  ...including Architects, Tech Leads, Product, and Design...  ...qualitySupport healthy system operations and ensure high...  ...CI/CD pipelines and deployment processesExperience... 
    Operations
    Work at office
    Local area
    2 days per week
    3 days per week

    Forge Global

    San Francisco, CA
    1 day ago
  •  ...economy. But the payments infrastructure on which it runs is...  ...we aim to center our customers in every way -...  ...at Cruise Automation, leading ops and growth for Uber...  ...and scale Own and operate our Kubernetes (EKS)...  ...rollout strategies Tune deployments and cluster... 
    Operations
    Work at office

    AtoB

    San Francisco, CA
    20 hours ago
  • $207k - $300k

    Lead a dedicated development team responsible for...  ...requirements and infrastructure needs.Design, guide,...  ...is growing every day. Operating with scale and speed,...  ...budget and oversee the deployment of large-scale projects...  .... We empower Google customers with breakthrough... 
    Operations
    Worldwide

    Google

    San Francisco, CA
    1 day ago
  • $168k - $200k

     ...the best-in-class technology infrastructure to power a global private...  ...from investors, demand from customers, and a need to grow our...  ...Engineering, Security, and Operations to clarify requirements and...  ...reviews, and support production deployments. Create concise, articulate... 
    Operations
    Work at office
    Local area
    2 days per week
    3 days per week

    Forge Global

    San Francisco, CA
    3 days ago
  • $196k - $294k

     ...Vercel is the agentic infrastructure company. We free people...  ...built, extended, and operated by agents.We are building...  ..., supporting our customers, growing our community...  ...powers Vercel’s build and deployment lifecycle — from...  ...and developer delight.Lead projects end-to-end: from... 
    Operations
    Work from home
    Worldwide
    Flexible hours

    Vercel

    San Francisco, CA
    2 days ago
  • $300 per month

     ...only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from...  ...work of your career, help our customers and partners advance their AI...  ...a service.The DCIE team owns, deployment maintenance, observability, critical... 
    Operations
    Temporary work

    Crusoe

    San Francisco, CA
    3 days ago
  •  ...transforming how enterprises operate. We build intelligent...  ...moving. We’re backed by leading investors including...  ...As a Software Engineer, Infrastructure, you’ll build and scale the...  ...and supporting self-hosted deployments for enterprise customers who require on-premises or... 
    Operations

    Serval

    San Francisco, CA
    3 days ago
  •  ...Machine Learning for operations, risk, and safety. We...  ...billion annually. Our customers include Fortune 500...  ..., backed by industry-leading VCs. About the Role...  ...engineer to own the ML Infrastructure that powers how Voxel...  ...Own the train-to-deploy handoff - export trained... 
    Operations
    Work at office
    Flexible hours

    Voxel Labs

    San Francisco, CA
    3 days ago
  • $10k

     ...Ramp is a financial operations platform designed to save...  ...founders or executives of leading companies. The Ramp...  ...versions of our database infrastructure. They have been accountable for deploying production databases and...  ...to think through customer requirements and come up... 
    Operations
    Full time
    Work at office
    Home office
    Relocation package
    Flexible hours

    RAMP

    San Francisco, CA
    20 hours ago
  • $401k

     ...in what we build but operational in how we execute, and...  ...Engineer to join the Infrastructure Security (InfraSec)...  ...backbone of OpenAI’s customer and supercomputing environment...  ...developer ergonomics.Lead cross-functional...  ...an AI research and deployment company dedicated to... 
    Operations
    Work at office
    Local area
    Remote work
    Flexible hours

    OpenAI

    San Francisco, CA
    2 days ago
  • $180k - $240k

     ...FranciscoInfrastructure - Cloud Infrastructure /Full time /HybridWanna...  ...to get to orbit. We operate satellites, fly customer payloads, and handle entire...  ..., and environments.Lead efforts to automate and optimize...  ...provisioning (IaC), and deployment workflows.Own and evolve... 
    Operations
    Full time
    Temporary work

    Loft Orbital

    San Francisco, CA
    3 days ago
  • $207k - $300k

     ...Engineer Manager II, Infrastructure Share Software Engineer...  ...growing every day. Operating with scale and speed,...  ...and oversee the deployment of large-scale projects...  ...possible. We empower Google customers with breakthrough...  ...the future of world-leading hyperscale computing,... 
    Operations
    Worldwide

    Jobleads-US

    San Francisco, CA
    10 hours ago
  • $199k - $275k

     ...looking for a Technical Lead Manager (TLM) to lead...  ...Every merge and every deployment at Chime runs through...  ...Under Chime's Builder Operating Model, the TLM is a...  ...teams, security, Core Infrastructure, and other Engineering...  ...team's output. ~ A customer-obsessed mindset toward... 
    Operations
    Full time
    Work at office
    Local area
    Remote work

    Jobleads-US

    San Francisco, CA
    1 day ago
  • $133k - $184k

     ...roleThis role is on the Infrastructure Platform team, where...  ...supporting Chime's customer-serving products. As...  ...scaling, monitoring, and operations.Build tools to enable...  ...repo and continuous deployments.Partner with product...  ...with some of this tech - AWS services, Kubernetes... 
    Operations
    Full time
    Work at office
    Local area
    Remote work

    Chime

    San Francisco, CA
    4 days ago
  •  ...ship. Our robots are already operating in real homes and businesses...  ...With a growing team, strong customer demand, and capital for...  ...the companion app all run on infrastructure you'll own. As a Cloud Infrastructure...  ...from architecture through deployment. Build the real-time... 
    Operations
    Work from home

    Weave Robotics

    San Francisco, CA
    20 hours ago
  • $120k - $250k

     ...time. We're looking for an infrastructure engineer who's bothered by a...  ...product built on it. A deployment that needs someone to supervise...  ...bring product releases and customer commitments forward by weeks...  ...constraints, and make common operations straightforward without hiding... 
    Operations
    Full time
    Work at office
    Local area
    Immediate start
    Remote work
    Visa sponsorship
    Relocation package

    Jobleads-US

    San Francisco, CA
    1 day ago
  • $63k - $140k

     ...Amazon Web Services, Cloud Infrastructure Experienced Associate you will...  ...to troubleshooting and deployment activities, and build familiarity...  ...design, deployment, and operational support for secure and...  ...accessing sensitive company or customer information, handling... 
    Operations
    Full time
    Internship
    H1b

    PwC

    San Francisco, CA
    3 days ago
  •  ...needs it. Our customers are the ones building...  ...Francisco and backed by leading investors including...  ...own and architect core infrastructure systems that power our...  ...standards for system design, deployment, reliability, and infrastructure operations. Required... 
    Operations
    Local area

    AfterQuery

    San Francisco, CA
    1 day ago
  • $133k - $184k

     ...This role is on the Infrastructure Platform team , where...  ...infrastructure supporting Chime's customer-serving products. As...  ..., monitoring, and operations. Build tools to...  ...repo and continuous deployments. Partner with product...  ...with some of this tech - AWS services, Kubernetes... 
    Operations
    Full time
    Work at office
    Local area
    Remote work

    Engg

    San Francisco, CA
    4 days ago
  •  ...performing team that delivers infrastructure and performance...  ...companies. As a Lead Infrastructure Engineer...  ...storage engineering, deployment practices, integration...  ...LACP, port-channels.Operational experience with network...  ...businesses, clients, customers and employees up for... 

    JP Morgan Chase

    San Francisco, CA
    3 days ago
  •  ...meaningful impact.As a Senior Lead Infrastructure Engineer at JPMorgan Chase...  ...outputs and handling operational data according to sensitivity...  ...databases, storage engineering, deployment practices, integration,...  ...setting our businesses, clients, customers and employees up for... 

    JP Morgan Chase

    San Francisco, CA
    2 days ago
  • $160k - $195k

     ...only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from...  ...work of your career, help our customers and partners advance their AI...  ...version-controlled, and auditable deployments Plan, rack, and commission... 
    Operations
    Full time
    Temporary work
    Work at office

    Crusoe

    San Francisco, CA
    12 days ago
  • $200k - $260k

     ...'s best engineers and operators. If you are obsessed with...  ...Engineer to join our Infrastructure Team , focused on...  ...testing, field testing, deployment, monitoring, and...  ...the organization. Lead security architecture...  ...readiness reviews for custom business applications... 
    Operations
    Local area
    Flexible hours

    Nimble Robotics

    San Francisco, CA
    28 days ago
  • $350k

     ...a rapidly growing AI infrastructure provider delivering large...  ...partners with leading AI companies and infrastructure...  ..., covering deployment, GPU health, distributed...  ...infrastructure reliability, operational standards, and...  ...scalability Participate in customer-facing technical... 
    Operations
    Full time
    San Francisco, CA
    more than 2 months ago
  • $230k - $405k

    About the Team:Compute Infrastructure builds the platform that turns enormous amounts of compute...  ...AI. We design, provision, schedule, operate, and optimize the systems that connect...  ...impactAbout OpenAIOpenAI is an AI research and deployment company dedicated to ensuring that... 
    Operations
    Work at office
    Local area
    Flexible hours

    OpenAI

    San Francisco, CA
    20 hours ago
  •  ...Senior Software Engineer - Infrastructure Role Overview As a Senior Software...  ...engineering organization and lead major technical initiatives....  ...Define standards for system design, deployment, reliability, and infrastructure operations. Required Qualifications Strong... 
    Operations

    AfterQuery

    San Francisco, CA
    3 days ago
  • $61k - $101k

     ...training or certification in infrastructure engineering concepts, along...  ..., storage engineering, deployment practices, integration, automation...  ...-channels. We require operational experience in network...  ...set our businesses, clients, customers, and employees up for success... 
    Full time

    J.P. Morgan

    San Francisco, CA
    3 days ago
  • $168k - $227k

     ...Technical Lead Manager $168k – $227k CAD San Francisco...  ...ship applied AI that customers directly interact...  ...building and operating production systems, including...  ...design through deployment, monitoring, and iteration...  ...conversational AI infrastructure such as telephony, messaging... 
    Operations
    Full time
    Relocation
    Visa sponsorship
    1 day per week

    Point Defunct

    San Francisco, CA
    20 hours ago
  • $165k - $226.6k

     ...building the trusted, neutral infrastructure that enables organizations...  ...for builders and owners who operate with speed and urgency and execute...  ...projects on a modern tech stack (Redis, Docker, Terraform...  ...some of the largest cloud deployments in the enterprise worldBachelor... 
    Operations
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    12 hours ago
  • $180k - $250k

     ....We’re looking for an Infrastructure Engineer to own the reliability...  ...you design and operate.Our platform runs in a...  ....Improve CI/CD, deployments, and developer workflows...  ...risks before they impact customers.What we are looking...  ...make decisions.Tools and tech stackRuby on... 

    Thatch Health

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Tech Lead, Deployment & Operations - Custom Infrastructure. Be the first to apply!