Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Network Operations Engineer, AI Networking

OpenAI

About the TeamOpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads.About the RoleWe are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network.The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil.Key ResponsibilitiesOwn the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers.Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR).Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks.Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact.Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance.Support new AI cluster deployments, data center expansions, and infrastructure migrations in partnership with deployment and engineering teams.Partner with cloud service providers (CSPs), colocation providers, Smart Hands teams, and hardware vendors to maintain production infrastructure.Perform root-cause analysis (RCA) for production incidents and drive permanent corrective actions that eliminate recurring issues.Build and maintain monitoring, telemetry, dashboards, and alerting to improve network observability and proactive issue detection.Develop and improve operational runbooks, playbooks, troubleshooting documentation, and standard operating procedures.Automate repetitive operational tasks using Python and infrastructure automation frameworks to reduce toil and improve efficiency.Continuously identify opportunities to improve service reliability, scalability, operational maturity, and engineering efficiency.QualificationsBachelor’s degree in Computer Science, Network Engineering, or a related discipline, or equivalent practical experience.5+ years of experience operating large-scale data center, cloud, AI, or HPC network infrastructure.Experience supporting production network environments with high-availability requirements.Hands-on experience with one or more of the following platforms: Cisco NX-OS, Arista EOS, NVIDIA Spectrum / Cumulus Linux, or Juniper JunOS.Strong knowledge of Layer 2 and Layer 3 networking, BGP, OSPF, ECMP, MLAG, LACP, VRFs, and VLANs.Experience troubleshooting physical infrastructure, including fiber optics, transceivers, DAC/AOC cables, and high-speed Ethernet links.Experience performing software upgrades, hardware maintenance, and production change management.Excellent analytical and troubleshooting skills, with the ability to communicate technical risk clearly across teams.Preferred SkillsExperience operating AI or High-Performance Computing (HPC) network environments.Experience with NVIDIA AI networking technologies and GPU infrastructure.Experience supporting RoCE v2 or RDMA-based Ethernet fabrics, with a strong understanding of Priority Flow Control (PFC), Explicit Congestion Notification (ECN), Data Center Quantized Congestion Notification (DCQCN), Quality of Service (QoS), and lossless Ethernet networking.Experience supporting 100G, 200G, 400G, and 800G Ethernet networks.Experience with GPU platforms including NVIDIA HGX, DGX, GB200, or equivalent AI infrastructure.Experience supporting distributed storage environments such as VAST, DDN, or similar technologies.Experience working with cloud service providers such as AWS, Azure, or Google Cloud, and with third-party colocation providers.Experience with network monitoring and telemetry technologies, including Prometheus, Grafana, gNMI, streaming telemetry, SNMP, or similar tools.Experience developing automation using Python, Git, REST APIs, Terraform, or similar automation frameworks.Work Environment and On-CallParticipate in a 24x7 on-call rotation supporting mission-critical AI infrastructure.Support time-sensitive production incidents, maintenance windows, capacity expansions, and network changes with a focus on service availability and minimal customer impact.This role requires up to 30% travel to data center locations for new turnups and acceptance activities, as needed.About OpenAIOpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Network Operations Engineer, AI Networking in San Francisco, CA vacancy
  • $10k - $20k

     ...centres, helping customers streamline operations, improve customer experience, and stay...  ...services. The company is looking for a Network Automation Engineer to design and build high-speed, low-...  ...00G connectivity, supporting advanced AI and LLM workloads in a high-... 
    Suggested
    Full time
    San Francisco, CA
    23 days ago
  • $193k - $234k

     ...intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons...  ...a high-energy, detail-oriented Staff Network Production Engineer to lead the physical and logical implementation... 
    Suggested
    Temporary work
    Remote work

    Crusoe

    San Francisco, CA
    a month ago
  • $200k

     ...Ready to architect the high-speed networks powering the AI era? Join a trailblazing leader in GPU-accelerated computing, designing and operating the network fabric that supports some of the most demanding computational workloads on the planet. Gain exposure to... 
    Suggested
    Full time
    San Francisco, CA
    more than 2 months ago
  •  ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of...  ...highly capable software, hardware and network engineers building one of the largest AI training...  ...automation underneath all of it, while operating the existing network flawlessly for all... 
    Suggested
    Local area
    Flexible hours

    Lambda

    San Francisco, CA
    2 days ago
  •  ...Principal Network Engineer Houston; New York; San Francisco; Seattle Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI...  ...future. About The Role The Network Operations and Engineering teams at Nscale... 
    Suggested
    Flexible hours

    Nscale

    San Francisco, CA
    4 days ago
  • Network Engineer Reports To: Director of Network Engineering Location: San Francisco, California Position Summary The position will focus on the core functions of network and security support, implementation, analysis, maintenance, administration and reporting for Retail... 

    Software Technology Inc

    San Francisco, CA
    5 days ago
  • Berkeley Lab is seeking a Network Engineer to join the team that powers AI-driven network operations and HPC/AI workflows. You’ll help design, implement, and operate automation and observability for a 1 Tb/s border network and an 800G/400G data center backbone supporting... 

    Berkeley Lab

    Berkeley, CA
    3 days ago
  • $195k - $235k

     ...intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons...  ...Role:Crusoe Cloud is seeking a Staff Network Production Operations Engineer to help own production reliability... 
    Temporary work
    Worldwide

    Crusoe

    San Francisco, CA
    21 hours ago
  • Lawrence Berkeley National Laboratory's NERSC is seeking a Network Engineer, Platform, Automation & HPC/AI to advance the 1 Tb/s border network and an 800G/400G data center backbone supporting HPC workloads and a wide user base across scientific computing. This role spans... 

    Lawrence Berkeley National Laboratory

    Berkeley, CA
    2 days ago
  • $195k - $235k

     ...intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons...  ...: Crusoe Cloud is seeking a Staff Network Production Operations Engineer to help own production reliability... 
    Temporary work
    Worldwide

    Crusoe

    San Francisco, CA
    8 days ago
  •  ...Gimlet is building the next generation of AI infrastructure: large-scale AI...  ...About this Role Gimlet Labs is seeking a Network Engineer to design, build, and scale the network...  ...designed, deployed, interconnected, and operated. The ideal candidate has deep technical... 

    Gimlet Labs

    San Francisco, CA
    3 days ago
  • $100k - $135k

     ...logistics. We’re looking for a skilled Network Engineer to design, deploy, and maintain robust...  ...$225,000 1 week ago Network Engineer, Operations and Support (Labs) Menlo Park, CA $147,...  ...directly into each article, started with the help of AI. #J-18808-Ljbffr Insight Global
    Full time
    Internship
    Remote work
    3 days per week

    Insight Global

    Brisbane, CA
    2 days ago
  • $150k - $190k

    Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we...  ..., and large enterprises. Lightning AI operates globally with offices in New York City,...  ...Looking For Lightning AI is seeking a Senior Network Engineer with hands-on Cumulus Linux expertise to... 
    Work at office
    Remote work
    Work from home
    Flexible hours
    2 days per week

    Lightning AI

    San Francisco, CA
    1 day ago
  • $100k - $130k

    Join to apply for the Senior Network Infrastructure Engineer role at Pliancy Join to apply for the Senior...  .... In addition to managing day-to-day operations, you will help play a key role in maintaining...  ...into each article, started with the help of AI. #J-18808-Ljbffr Pliancy
    Full time
    For contractors
    For subcontractor
    Work at office
    Remote work
    Relocation

    Pliancy

    San Francisco, CA
    2 days ago
  • $350k

     ...the knowledge and tools to make AI work for their unique needs and goals. We are scientists, engineers, and builders who’ve created...  ...the Role We're looking for a network engineer to own the lowest layers...  ...(Python or Rust). Experience operating large‑scale clusters and... 
    Local area
    Visa sponsorship
    Relocation package

    Thinking Machines Lab

    San Francisco, CA
    4 days ago
  • About the Team OpenAI's Network Engineering team within IT and Security advances the mission of deploying...  ...network services. We build and operate the connectivity that supports OpenAI's...  ...user-centered design, we enable impactful AI research, corporate operations, and... 
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    5 days ago
  • Beam is an ultrafast AI inference platform. We built a serverless runtime...  ...Deploy and validate data center network infrastructure (front‑end, back‑end...  ...need them. Partner with DC Operations, ICT, Hardware, and Network Engineering to identify blockers early, elevate... 

    Beam

    San Francisco, CA
    4 days ago
  •  ...Superintelligence CloudLambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers....  ...our business. We partner across the company—Finance, GTM, Engineering, and People—to implement tools, automate workflows, and ensure... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda

    San Francisco, CA
    3 days ago
  • Senior Network Engineer Houston; New York; San Francisco; Seattle About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI...  ...the design, deployment, and ongoing operation of all front-end networking services... 

    Nscale

    San Francisco, CA
    5 days ago
  • $200k - $240k

    Senior Network Engineer Altruist is the modern custodian built exclusively for independent financial...  ...are designed, segmented, secured, and operated — from cloud routing topology and...  ...shape much of this roadmap, and bringing an AI-native approach to the work. You will set... 
    Work at office
    Immediate start
    3 days per week

    Altruist

    San Francisco, CA
    5 days ago
  •  ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving...  ...'s high performance cloud network Work on deploying and...  ...configuration management and operation Work with internal and external...  ...on-call rotation for Network Engineering team You Have 10+ years of... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda

    San Francisco, CA
    5 days ago
  • We are an applied AI lab building end-to-end software agents....  ...Devin, the first AI software engineer. Our team is extremely talent...  ...the role We're looking for a Network Engineer to own the reliability...  ...networks. Strong candidates have operated networks in a hypergrowth... 
    For contractors
    Work at office
    Immediate start
    Remote work

    Cognition Corp

    San Francisco, CA
    2 days ago
  • $70 - $76 per hour

     ...critical internal systems and networks during a transitional period....  ...triage issues, maintain operational stability, and ensure a seamless...  ...Get notified about new Network Engineer jobs in San Francisco, CA ....  ...article, started with the help of AI. #J-18808-Ljbffr Milestone... 
    Permanent employment
    Full time
    Contract work
    Work at office
    Immediate start
    Remote work

    Milestone Technologies

    San Francisco, CA
    2 days ago
  • Network Engineer Gimlet is building the first multi-silicon neocloud designed for fast, efficient AI inference. We combine large-scale compute infrastructure with an execution platform...  ...centers are designed, deployed, and operated as the platform scales. You will work across... 

    Gimlet Labs

    San Francisco, CA
    4 days ago
  • $300 per month

     ...intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to...  ...About the Role At Crusoe, our Production Engineering team ensures the reliability and scalability of... 
    Temporary work

    Crusoe

    San Francisco, CA
    25 days ago
  •  ...Specter Data Operations Engineer Specter's mission is to help automate the physical world. Today, we build video sensors with state-of-the-art AI agents that answer any question, anywhere in their environments. Our systems can automatically detect and reason about any... 

    Specter Services LLC

    San Francisco, CA
    3 days ago
  • $94.4k - $224.6k

     ...and, Interactive, Technology, and Operations services, all powered by the world’s largest network of Advanced Technology and...  ...Visit us at A successful Network Engineer Architect combines deep technical...  ...designing infrastructure to support AI, analytics, data-intensive, and... 
    Work experience placement
    Live in
    Work at office
    Local area
    Remote work

    Accenture

    San Francisco, CA
    1 day ago
  • $300 per month

     ...the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from...  ...are seeking Senior Software Engineers to design and develop internal...  ...and configuration of servers, network switches, power delivery units... 
    Full time
    Temporary work

    Crusoe

    San Francisco, CA
    21 hours ago
  • $80 - $85 per hour

     ...us regarding this posting.Overview LABUR is partnering with a client to hire a Senior AI Platform & Production Engineer for a hands-on engineering role focused on building, operating, and continuously improving shared Enterprise AI platform capabilities and the production... 
    Remote work

    Labur

    San Francisco, CA
    6 days ago
  •  ...for a career glow up? As  Senior Manager, Network Retail & Corporate Engineering , you'll be driving the vision, execution, and operational excellence of Sephora's global retail and...  ...observability, automation, and GenAI/Agentic AI-driven operational capabilities to improve... 
    Hourly pay
    Full time
    Work at office

    Sephora

    San Francisco, CA
    more than 2 months ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Network Operations Engineer, AI Networking. Be the first to apply!