Network Operations Engineer, AI Networking
OpenAI
About the TeamOpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads.About the RoleWe are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network.The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil.Key ResponsibilitiesOwn the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers.Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR).Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks.Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact.Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance.Support new AI cluster deployments, data center expansions, and infrastructure migrations in partnership with deployment and engineering teams.Partner with cloud service providers (CSPs), colocation providers, Smart Hands teams, and hardware vendors to maintain production infrastructure.Perform root-cause analysis (RCA) for production incidents and drive permanent corrective actions that eliminate recurring issues.Build and maintain monitoring, telemetry, dashboards, and alerting to improve network observability and proactive issue detection.Develop and improve operational runbooks, playbooks, troubleshooting documentation, and standard operating procedures.Automate repetitive operational tasks using Python and infrastructure automation frameworks to reduce toil and improve efficiency.Continuously identify opportunities to improve service reliability, scalability, operational maturity, and engineering efficiency.QualificationsBachelor’s degree in Computer Science, Network Engineering, or a related discipline, or equivalent practical experience.5+ years of experience operating large-scale data center, cloud, AI, or HPC network infrastructure.Experience supporting production network environments with high-availability requirements.Hands-on experience with one or more of the following platforms: Cisco NX-OS, Arista EOS, NVIDIA Spectrum / Cumulus Linux, or Juniper JunOS.Strong knowledge of Layer 2 and Layer 3 networking, BGP, OSPF, ECMP, MLAG, LACP, VRFs, and VLANs.Experience troubleshooting physical infrastructure, including fiber optics, transceivers, DAC/AOC cables, and high-speed Ethernet links.Experience performing software upgrades, hardware maintenance, and production change management.Excellent analytical and troubleshooting skills, with the ability to communicate technical risk clearly across teams.Preferred SkillsExperience operating AI or High-Performance Computing (HPC) network environments.Experience with NVIDIA AI networking technologies and GPU infrastructure.Experience supporting RoCE v2 or RDMA-based Ethernet fabrics, with a strong understanding of Priority Flow Control (PFC), Explicit Congestion Notification (ECN), Data Center Quantized Congestion Notification (DCQCN), Quality of Service (QoS), and lossless Ethernet networking.Experience supporting 100G, 200G, 400G, and 800G Ethernet networks.Experience with GPU platforms including NVIDIA HGX, DGX, GB200, or equivalent AI infrastructure.Experience supporting distributed storage environments such as VAST, DDN, or similar technologies.Experience working with cloud service providers such as AWS, Azure, or Google Cloud, and with third-party colocation providers.Experience with network monitoring and telemetry technologies, including Prometheus, Grafana, gNMI, streaming telemetry, SNMP, or similar tools.Experience developing automation using Python, Git, REST APIs, Terraform, or similar automation frameworks.Work Environment and On-CallParticipate in a 24x7 on-call rotation supporting mission-critical AI infrastructure.Support time-sensitive production incidents, maintenance windows, capacity expansions, and network changes with a focus on service availability and minimal customer impact.This role requires up to 30% travel to data center locations for new turnups and acceptance activities, as needed.About OpenAIOpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an
$10k - $20k
...centres, helping customers streamline operations, improve customer experience, and stay... ...services. The company is looking for a Network Automation Engineer to design and build high-speed, low-... ...00G connectivity, supporting advanced AI and LLM workloads in a high-...SuggestedFull time$193k - $234k
...intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons... ...a high-energy, detail-oriented Staff Network Production Engineer to lead the physical and logical implementation...SuggestedTemporary workRemote work$200k
...Ready to architect the high-speed networks powering the AI era? Join a trailblazing leader in GPU-accelerated computing, designing and operating the network fabric that supports some of the most demanding computational workloads on the planet. Gain exposure to...SuggestedFull time- ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of... ...highly capable software, hardware and network engineers building one of the largest AI training... ...automation underneath all of it, while operating the existing network flawlessly for all...SuggestedLocal areaFlexible hours
- ...Principal Network Engineer Houston; New York; San Francisco; Seattle Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI... ...future. About The Role The Network Operations and Engineering teams at Nscale...SuggestedFlexible hours
- Network Engineer Reports To: Director of Network Engineering Location: San Francisco, California Position Summary The position will focus on the core functions of network and security support, implementation, analysis, maintenance, administration and reporting for Retail...
- Berkeley Lab is seeking a Network Engineer to join the team that powers AI-driven network operations and HPC/AI workflows. You’ll help design, implement, and operate automation and observability for a 1 Tb/s border network and an 800G/400G data center backbone supporting...
$195k - $235k
...intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons... ...Role:Crusoe Cloud is seeking a Staff Network Production Operations Engineer to help own production reliability...Temporary workWorldwide- Lawrence Berkeley National Laboratory's NERSC is seeking a Network Engineer, Platform, Automation & HPC/AI to advance the 1 Tb/s border network and an 800G/400G data center backbone supporting HPC workloads and a wide user base across scientific computing. This role spans...
$195k - $235k
...intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons... ...: Crusoe Cloud is seeking a Staff Network Production Operations Engineer to help own production reliability...Temporary workWorldwide- ...Gimlet is building the next generation of AI infrastructure: large-scale AI... ...About this Role Gimlet Labs is seeking a Network Engineer to design, build, and scale the network... ...designed, deployed, interconnected, and operated. The ideal candidate has deep technical...
$100k - $135k
...logistics. We’re looking for a skilled Network Engineer to design, deploy, and maintain robust... ...$225,000 1 week ago Network Engineer, Operations and Support (Labs) Menlo Park, CA $147,... ...directly into each article, started with the help of AI. #J-18808-Ljbffr Insight GlobalFull timeInternshipRemote work3 days per week$150k - $190k
Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we... ..., and large enterprises. Lightning AI operates globally with offices in New York City,... ...Looking For Lightning AI is seeking a Senior Network Engineer with hands-on Cumulus Linux expertise to...Work at officeRemote workWork from homeFlexible hours2 days per week$100k - $130k
Join to apply for the Senior Network Infrastructure Engineer role at Pliancy Join to apply for the Senior... .... In addition to managing day-to-day operations, you will help play a key role in maintaining... ...into each article, started with the help of AI. #J-18808-Ljbffr PliancyFull timeFor contractorsFor subcontractorWork at officeRemote workRelocation$350k
...the knowledge and tools to make AI work for their unique needs and goals. We are scientists, engineers, and builders who’ve created... ...the Role We're looking for a network engineer to own the lowest layers... ...(Python or Rust). Experience operating large‑scale clusters and...Local areaVisa sponsorshipRelocation package- About the Team OpenAI's Network Engineering team within IT and Security advances the mission of deploying... ...network services. We build and operate the connectivity that supports OpenAI's... ...user-centered design, we enable impactful AI research, corporate operations, and...Work at officeRelocation package
- Beam is an ultrafast AI inference platform. We built a serverless runtime... ...Deploy and validate data center network infrastructure (front‑end, back‑end... ...need them. Partner with DC Operations, ICT, Hardware, and Network Engineering to identify blockers early, elevate...
- ...Superintelligence CloudLambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers.... ...our business. We partner across the company—Finance, GTM, Engineering, and People—to implement tools, automate workflows, and ensure...Work at officeLocal areaWork from homeFlexible hours
- Senior Network Engineer Houston; New York; San Francisco; Seattle About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI... ...the design, deployment, and ongoing operation of all front-end networking services...
$200k - $240k
Senior Network Engineer Altruist is the modern custodian built exclusively for independent financial... ...are designed, segmented, secured, and operated — from cloud routing topology and... ...shape much of this roadmap, and bringing an AI-native approach to the work. You will set...Work at officeImmediate start3 days per week- ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving... ...'s high performance cloud network Work on deploying and... ...configuration management and operation Work with internal and external... ...on-call rotation for Network Engineering team You Have 10+ years of...Work at officeLocal areaWork from homeFlexible hours
- We are an applied AI lab building end-to-end software agents.... ...Devin, the first AI software engineer. Our team is extremely talent... ...the role We're looking for a Network Engineer to own the reliability... ...networks. Strong candidates have operated networks in a hypergrowth...For contractorsWork at officeImmediate startRemote work
$70 - $76 per hour
...critical internal systems and networks during a transitional period.... ...triage issues, maintain operational stability, and ensure a seamless... ...Get notified about new Network Engineer jobs in San Francisco, CA .... ...article, started with the help of AI. #J-18808-Ljbffr Milestone...Permanent employmentFull timeContract workWork at officeImmediate startRemote work- Network Engineer Gimlet is building the first multi-silicon neocloud designed for fast, efficient AI inference. We combine large-scale compute infrastructure with an execution platform... ...centers are designed, deployed, and operated as the platform scales. You will work across...
$300 per month
...intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to... ...About the Role At Crusoe, our Production Engineering team ensures the reliability and scalability of...Temporary work- ...Specter Data Operations Engineer Specter's mission is to help automate the physical world. Today, we build video sensors with state-of-the-art AI agents that answer any question, anywhere in their environments. Our systems can automatically detect and reason about any...
$94.4k - $224.6k
...and, Interactive, Technology, and Operations services, all powered by the world’s largest network of Advanced Technology and... ...Visit us at A successful Network Engineer Architect combines deep technical... ...designing infrastructure to support AI, analytics, data-intensive, and...Work experience placementLive inWork at officeLocal areaRemote work$300 per month
...the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from... ...are seeking Senior Software Engineers to design and develop internal... ...and configuration of servers, network switches, power delivery units...Full timeTemporary work$80 - $85 per hour
...us regarding this posting.Overview LABUR is partnering with a client to hire a Senior AI Platform & Production Engineer for a hands-on engineering role focused on building, operating, and continuously improving shared Enterprise AI platform capabilities and the production...Remote work- ...for a career glow up? As Senior Manager, Network Retail & Corporate Engineering , you'll be driving the vision, execution, and operational excellence of Sephora's global retail and... ...observability, automation, and GenAI/Agentic AI-driven operational capabilities to improve...Hourly payFull timeWork at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Network Operations Engineer, AI Networking. Be the first to apply!
- lead network engineer San Francisco, CA
- cisco network engineer San Francisco, CA
- production network engineer San Francisco, CA
- network engineer San Francisco, CA
- network engineer full time San Francisco, CA
- remote cisco network engineer San Francisco, CA
- network consulting engineer San Francisco, CA
- network engineer internship San Francisco, CA
- network applications engineer San Francisco, CA
- principal network engineer San Francisco, CA


