Senior Network Engineer - GPU Cluster Networking
Advanced Micro Devices , Inc.
Senior Network Engineer At AMD, we believe technology has the power to solve the world's most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. Whether you're designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we'll advance your career. We are seeking a Senior Network Engineer to join the AMD IT System Engineering team. This role is responsible for the architecture, deployment, optimization, automation, and production operation of high-performance backend networks supporting large-scale AMD GPU clusters. The engineer will own the network path from the GPU server and NIC through the data center switching fabric, ensuring that distributed AI training, large language model, inference, and HPC workloads receive predictable bandwidth, low latency, and reliable collective communication performance. The ideal candidate will have experience designing, scaling, and operating backend network infrastructure for GPU clusters with approximately 10,000 or more GPUs, or comparable hyperscale AI and HPC environments. The primary focus of this position is high-speed Ethernet and RoCEv2 networking for AMD Instinct accelerator clusters. You will work across switches, NICs, optics, RDMA, Linux networking, PCIe and NUMA topology, ROCm, RCCL, SLURM, Kubernetes, storage networks, automation platforms, and observability systems. You will partner with AMD AI engineering, network engineering, data center, storage, security, platform, and application teams to ensure the backend network fabric is not a bottleneck to GPU workload performance. You are a highly experienced, hands-on network engineer with deep expertise in data center networking, RDMA, RoCEv2, and large-scale GPU cluster fabrics with approximately 10,000 or more GPUs. You understand how distributed GPU workloads generate traffic across the backend network and how application performance is affected by network topology, congestion, GPU-to-NIC locality, routing, switch buffering, traffic-class configuration, and collective communication patterns. You take responsibility for end-to-end outcomes, including architecture, implementation, qualification, production deployment, monitoring, incident response, capacity planning, and continuous improvement. You use telemetry and repeatable performance testing to validate designs and make data-driven engineering decisions. You are comfortable leading complex technical initiatives, mentoring engineers, documenting architecture and operating standards, and working across globally distributed organizations. Key Responsibilities: Architect, deploy, operate, and continuously improve high-performance backend networks for large-scale AMD Instinct GPU clusters. Design network fabrics capable of supporting AI and HPC environments ranging from individual GPU racks to clusters containing 10,000 or more GPUs. Own the backend network architecture from the GPU server and network interface card through the leaf-spine switching fabric. Design and optimize high-speed Ethernet fabrics using RoCEv2 and 100/200/400 GbE technologies. Develop scalable network topologies, including leaf-spine, Clos, fat-tree, rail-optimized, multi-plane, and non-blocking fabric architectures. Perform network topology modeling, oversubscription analysis, traffic-flow analysis, bandwidth planning, port-capacity planning, failure-domain analysis, and long-term growth forecasting. Configure, tune, validate, and troubleshoot lossless or near-lossless RoCEv2 environments, including PFC, ECN, DCQCN, QoS, ECMP, Switch buffer and queue management, DSCP and priority mappings Design and operate routing and switching environments using technologies such as BGP, ECMP, VLAN, VRF, EVPN, and VXLAN. Optimize end-to-end communication performance across GPUs, NICs, switches, CPUs, PCIe devices, storage systems, and the Linux networking stack. Lead production incident response, root-cause analysis, corrective actions, and preventive engineering improvements for GPU cluster networks. Plan and execute network expansions, cluster scale-outs, switch replacements, capacity upgrades, and fabric migrations Preferred Experience: Significant experience designing, deploying, and operating production data center networks for AI, GPU, HPC, cloud, or other large-scale distributed computing environments. Experience designing, scaling, or operating backend network infrastructure for GPU clusters containing approximately 10,000 or more GPUs, or similarly sized hyperscale compute environments. Deep knowledge of data center networking fundamentals; Routing and switching, VLANs and subnetting, BGP and ECMP, Quality of Service, MTU configuration, Switch buffering, Network segmentation Strong hands-on experience with RDMA and RoCEv2 in production environments. Demonstrated experience configuring, tuning, and troubleshooting PFC, ECN, DCQCN, QoS, switch buffers, NIC queues, RDMA traffic classes, and lossless or near-lossless Ethernet. Strong understanding of leaf-spine, Clos, fat-tree, rail-optimized, and multi-plane network architectures. Experience with network routing technologies such as BGP and ECMP and overlay technologies such as EVPN and VXLAN. Strong understanding of GPU cluster topology, including GPU-to-GPU, GPU-to-NIC, CPU-to-NIC, PCIe, NUMA, and network locality. Experience building monitoring and observability solutions using Prometheus, Grafana, streaming telemetry, gNMI, SNMP, sFlow, or equivalent platforms. Experience with Juniper data center switching platforms and Junos OS, including configuration and troubleshooting Experience with AMD Instinct accelerators, ROCm, RCCL, and AMD GPU software environments. Experience with AMD Pensando AI NICs, SmartNICs, DPUs, or other AMD Pensando networking technologies. Experience designing backend networks specifically for large language model training and other communication-intensive distributed AI workloads. Experience with Ethernet fabric technologies such as BGP, EVPN, VXLAN, and modern leaf-spine data center architectures. Academic Credits: Bachelor's or Master's degree in Computer Engineering, or a related field, or equivalent practical experience. Location: San Jose, CA OR Austin, TX This role is not eligible for visa sponsorship. Advanced Micro Devices , Inc.
$184k - $287.5k
...computing, known for inventing the GPU and driving breakthroughs in... ...that enable researchers and engineers to develop the next... ...globally.We are looking for a senior networking engineer to lead the network... ...and we are scaling it toward clusters of ten thousand nodes and...SeniorFull timeRemote work- Senior Principal Network Engineer Austin, Texas, United States Graphcore is one of the world's leading innovators in Artificial Intelligence compute... ...performance computing (HPC) network fabrics supporting GPU clusters. You will work across hardware, networking, and AI...SeniorFlexible hours
$224k - $356.5k
...next era of computing. An era in which our GPU acts as the brains of computers, robots,... ..., and parallelism efficiency on edge cluster configurationsProduce performance analysis... ...MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience...SeniorFull timeLocal area$168k - $270.25k
...computing. An era in which our GPU acts as the brains of computers... ...Enterprise Experience (NVEX) Solutions Engineering team is looking for a senior Computer or Software Engineer.... ...expertise in ground-breaking network technology used in AI clusters. Our software engineers connect...SeniorFull timeWeekend work$140k - $224.25k
The NVIDIA Experience (NVEX) Solutions Engineering team is looking for a senior Computer or Software Engineer who is ready to become an authority in ground-breaking network technology used in AI clusters. Our team of software engineers bridge the gap between the customer...SeniorFull timeWeekend work$114.6k - $234.6k
AI2NE strives to be a global leader in the RDMA cluster networking domain and enable seamless, accelerated High-Performance Compute (HPC),... ...as the job remains posted.Career Level - IC4Collaborate with engineers from L1 optical engineering team, network design, delivery and...SeniorTemporary workFlexible hours- ..., and play a key role in supporting the network technologies that power next-generation... ...networks that enable AMD's most advanced engineering initiatives.In this role, you will collaborate... ...connectivity for servers, racks, clusters, and validation environments.Install, configure...SeniorVisa sponsorshipAfternoon shift
$184k - $287.5k
...GB300 GPUs. NVIDIA seeks a Senior Software Engineer for our CSP (Cloud Service... ...of multi-rack, multi-tenant clusters: scheduler behavior, container... ...-services that expose new GPU capabilities.Drive joint... ...cloud-native stacks across networking (RDMA/RoCE), storage, and control...SeniorFull timeRemote work$224k - $356.5k
...to-end behavior across GPUs, networking, storage, and software stacks. We are seeking a Senior Performance Engineer to characterize workloads,... ...engineers, platform teams, and GPU architects to validate and... ...Experience analyzing large-scale AI clusters or distributed training and...SeniorFull timeRemote work$224k - $356.5k
...Our data-center platforms bring together GPUs, CPUs, networking, systems, and software to solve some of the world’s... ...exciting computing problems. We are looking for a Senior System Software Engineer to join NVIDIA’s GPU Performance and Power Management Software team. You...SeniorFull timeRemote work$170k - $230k
...DescriptionWe are hiring a Senior Platform Engineer to join the Autonomous Vehicle... ...of production-grade clusters.Required: Strong proficiency... ...hardware-level performance (GPU passthrough) and clean cloud... ...broader GCP ecosystem, including networking and IAM; equivalent...SeniorFull timeWork experience placementWork at officeLocal areaWork from homeFlexible hours- ...Position Overview We are seeking a Senior GPU Systems & Fabric Engineer to serve as the critical bridge between our physical GPU/network infrastructure and the Kubernetes abstraction... ...tenant inference workloads and maximize cluster utilization. Profile and tune kernel-...SeniorFull timeLocal area
$126.2k - $264.1k
Job Description What You’ll Do As a Senior Principal Network Reliability Engineer, you will: Provide technical leadership for the deployment, operation,... ...critical services. Experience supporting AI infrastructure, GPU networking, or large-scale data center deployments....SeniorTemporary workFlexible hours$184k - $287.5k
Senior Software Engineer NVIDIA is at the forefront of the generative AI revolution... ...workloads across NVIDIA GPU platforms at the largest... ...capabilities that keep large clusters productive. This is a hands... ...performance across compute, memory, networking, and communication layers...Senior- ...Senior Network Engineer Location: Remote, USA (must currently reside in the US for consideration) Contract Position: 12-18months+ We are seeking a highly motivated and skilled Senior Network Engineer with strong experience in network design and security. The...SeniorContract workRemote work
- ...Enterprise Network Engineer - Must be bilingual - Korean/English Seeking an experienced Enterprise Network Engineer fluent in Korean and English to join a high-availability manufacturing network environment. This position involves collaborating with stakeholders...Senior
$103.62k - $172.63k
...need at LPL Financial to shape your success while helping clients pursue their financial goals. Job Overview:LPL is seeking a Senior Network Engineer to join our Datacenter Network Engineering team. This individual will be responsible for the design, engineering, and...SeniorFull timeWork experience placementWork from home- ...Senior Level Network Engineer We are looking for an experienced Senior Network Engineer to execute network assessment process, topology, and optimization. Job Summary: The Senior Network Engineer is responsible for executing data center network assessments and...SeniorRemote workVisa sponsorship
$150k - $217k
...responsible for development of solutions for Google's Enterprise network, executing new features qualification and validation through... ...technical and communication aspects.Coach and collaborate with peer engineers towards achieving common goals in assigned projects.Minimum...Senior- ...enhance maritime operations through autonomous and intelligent platforms.We are seeking an experienced and results-driven Senior Network Engineer to join our IT team and lead the design, deployment, and management of mission-critical network infrastructure. The ideal candidate...SeniorPermanent employmentTemporary work
$150k - $217k
Lead the requirement analysis, engineering design, and development of solutions for Google’s Enterprise network infrastructure, executing new feature qualification, deployment, and optimization.Deliver secure and dependable network services by scaling systems sustainably...SeniorWorldwide- Bitdeer Technologies Group in the United States is seeking a senior network engineer to design and operate a multi-tenant, multi-region GPU cloud network supporting distributed AI training. You will implement tenant isolation, high-performance interconnects, and cross-...Senior
$186.07k - $218.9k
...Lead on the Exchange team within Institutional, you'll own the network infrastructure, deployment, and operational tooling behind a unified... ...systems, and participate in an on-call rotation. Mentor engineers and partner cross-functionally with SRE/Infra, Product, and...SeniorLocal area- ...Supply Planning to lead a team of planners overseeing data center GPU programs. You will translate demand, capacity, and material... ...making, and process improvements to scale planning across multi-site networks and ensure on-time delivery and optimal inventory. #J-18808-...Senior
- ...a highly skilled and motivated Cloud Senior DevOps Engineer to join our AI Cloud team. In this high... ...computing resources (e.g., GPU clusters) to support high-performance AI workloads... ...), or Cloud Infrastructure roles. Networking & OS: Expert-level knowledge of Linux...SeniorFull timeLocal area
$70 - $80 per hour
...Immediate need for a talented Senior Network Engineer This is a 12 months contract opportunity with long-term potential and is located in Round Rock, TX (Onsite). Please review the job description below and contact me ASAP if you are interested. Job ID: 26-245...SeniorContract workLocal areaImmediate start- ...Network Engineer Austin, TX Delart is home to a team of world-class engineers and project leaders dedicated to developing the next generation of advanced networking technologies, consumer devices, and innovative technology solutions. Trusted by some of the world...SeniorPermanent employmentTemporary workFlexible hours
- ...and high-performance computing products within the Data Center GPU group. You will drive cross-functional execution from SoC planning... ...deployment, partnering with silicon, firmware, software, networking, and manufacturing teams. The role requires strong leadership,...Senior
- A tech company in Austin, Texas is seeking an experienced Network Engineer to join its Engineering team. The ideal candidate will be responsible for implementing network solutions, supporting multi-site designs, and maintaining network security. Required qualifications...Senior
$112.5k - $202.5k
...excited about working with cutting-edge network technologies? Do you enjoy acting quickly... ...to solve problems? Join our Network Engineering team The Network Engineering team is... ...services. Partner with the best As a Senior Network Engineer you will lead efforts...SeniorWork experience placementWork at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Network Engineer - GPU Cluster Networking. Be the first to apply!
- network automation engineer Austin, TX
- cisco ccnp network engineer Austin, TX
- ccna network engineer Austin, TX
- network engineer full time Austin, TX
- network engineer - transport Austin, TX
- network infrastructure engineer Austin, TX
- work from home network engineer Austin, TX
- network engineer - remote Austin, TX
- senior network engineer remote Austin, TX
- wireless network engineer Austin, TX


