Senior Network Engineer - GPU Cluster Networking
AMD
ADVANCE YOUR CAREER. ADVANCE THE WORLD. At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career.THE ROLE:We are seeking a Senior Network Engineer to join the AMD IT System Engineering team.This role is responsible for the architecture, deployment, optimization, automation, and production operation of high-performance backend networks supporting large-scale AMD GPU clusters. The engineer will own the network path from the GPU server and NIC through the data center switching fabric, ensuring that distributed AI training, large language model, inference, and HPC workloads receive predictable bandwidth, low latency, and reliable collective communication performance.The ideal candidate will have experience designing, scaling, and operating backend network infrastructure for GPU clusters with approximately 10,000 or more GPUs, or comparable hyperscale AI and HPC environments.The primary focus of this position is high-speed Ethernet and RoCEv2 networking for AMD Instinct accelerator clusters. You will work across switches, NICs, optics, RDMA, Linux networking, PCIe and NUMA topology, ROCm, RCCL, SLURM, Kubernetes, storage networks, automation platforms, and observability systems.You will partner with AMD AI engineering, network engineering, data center, storage, security, platform, and application teams to ensure the backend network fabric is not a bottleneck to GPU workload performance.THE PERSON:You are a highly experienced, hands-on network engineer with deep expertise in data center networking, RDMA, RoCEv2, and large-scale GPU cluster fabrics with approximately 10,000 or more GPUs,.You understand how distributed GPU workloads generate traffic across the backend network and how application performance is affected by network topology, congestion, GPU-to-NIC locality, routing, switch buffering, traffic-class configuration, and collective communication patterns. You take responsibility for end-to-end outcomes, including architecture, implementation, qualification, production deployment, monitoring, incident response, capacity planning, and continuous improvement. You use telemetry and repeatable performance testing to validate designs and make data-driven engineering decisions.You are comfortable leading complex technical initiatives, mentoring engineers, documenting architecture and operating standards, and working across globally distributed organizations.KEY RESPONSIBILITIES:Architect, deploy, operate, and continuously improve high-performance backend networks for large-scale AMD Instinct GPU clusters.Design network fabrics capable of supporting AI and HPC environments ranging from individual GPU racks to clusters containing 10,000 or more GPUs.Own the backend network architecture from the GPU server and network interface card through the leaf-spine switching fabric.Design and optimize high-speed Ethernet fabrics using RoCEv2 and 100/200/400 GbE technologies.Develop scalable network topologies, including leaf-spine, Clos, fat-tree, rail-optimized, multi-plane, and non-blocking fabric architectures.Perform network topology modeling, oversubscription analysis, traffic-flow analysis, bandwidth planning, port-capacity planning, failure-domain analysis, and long-term growth forecasting.Configure, tune, validate, and troubleshoot lossless or near-lossless RoCEv2 environments, including PFC, ECN, DCQCN, QoS, ECMP, Switch buffer and queue management, DSCP and priority mappingsDesign and operate routing and switching environments using technologies such as BGP, ECMP, VLAN, VRF, EVPN, and VXLAN.Optimize end-to-end communication performance across GPUs, NICs, switches, CPUs, PCIe devices, storage systems, and the Linux networking stack.Lead production incident response, root-cause analysis, corrective actions, and preventive engineering improvements for GPU cluster networks.Plan and execute network expansions, cluster scale-outs, switch replacements, capacity upgrades, and fabric migrationsPREFERRED EXPERIENCE:Significant experience designing, deploying, and operating production data center networks for AI, GPU, HPC, cloud, or other large-scale distributed computing environments.Experience designing, scaling, or operating backend network infrastructure for GPU clusters containing approximately 10,000 or more GPUs, or similarly sized hyperscale compute environments.Deep knowledge of data center networking fundamentals; Routing and switching, VLANs and subnetting, BGP and ECMP, Quality of Service, MTU configuration, Switch buffering, Network segmentationStrong hands-on experience with RDMA and RoCEv2 in production environments.Demonstrated experience configuring, tuning, and troubleshooting PFC, ECN, DCQCN, QoS, switch buffers, NIC queues, RDMA traffic classes, and lossless or near-lossless Ethernet.Strong understanding of leaf-spine, Clos, fat-tree, rail-optimized, and multi-plane network architectures.Experience with network routing technologies such as BGP and ECMP and overlay technologies such as EVPN and VXLAN.Strong understanding of GPU cluster topology, including GPU-to-GPU, GPU-to-NIC, CPU-to-NIC, PCIe, NUMA, and network locality.Experience building monitoring and observability solutions using Prometheus, Grafana, streaming telemetry, gNMI, SNMP, sFlow, or equivalent platforms.Experience with Juniper data center switching platforms and Junos OS, including configuration and troubleshootingExperience with AMD Instinct accelerators, ROCm, RCCL, and AMD GPU software environments.Experience with AMD Pensando AI NICs, SmartNICs, DPUs, or other AMD Pensando networking technologies.Strong hands-on experience with Juniper data center switching platforms and Junos OS, including configuration and troubleshooting Experience designing backend networks specifically for large language model training and other communication-intensive distributed AI workloads.Experience with Ethernet fabric technologies such as BGP, EVPN, VXLAN, and modern leaf-spine data center architectures.ACADEMIC CREDITALS:Bachelor’s or Master’s degree in Computer Engineering, or a related field, or equivalent practical experience.LOCATION:San Jose, CA OR Austin, TXThis role is not eligible for visa sponsorship.#LI-MF2#LI-HYBRIDBenefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.
- ...technologies.As a Technical Marketing Engineer (TME) within the Software Product Management... ...organization for AMD’s Data Center GPU Business Unit, you will play a critical... ...design, deploy, and operate AMD-powered GPU clusters & networks to unlock performance and value across...Suggested
- Job Title: GPU Network Engineer Job Location: Sunnyvale, CA (hybrid)Job Salary: 200k-250k + BenefitsRequirements: Data center, HPC networking... ...build, and operate the high-speed fabrics that connect our GPU clusters. Must be local to Sunnyvale, CA or willing to relocate for...SuggestedLocal areaRelocation
- ...superintelligence. One person, one GPU. If you'd like to build... ...'s high performance cloud network Work on deploying and configuring... ...for new and existing clusters Ensure high availability of... ...on-call rotation for Network Engineering team You Have 10+...SeniorWork at officeLocal areaWork from homeFlexible hours
$200k - $250k
...months. The organization has built the leading network automation platform for AI clouds, with 35+ live GPU cluster deployments globally, sitting at the centre... .... This is a great opportunity for a Senior Network Engineer to move beyond traditional networking and guide...SeniorFull timeRemote work$146.7k - $339.3k
What you can expectThis role leads global network infrastructure and data center... ...team of 12 globally distributed network engineers and data center technicians. The team is... ...network topology and traffic engineering for GPU clusters (100+ GPUs), collective communications...SeniorFull timeWork at officeRemote work$152k - $241.5k
...Computing and Visualization. The GPU, our invention, serves as the... ...Communications Libraries and Networking team at NVIDIA. We deliver... ...for a motivated Performance engineer to influence the roadmap of our... ...multi-GPU and multi-node clusters.Study the interaction of our...SeniorFull time$224k - $356.5k
...next era of computing. An era in which our GPU acts as the brains of computers, robots,... ..., and parallelism efficiency on edge cluster configurationsProduce performance analysis... ...MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience...SeniorFull timeLocal area$160k - $170k
...provider of advanced server, storage, and networking solutions for Data Center, Cloud... ...seek talented, passionate, and committed engineers, technologists, and business leaders to... ...and lossless transport for large-scale GPU clusters.Manage, maintain, and support hands-on...SeniorWorldwide$131k - $175k
Company DescriptionArista Networks is an industry leader... ...awards, such as Best Engineering Team, Best Company for... ...You’ll Work With As a Senior Rack Solution Engineer... ...-scale AI and cloud clusters.What You’ll Do Lead the... ..., into high-density GPU environments, ensuring...SeniorRemote workFlexible hours$168k - $270.25k
...computing. An era in which our GPU acts as the brains of computers... ...Enterprise Experience (NVEX) Solutions Engineering team is looking for a senior Computer or Software Engineer.... ...expertise in ground-breaking network technology used in AI clusters. Our software engineers connect...SeniorFull timeWeekend work$140k - $224.25k
The NVIDIA Experience (NVEX) Solutions Engineering team is looking for a senior Computer or Software Engineer who is ready to become an authority in ground-breaking network technology used in AI clusters. Our team of software engineers bridge the gap between the customer...SeniorFull timeWeekend work- ..., Mirantis empowers platform engineering teams to deliver composable,... ...Mirantis delivers the automation, GPU orchestration, and policy-... ...accelerated compute: GPU clusters, bare metal, and managed AI infrastructure... ...to GPU architecture, networking, and orchestration. Own...Senior
- ...career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance... ...differences.Multi-server inference networking: Profile and optimize distributed... ...tensor parallelism across multi-node clusters. Analyze network-level bottlenecks...Senior
$184k - $287.5k
...scientists, researchers, and engineers can push the boundaries... ...a highly motivated Senior Solutions Architect to join the Cluster Design and Architecture team with a focus on networking technologies. As AI workloads... ...internal engineering efforts in GPU cluster building and...SeniorFull time$184k - $287.5k
We are looking for a Senior Solutions Architect specializing... ...pivotal technical expert uniting engineering, field teams, and customers... ...across interconnected GPU, CPU, and networking systems.Complete and maintain... ...-test high-performance clusters and establish performance baselines...SeniorFull time$128k - $201.25k
...computing. An era in which our GPU acts as the brains of computers,... ...lasting impact on the world.As a Senior Technical Marketing Engineer for Datacenter Networking, you will join a dedicated team... ...up, and maintain large storage clusters. This includes monitoring, logging...SeniorFull timeWork at office- DDN is seeking a Senior Software Engineering Manager to lead the engineering organization responsible... ...large-scale LLM inference across GPU clusters.In this role, you will lead geographically... ..., distributed caching, RDMA networking, GPUDirect Storage, NVIDIA BlueField...Senior
$256k - $414k
...scale.We are looking for a Senior Manager to lead the design,... ...operations of high-performance networking for GPU-based cloud infrastructure.... ...Oversee the design of intra-cluster and inter-cluster connectivity... ...Science or a related engineering field (or equivalent experience...SeniorFull timeLocal area$168k - $264.5k
NVIDIA is looking for a Senior Network Engineer to develop a cloud network infrastructure. The goal is to craft a reliable, scalable and efficient... ..., High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers...SeniorFull time$184k - $287.5k
...GB300 GPUs. NVIDIA seeks a Senior Software Engineer for our CSP (Cloud Service... ...of multi-rack, multi-tenant clusters: scheduler behavior, container... ...-services that expose new GPU capabilities.Drive joint... ...cloud-native stacks across networking (RDMA/RoCE), storage, and control...SeniorFull timeRemote work$224k - $356.5k
...to-end behavior across GPUs, networking, storage, and software stacks. We are seeking a Senior Performance Engineer to characterize workloads,... ...engineers, platform teams, and GPU architects to validate and... ...Experience analyzing large-scale AI clusters or distributed training and...SeniorFull timeRemote work$176k - $276k
NVIDIA is looking for an experienced HPC-AI Engineer to join the Networking Clusters Solutions Infrastructure team. we are focused on building supercomputers... ...the latest breakthroughs in artificial intelligence and GPU computing. Provide insights on at-scale system design...SeniorFull time$184k - $287.5k
We are seeking a Senior DevOps / Cloud Simulation Infrastructure Engineer to own the complete end-to-end cloud... ...deploying a robust, multi-GPU pipeline that supports structural... ...function-to-function networking, gRPC bottlenecks, and in-cluster proxy behavior.Artifact & Evidence...SeniorFull timeLocal area$184k - $287.5k
...workloads. We are looking for a Senior Software Engineer to lead the bring-up,... ...inference workloads across NVIDIA GPU platforms at the largest... ...that keep large clusters productive. This is a hands... ...performance across compute, memory, networking, and communication layers...SeniorFull timeRemote work$152k - $241.5k
...era of computing. An era in which our GPU acts as the brains of computers,... ...work.We are searching for outstanding senior system software engineer to join the NVIDIA's GPU Diagnostics... ...along with strong collaborative and networking abilities.Proven ability to thrive in...SeniorFull time- ...superintelligence. One person, one GPU.If you'd like to build the... ...day is currently Tuesday.Engineering at Lambda is responsible for... ...deploy, and operate Kubernetes clusters across AWS and Lambda's bare-... ...standards for resource management, networking, and RBAC across the platform...SeniorWork at officeLocal areaWork from homeFlexible hours
$184k - $287.5k
...era of computing. An era in which our GPU acts as the brains of computers, robots... ...world.We are looking for a dedicated engineer for the Senior Systems Software Engineer role, focusing... ...builds for NVIDIA GPUs, CPUs, and networking hardware. Engage early with HW/FW/SW/platform...SeniorFull timeRemote work$137k - $156k
...provider of advanced server, storage, and networking solutions for Data Center, Cloud... ...talented, passionate, and committed engineers, technologists, and business leaders... ...is seeking an experienced Senior Systems Engineer / GPU Platforms to support the bring-up, qualification...SeniorWorldwide$184k - $287.5k
...boundaries of innovation and engineering? At NVIDIA, we lead the... ...systems.As a Senior Hardware Systems Engineer... ...power distribution, and cluster‑level cooling systems.What... ...platforms such as LPU, GPU, TPU, or custom... ...designs, such as new power networks, liquid cooling, or optical...SeniorFull time$176k - $276k
Production engineering is a field that involves crafting, building, and... ..., data management, systems, networking, coding, database management,... ...internal and external-facing GPU cloud services meet reliability... ...support large-scale storage clusters, ensuring scalability, high...SeniorFull timeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Network Engineer - GPU Cluster Networking. Be the first to apply!
- network engineer San Jose, CA
- juniper network engineer San Jose, CA
- network engineer level San Jose, CA
- senior network engineer San Jose, CA
- network infrastructure engineer San Jose, CA
- senior network engineer remote San Jose, CA
- network implementation engineer San Jose, CA
- network engineer full time San Jose, CA
- network software engineer San Jose, CA
- remote cisco network engineer San Jose, CA


