Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Network Engineer - GPU Cluster Networking

Advanced Micro Devices , Inc.

Senior Network Engineer At AMD, we believe technology has the power to solve the world's most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. Whether you're designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we'll advance your career. We are seeking a Senior Network Engineer to join the AMD IT System Engineering team. This role is responsible for the architecture, deployment, optimization, automation, and production operation of high-performance backend networks supporting large-scale AMD GPU clusters. The engineer will own the network path from the GPU server and NIC through the data center switching fabric, ensuring that distributed AI training, large language model, inference, and HPC workloads receive predictable bandwidth, low latency, and reliable collective communication performance. The ideal candidate will have experience designing, scaling, and operating backend network infrastructure for GPU clusters with approximately 10,000 or more GPUs, or comparable hyperscale AI and HPC environments. The primary focus of this position is high-speed Ethernet and RoCEv2 networking for AMD Instinct accelerator clusters. You will work across switches, NICs, optics, RDMA, Linux networking, PCIe and NUMA topology, ROCm, RCCL, SLURM, Kubernetes, storage networks, automation platforms, and observability systems. You will partner with AMD AI engineering, network engineering, data center, storage, security, platform, and application teams to ensure the backend network fabric is not a bottleneck to GPU workload performance. You are a highly experienced, hands-on network engineer with deep expertise in data center networking, RDMA, RoCEv2, and large-scale GPU cluster fabrics with approximately 10,000 or more GPUs. You understand how distributed GPU workloads generate traffic across the backend network and how application performance is affected by network topology, congestion, GPU-to-NIC locality, routing, switch buffering, traffic-class configuration, and collective communication patterns. You take responsibility for end-to-end outcomes, including architecture, implementation, qualification, production deployment, monitoring, incident response, capacity planning, and continuous improvement. You use telemetry and repeatable performance testing to validate designs and make data-driven engineering decisions. You are comfortable leading complex technical initiatives, mentoring engineers, documenting architecture and operating standards, and working across globally distributed organizations. Key Responsibilities: Architect, deploy, operate, and continuously improve high-performance backend networks for large-scale AMD Instinct GPU clusters. Design network fabrics capable of supporting AI and HPC environments ranging from individual GPU racks to clusters containing 10,000 or more GPUs. Own the backend network architecture from the GPU server and network interface card through the leaf-spine switching fabric. Design and optimize high-speed Ethernet fabrics using RoCEv2 and 100/200/400 GbE technologies. Develop scalable network topologies, including leaf-spine, Clos, fat-tree, rail-optimized, multi-plane, and non-blocking fabric architectures. Perform network topology modeling, oversubscription analysis, traffic-flow analysis, bandwidth planning, port-capacity planning, failure-domain analysis, and long-term growth forecasting. Configure, tune, validate, and troubleshoot lossless or near-lossless RoCEv2 environments, including PFC, ECN, DCQCN, QoS, ECMP, Switch buffer and queue management, DSCP and priority mappings Design and operate routing and switching environments using technologies such as BGP, ECMP, VLAN, VRF, EVPN, and VXLAN. Optimize end-to-end communication performance across GPUs, NICs, switches, CPUs, PCIe devices, storage systems, and the Linux networking stack. Lead production incident response, root-cause analysis, corrective actions, and preventive engineering improvements for GPU cluster networks. Plan and execute network expansions, cluster scale-outs, switch replacements, capacity upgrades, and fabric migrations Preferred Experience: Significant experience designing, deploying, and operating production data center networks for AI, GPU, HPC, cloud, or other large-scale distributed computing environments. Experience designing, scaling, or operating backend network infrastructure for GPU clusters containing approximately 10,000 or more GPUs, or similarly sized hyperscale compute environments. Deep knowledge of data center networking fundamentals; Routing and switching, VLANs and subnetting, BGP and ECMP, Quality of Service, MTU configuration, Switch buffering, Network segmentation Strong hands-on experience with RDMA and RoCEv2 in production environments. Demonstrated experience configuring, tuning, and troubleshooting PFC, ECN, DCQCN, QoS, switch buffers, NIC queues, RDMA traffic classes, and lossless or near-lossless Ethernet. Strong understanding of leaf-spine, Clos, fat-tree, rail-optimized, and multi-plane network architectures. Experience with network routing technologies such as BGP and ECMP and overlay technologies such as EVPN and VXLAN. Strong understanding of GPU cluster topology, including GPU-to-GPU, GPU-to-NIC, CPU-to-NIC, PCIe, NUMA, and network locality. Experience building monitoring and observability solutions using Prometheus, Grafana, streaming telemetry, gNMI, SNMP, sFlow, or equivalent platforms. Experience with Juniper data center switching platforms and Junos OS, including configuration and troubleshooting Experience with AMD Instinct accelerators, ROCm, RCCL, and AMD GPU software environments. Experience with AMD Pensando AI NICs, SmartNICs, DPUs, or other AMD Pensando networking technologies. Experience designing backend networks specifically for large language model training and other communication-intensive distributed AI workloads. Experience with Ethernet fabric technologies such as BGP, EVPN, VXLAN, and modern leaf-spine data center architectures. Academic Credits: Bachelor's or Master's degree in Computer Engineering, or a related field, or equivalent practical experience. Location: San Jose, CA OR Austin, TX This role is not eligible for visa sponsorship. Advanced Micro Devices , Inc.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior Network Engineer - GPU Cluster Networking in Austin, TX vacancy
  • $184k - $287.5k

     ...computing, known for inventing the GPU and driving breakthroughs in...  ...that enable researchers and engineers to develop the next...  ...globally.We are looking for a senior networking engineer to lead the network...  ...and we are scaling it toward clusters of ten thousand nodes and... 
    Senior
    Full time
    Remote work

    Nvidia

    Austin, TX
    3 days ago
  • Senior Principal Network Engineer Austin, Texas, United States Graphcore is one of the world's leading innovators in Artificial Intelligence compute...  ...performance computing (HPC) network fabrics supporting GPU clusters. You will work across hardware, networking, and AI... 
    Senior
    Flexible hours

    Graphcore

    Austin, TX
    3 days ago
  • $224k - $356.5k

     ...next era of computing. An era in which our GPU acts as the brains of computers, robots,...  ..., and parallelism efficiency on edge cluster configurationsProduce performance analysis...  ...MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience... 
    Senior
    Full time
    Local area

    Nvidia

    Austin, TX
    2 days ago
  • $168k - $270.25k

     ...computing. An era in which our GPU acts as the brains of computers...  ...Enterprise Experience (NVEX) Solutions Engineering team is looking for a senior Computer or Software Engineer....  ...expertise in ground-breaking network technology used in AI clusters. Our software engineers connect... 
    Senior
    Full time
    Weekend work

    Nvidia

    Austin, TX
    2 days ago
  • $140k - $224.25k

    The NVIDIA Experience (NVEX) Solutions Engineering team is looking for a senior Computer or Software Engineer who is ready to become an authority in ground-breaking network technology used in AI clusters. Our team of software engineers bridge the gap between the customer... 
    Senior
    Full time
    Weekend work

    Nvidia

    Austin, TX
    1 day ago
  • $114.6k - $234.6k

    AI2NE strives to be a global leader in the RDMA cluster networking domain and enable seamless, accelerated High-Performance Compute (HPC),...  ...as the job remains posted.Career Level - IC4Collaborate with engineers from L1 optical engineering team, network design, delivery and... 
    Senior
    Temporary work
    Flexible hours

    Oracle Corporation

    Austin, TX
    7 hours ago
  •  ..., and play a key role in supporting the network technologies that power next-generation...  ...networks that enable AMD's most advanced engineering initiatives.In this role, you will collaborate...  ...connectivity for servers, racks, clusters, and validation environments.Install, configure... 
    Senior
    Visa sponsorship
    Afternoon shift

    AMD

    Austin, TX
    2 days ago
  • $184k - $287.5k

     ...GB300 GPUs. NVIDIA seeks a Senior Software Engineer for our CSP (Cloud Service...  ...of multi-rack, multi-tenant clusters: scheduler behavior, container...  ...-services that expose new GPU capabilities.Drive joint...  ...cloud-native stacks across networking (RDMA/RoCE), storage, and control... 
    Senior
    Full time
    Remote work

    Nvidia

    Austin, TX
    2 days ago
  • $224k - $356.5k

     ...to-end behavior across GPUs, networking, storage, and software stacks. We are seeking a Senior Performance Engineer to characterize workloads,...  ...engineers, platform teams, and GPU architects to validate and...  ...Experience analyzing large-scale AI clusters or distributed training and... 
    Senior
    Full time
    Remote work

    Nvidia

    Austin, TX
    3 days ago
  • $224k - $356.5k

     ...Our data-center platforms bring together GPUs, CPUs, networking, systems, and software to solve some of the world’s...  ...exciting computing problems. We are looking for a Senior System Software Engineer to join NVIDIA’s GPU Performance and Power Management Software team. You... 
    Senior
    Full time
    Remote work

    Nvidia

    Austin, TX
    19 hours ago
  • $170k - $230k

     ...DescriptionWe are hiring a Senior Platform Engineer to join the Autonomous Vehicle...  ...of production-grade clusters.Required: Strong proficiency...  ...hardware-level performance (GPU passthrough) and clean cloud...  ...broader GCP ecosystem, including networking and IAM; equivalent... 
    Senior
    Full time
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    General Motors

    Austin, TX
    2 days ago
  •  ...Position Overview We are seeking a Senior GPU Systems & Fabric Engineer to serve as the critical bridge between our physical GPU/network infrastructure and the Kubernetes abstraction...  ...tenant inference workloads and maximize cluster utilization. Profile and tune kernel-... 
    Senior
    Full time
    Local area

    Bitdeer Technologies Group

    Austin, TX
    a month ago
  • $126.2k - $264.1k

    Job Description What You’ll Do As a Senior Principal Network Reliability Engineer, you will: Provide technical leadership for the deployment, operation,...  ...critical services. Experience supporting AI infrastructure, GPU networking, or large-scale data center deployments.... 
    Senior
    Temporary work
    Flexible hours

    Oracle

    Austin, TX
    5 days ago
  • $184k - $287.5k

    Senior Software Engineer NVIDIA is at the forefront of the generative AI revolution...  ...workloads across NVIDIA GPU platforms at the largest...  ...capabilities that keep large clusters productive. This is a hands...  ...performance across compute, memory, networking, and communication layers... 
    Senior

    NVIDIA

    Austin, TX
    3 days ago
  •  ...Senior Network Engineer Location: Remote, USA (must currently reside in the US for consideration) Contract Position: 12-18months+ We are seeking a highly motivated and skilled Senior Network Engineer with strong experience in network design and security. The... 
    Senior
    Contract work
    Remote work

    Genius Road

    Austin, TX
    4 days ago
  •  ...Enterprise Network Engineer - Must be bilingual - Korean/English Seeking an experienced Enterprise Network Engineer fluent in Korean and English to join a high-availability manufacturing network environment. This position involves collaborating with stakeholders... 
    Senior

    Saxon Global

    Austin, TX
    17 hours ago
  • $103.62k - $172.63k

     ...need at LPL Financial to shape your success while helping clients pursue their financial goals. Job Overview:LPL is seeking a Senior Network Engineer to join our Datacenter Network Engineering team. This individual will be responsible for the design, engineering, and... 
    Senior
    Full time
    Work experience placement
    Work from home

    LPL Financial

    Austin, TX
    1 day ago
  •  ...Senior Level Network Engineer We are looking for an experienced Senior Network Engineer to execute network assessment process, topology, and optimization.  Job Summary: The Senior Network Engineer is responsible for executing data center network assessments and... 
    Senior
    Remote work
    Visa sponsorship

    VICTORY

    Austin, TX
    4 days ago
  • $150k - $217k

     ...responsible for development of solutions for Google's Enterprise network, executing new features qualification and validation through...  ...technical and communication aspects.Coach and collaborate with peer engineers towards achieving common goals in assigned projects.Minimum... 
    Senior

    Google

    Austin, TX
    3 days ago
  •  ...enhance maritime operations through autonomous and intelligent platforms.We are seeking an experienced and results-driven Senior Network Engineer to join our IT team and lead the design, deployment, and management of mission-critical network infrastructure. The ideal candidate... 
    Senior
    Permanent employment
    Temporary work

    Saronic Technologies

    Austin, TX
    4 days ago
  • $150k - $217k

    Lead the requirement analysis, engineering design, and development of solutions for Google’s Enterprise network infrastructure, executing new feature qualification, deployment, and optimization.Deliver secure and dependable network services by scaling systems sustainably... 
    Senior
    Worldwide

    Google

    Austin, TX
    3 days ago
  • Bitdeer Technologies Group in the United States is seeking a senior network engineer to design and operate a multi-tenant, multi-region GPU cloud network supporting distributed AI training. You will implement tenant isolation, high-performance interconnects, and cross-... 
    Senior

    Bitdeer (NASDAQ: BTDR)

    Austin, TX
    6 days ago
  • $186.07k - $218.9k

     ...Lead on the Exchange team within Institutional, you'll own the network infrastructure, deployment, and operational tooling behind a unified...  ...systems, and participate in an on-call rotation. Mentor engineers and partner cross-functionally with SRE/Infra, Product, and... 
    Senior
    Local area

    Coinbase

    Austin, TX
    3 days ago
  •  ...Supply Planning to lead a team of planners overseeing data center GPU programs. You will translate demand, capacity, and material...  ...making, and process improvements to scale planning across multi-site networks and ensure on-time delivery and optimal inventory. #J-18808-... 
    Senior

    Advanced Micro Devices, Inc.

    Austin, TX
    4 days ago
  •  ...a highly skilled and motivated Cloud Senior DevOps Engineer to join our AI Cloud team. In this high...  ...computing resources (e.g., GPU clusters) to support high-performance AI workloads...  ...), or Cloud Infrastructure roles. Networking & OS: Expert-level knowledge of Linux... 
    Senior
    Full time
    Local area

    Bitdeer Technologies Group

    Austin, TX
    a month ago
  • $70 - $80 per hour

     ...Immediate need for a talented Senior Network Engineer This is a 12 months contract opportunity with long-term potential and is located in Round Rock, TX (Onsite). Please review the job description below and contact me ASAP if you are interested. Job ID: 26-245... 
    Senior
    Contract work
    Local area
    Immediate start

    Pyramid Consulting

    Round Rock, TX
    17 hours ago
  •  ...Network Engineer Austin, TX Delart is home to a team of world-class engineers and project leaders dedicated to developing the next generation of advanced networking technologies, consumer devices, and innovative technology solutions. Trusted by some of the world... 
    Senior
    Permanent employment
    Temporary work
    Flexible hours

    Delart

    Austin, TX
    3 days ago
  •  ...and high-performance computing products within the Data Center GPU group. You will drive cross-functional execution from SoC planning...  ...deployment, partnering with silicon, firmware, software, networking, and manufacturing teams. The role requires strong leadership,... 
    Senior

    AMD

    Austin, TX
    5 days ago
  • A tech company in Austin, Texas is seeking an experienced Network Engineer to join its Engineering team. The ideal candidate will be responsible for implementing network solutions, supporting multi-site designs, and maintaining network security. Required qualifications... 
    Senior

    Sigma Information Group, Inc.

    Austin, TX
    4 days ago
  • $112.5k - $202.5k

     ...excited about working with cutting-edge network technologies? Do you enjoy acting quickly...  ...to solve problems? Join our Network Engineering team The Network Engineering team is...  ...services. Partner with the best As a Senior Network Engineer you will lead efforts... 
    Senior
    Work experience placement
    Work at office

    Akamai

    Austin, TX
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Network Engineer - GPU Cluster Networking. Be the first to apply!