Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Network Engineer - GPU Cluster Networking

AMD

ADVANCE YOUR CAREER. ADVANCE THE WORLD. At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career.THE ROLE:We are seeking a Senior Network Engineer to join the AMD IT System Engineering team.This role is responsible for the architecture, deployment, optimization, automation, and production operation of high-performance backend networks supporting large-scale AMD GPU clusters. The engineer will own the network path from the GPU server and NIC through the data center switching fabric, ensuring that distributed AI training, large language model, inference, and HPC workloads receive predictable bandwidth, low latency, and reliable collective communication performance.The ideal candidate will have experience designing, scaling, and operating backend network infrastructure for GPU clusters with approximately 10,000 or more GPUs, or comparable hyperscale AI and HPC environments.The primary focus of this position is high-speed Ethernet and RoCEv2 networking for AMD Instinct accelerator clusters. You will work across switches, NICs, optics, RDMA, Linux networking, PCIe and NUMA topology, ROCm, RCCL, SLURM, Kubernetes, storage networks, automation platforms, and observability systems.You will partner with AMD AI engineering, network engineering, data center, storage, security, platform, and application teams to ensure the backend network fabric is not a bottleneck to GPU workload performance.THE PERSON:You are a highly experienced, hands-on network engineer with deep expertise in data center networking, RDMA, RoCEv2, and large-scale GPU cluster fabrics with approximately 10,000 or more GPUs,.You understand how distributed GPU workloads generate traffic across the backend network and how application performance is affected by network topology, congestion, GPU-to-NIC locality, routing, switch buffering, traffic-class configuration, and collective communication patterns. You take responsibility for end-to-end outcomes, including architecture, implementation, qualification, production deployment, monitoring, incident response, capacity planning, and continuous improvement. You use telemetry and repeatable performance testing to validate designs and make data-driven engineering decisions.You are comfortable leading complex technical initiatives, mentoring engineers, documenting architecture and operating standards, and working across globally distributed organizations.KEY RESPONSIBILITIES:Architect, deploy, operate, and continuously improve high-performance backend networks for large-scale AMD Instinct GPU clusters.Design network fabrics capable of supporting AI and HPC environments ranging from individual GPU racks to clusters containing 10,000 or more GPUs.Own the backend network architecture from the GPU server and network interface card through the leaf-spine switching fabric.Design and optimize high-speed Ethernet fabrics using RoCEv2 and 100/200/400 GbE technologies.Develop scalable network topologies, including leaf-spine, Clos, fat-tree, rail-optimized, multi-plane, and non-blocking fabric architectures.Perform network topology modeling, oversubscription analysis, traffic-flow analysis, bandwidth planning, port-capacity planning, failure-domain analysis, and long-term growth forecasting.Configure, tune, validate, and troubleshoot lossless or near-lossless RoCEv2 environments, including PFC, ECN, DCQCN, QoS, ECMP, Switch buffer and queue management, DSCP and priority mappingsDesign and operate routing and switching environments using technologies such as BGP, ECMP, VLAN, VRF, EVPN, and VXLAN.Optimize end-to-end communication performance across GPUs, NICs, switches, CPUs, PCIe devices, storage systems, and the Linux networking stack.Lead production incident response, root-cause analysis, corrective actions, and preventive engineering improvements for GPU cluster networks.Plan and execute network expansions, cluster scale-outs, switch replacements, capacity upgrades, and fabric migrationsPREFERRED EXPERIENCE:Significant experience designing, deploying, and operating production data center networks for AI, GPU, HPC, cloud, or other large-scale distributed computing environments.Experience designing, scaling, or operating backend network infrastructure for GPU clusters containing approximately 10,000 or more GPUs, or similarly sized hyperscale compute environments.Deep knowledge of data center networking fundamentals; Routing and switching, VLANs and subnetting, BGP and ECMP, Quality of Service, MTU configuration, Switch buffering, Network segmentationStrong hands-on experience with RDMA and RoCEv2 in production environments.Demonstrated experience configuring, tuning, and troubleshooting PFC, ECN, DCQCN, QoS, switch buffers, NIC queues, RDMA traffic classes, and lossless or near-lossless Ethernet.Strong understanding of leaf-spine, Clos, fat-tree, rail-optimized, and multi-plane network architectures.Experience with network routing technologies such as BGP and ECMP and overlay technologies such as EVPN and VXLAN.Strong understanding of GPU cluster topology, including GPU-to-GPU, GPU-to-NIC, CPU-to-NIC, PCIe, NUMA, and network locality.Experience building monitoring and observability solutions using Prometheus, Grafana, streaming telemetry, gNMI, SNMP, sFlow, or equivalent platforms.Experience with Juniper data center switching platforms and Junos OS, including configuration and troubleshootingExperience with AMD Instinct accelerators, ROCm, RCCL, and AMD GPU software environments.Experience with AMD Pensando AI NICs, SmartNICs, DPUs, or other AMD Pensando networking technologies.Strong hands-on experience with Juniper data center switching platforms and Junos OS, including configuration and troubleshooting Experience designing backend networks specifically for large language model training and other communication-intensive distributed AI workloads.Experience with Ethernet fabric technologies such as BGP, EVPN, VXLAN, and modern leaf-spine data center architectures.ACADEMIC CREDITALS:Bachelor’s or Master’s degree in Computer Engineering, or a related field, or equivalent practical experience.LOCATION:San Jose, CA OR Austin, TXThis role is not eligible for visa sponsorship.#LI-MF2#LI-HYBRIDBenefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior Network Engineer - GPU Cluster Networking in San Jose, CA vacancy
  •  ...technologies.As a Technical Marketing Engineer (TME) within the Software Product Management...  ...organization for AMD’s Data Center GPU Business Unit, you will play a critical...  ...design, deploy, and operate AMD-powered GPU clusters & networks to unlock performance and value across... 
    Suggested

    AMD

    Santa Clara, CA
    4 days ago
  • Job Title: GPU Network Engineer Job Location: Sunnyvale, CA (hybrid)Job Salary: 200k-250k + BenefitsRequirements: Data center, HPC networking...  ...build, and operate the high-speed fabrics that connect our GPU clusters. Must be local to Sunnyvale, CA or willing to relocate for... 
    Suggested
    Local area
    Relocation

    CyberCoders

    San Jose, CA
    3 days ago
  •  ...superintelligence. One person, one GPU. If you'd like to build...  ...'s high performance cloud network Work on deploying and configuring...  ...for new and existing clusters Ensure high availability of...  ...on-call rotation for Network Engineering team You Have 10+... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Corporation

    San Jose, CA
    4 days ago
  • $200k - $250k

     ...months. The organization has built the leading network automation platform for AI clouds, with 35+ live GPU cluster deployments globally, sitting at the centre...  .... This is a great opportunity for a Senior Network Engineer to move beyond traditional networking and guide... 
    Senior
    Full time
    Remote work
    Santa Clara, CA
    3 days ago
  • $146.7k - $339.3k

    What you can expectThis role leads global network infrastructure and data center...  ...team of 12 globally distributed network engineers and data center technicians. The team is...  ...network topology and traffic engineering for GPU clusters (100+ GPUs), collective communications... 
    Senior
    Full time
    Work at office
    Remote work

    Zoom

    San Jose, CA
    5 hours ago
  • $152k - $241.5k

     ...Computing and Visualization. The GPU, our invention, serves as the...  ...Communications Libraries and Networking team at NVIDIA. We deliver...  ...for a motivated Performance engineer to influence the roadmap of our...  ...multi-GPU and multi-node clusters.Study the interaction of our... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $224k - $356.5k

     ...next era of computing. An era in which our GPU acts as the brains of computers, robots,...  ..., and parallelism efficiency on edge cluster configurationsProduce performance analysis...  ...MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience... 
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    1 day ago
  • $160k - $170k

     ...provider of advanced server, storage, and networking solutions for Data Center, Cloud...  ...seek talented, passionate, and committed engineers, technologists, and business leaders to...  ...and lossless transport for large-scale GPU clusters.Manage, maintain, and support hands-on... 
    Senior
    Worldwide

    Super Micro Computer

    San Jose, CA
    2 days ago
  • $131k - $175k

    Company DescriptionArista Networks is an industry leader...  ...awards, such as Best Engineering Team, Best Company for...  ...You’ll Work With As a Senior Rack Solution Engineer...  ...-scale AI and cloud clusters.What You’ll Do Lead the...  ..., into high-density GPU environments, ensuring... 
    Senior
    Remote work
    Flexible hours

    Arista Networks

    Santa Clara, CA
    3 days ago
  • $168k - $270.25k

     ...computing. An era in which our GPU acts as the brains of computers...  ...Enterprise Experience (NVEX) Solutions Engineering team is looking for a senior Computer or Software Engineer....  ...expertise in ground-breaking network technology used in AI clusters. Our software engineers connect... 
    Senior
    Full time
    Weekend work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $140k - $224.25k

    The NVIDIA Experience (NVEX) Solutions Engineering team is looking for a senior Computer or Software Engineer who is ready to become an authority in ground-breaking network technology used in AI clusters. Our team of software engineers bridge the gap between the customer... 
    Senior
    Full time
    Weekend work

    Nvidia

    Santa Clara, CA
    5 hours ago
  •  ..., Mirantis empowers platform engineering teams to deliver composable,...  ...Mirantis delivers the automation, GPU orchestration, and policy-...  ...accelerated compute: GPU clusters, bare metal, and managed AI infrastructure...  ...to GPU architecture, networking, and orchestration. Own... 
    Senior

    Mirantis

    San Jose, CA
    22 days ago
  •  ...career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance...  ...differences.Multi-server inference networking: Profile and optimize distributed...  ...tensor parallelism across multi-node clusters. Analyze network-level bottlenecks... 
    Senior

    AMD

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...scientists, researchers, and engineers can push the boundaries...  ...a highly motivated Senior Solutions Architect to join the Cluster Design and Architecture team with a focus on networking technologies. As AI workloads...  ...internal engineering efforts in GPU cluster building and... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

    We are looking for a Senior Solutions Architect specializing...  ...pivotal technical expert uniting engineering, field teams, and customers...  ...across interconnected GPU, CPU, and networking systems.Complete and maintain...  ...-test high-performance clusters and establish performance baselines... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $128k - $201.25k

     ...computing. An era in which our GPU acts as the brains of computers,...  ...lasting impact on the world.As a Senior Technical Marketing Engineer for Datacenter Networking, you will join a dedicated team...  ...up, and maintain large storage clusters. This includes monitoring, logging... 
    Senior
    Full time
    Work at office

    Nvidia

    Santa Clara, CA
    1 day ago
  • DDN is seeking a Senior Software Engineering Manager to lead the engineering organization responsible...  ...large-scale LLM inference across GPU clusters.In this role, you will lead geographically...  ..., distributed caching, RDMA networking, GPUDirect Storage, NVIDIA BlueField... 
    Senior

    DataDirect Networks

    Santa Clara, CA
    5 hours ago
  • $256k - $414k

     ...scale.We are looking for a Senior Manager to lead the design,...  ...operations of high-performance networking for GPU-based cloud infrastructure....  ...Oversee the design of intra-cluster and inter-cluster connectivity...  ...Science or a related engineering field (or equivalent experience... 
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    4 days ago
  • $168k - $264.5k

    NVIDIA is looking for a Senior Network Engineer to develop a cloud network infrastructure. The goal is to craft a reliable, scalable and efficient...  ..., High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...GB300 GPUs. NVIDIA seeks a Senior Software Engineer for our CSP (Cloud Service...  ...of multi-rack, multi-tenant clusters: scheduler behavior, container...  ...-services that expose new GPU capabilities.Drive joint...  ...cloud-native stacks across networking (RDMA/RoCE), storage, and control... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $224k - $356.5k

     ...to-end behavior across GPUs, networking, storage, and software stacks. We are seeking a Senior Performance Engineer to characterize workloads,...  ...engineers, platform teams, and GPU architects to validate and...  ...Experience analyzing large-scale AI clusters or distributed training and... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $176k - $276k

    NVIDIA is looking for an experienced HPC-AI Engineer to join the Networking Clusters Solutions Infrastructure team. we are focused on building supercomputers...  ...the latest breakthroughs in artificial intelligence and GPU computing. Provide insights on at-scale system design... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

    We are seeking a Senior DevOps / Cloud Simulation Infrastructure Engineer to own the complete end-to-end cloud...  ...deploying a robust, multi-GPU pipeline that supports structural...  ...function-to-function networking, gRPC bottlenecks, and in-cluster proxy behavior.Artifact & Evidence... 
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...workloads. We are looking for a Senior Software Engineer to lead the bring-up,...  ...inference workloads across NVIDIA GPU platforms at the largest...  ...that keep large clusters productive. This is a hands...  ...performance across compute, memory, networking, and communication layers... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

     ...era of computing. An era in which our GPU acts as the brains of computers,...  ...work.We are searching for outstanding senior system software engineer to join the NVIDIA's GPU Diagnostics...  ...along with strong collaborative and networking abilities.Proven ability to thrive in... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    5 hours ago
  •  ...superintelligence. One person, one GPU.If you'd like to build the...  ...day is currently Tuesday.Engineering at Lambda is responsible for...  ...deploy, and operate Kubernetes clusters across AWS and Lambda's bare-...  ...standards for resource management, networking, and RBAC across the platform... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  • $184k - $287.5k

     ...era of computing. An era in which our GPU acts as the brains of computers, robots...  ...world.We are looking for a dedicated engineer for the Senior Systems Software Engineer role, focusing...  ...builds for NVIDIA GPUs, CPUs, and networking hardware. Engage early with HW/FW/SW/platform... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $137k - $156k

     ...provider of advanced server, storage, and networking solutions for Data Center, Cloud...  ...talented, passionate, and committed engineers, technologists, and business leaders...  ...is seeking an experienced Senior Systems Engineer / GPU Platforms to support the bring-up, qualification... 
    Senior
    Worldwide

    Super Micro Computer

    San Jose, CA
    5 hours ago
  • $184k - $287.5k

     ...boundaries of innovation and engineering? At NVIDIA, we lead the...  ...systems.As a Senior Hardware Systems Engineer...  ...power distribution, and cluster‑level cooling systems.What...  ...platforms such as LPU, GPU, TPU, or custom...  ...designs, such as new power networks, liquid cooling, or optical... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $176k - $276k

    Production engineering is a field that involves crafting, building, and...  ..., data management, systems, networking, coding, database management,...  ...internal and external-facing GPU cloud services meet reliability...  ...support large-scale storage clusters, ensuring scalability, high... 
    Senior
    Full time
    Flexible hours

    Nvidia

    Santa Clara, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Network Engineer - GPU Cluster Networking. Be the first to apply!