Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Software Engineer - Cluster Networking

$184k - $287.5k

Nvidia

NVIDIA is a pioneer in accelerated computing, known for inventing the GPU and driving breakthroughs in gaming, computer graphics, high-performance computing, and artificial intelligence. Our technology powers everything from generative AI to autonomous systems, and we continue to shape the future of computing through innovation and collaboration. Within this mission, our team, Managed AI Research Superclusters (MARS), builds and scales the infrastructure, platforms, and tools that enable researchers and engineers to develop the next generation of AI/ML systems. By joining us, you'll help build solutions that power some of the most sophisticated computing workloads globally.We are looking for a senior networking engineer to lead the network architecture of our GPU superclusters. Our platform runs frontier model training across tens of thousands of GPUs on multiple clouds, and we are scaling it toward clusters of ten thousand nodes and beyond. At that size, networking stops being a configuration exercise and becomes the hardest engineering problem on the platform — and it is currently one of the areas where we most need depth.What you'll be doing:You will be the technical owner of how our clusters communicate internally and externally. This includes the CNI data plane, the overlay mesh, and the nodes connecting clusters across regions and providers.Own and evolve the Kubernetes networking architecture for GPU clusters running at multi-thousand-node scaleDesign, operate and scale the overlay network - CNI, mesh and VPN topologies (Tailscale, WireGuard), and the gateways that connect control and data planesDesign, operate and scale the L7 gateways/load balancers/tunnels (Envoy, Cloudflare)Find and eliminate scale ceilings: packet loss under load, control-plane saturation, IP address management exhaustion, and the failure modes that only appear above a few thousand nodesBuild the scale-test environments and validation suites that let us catch networking regressions before they reach production, rather than during a training runDiagnose hard, ambiguous problems across the stack - where a symptom in Slurm or a training job traces back to a mark collision, a stale route, or a saturated tunnelPartner with cloud and neocloud providers on network topology, requirements and capabilities as we bring up new clustersProvide senior technical judgement to a distributed team, and depth in the Custer Networking domain.What we need to see:BS/MS in Computer Science, Electrical Engineering or a related field, or equivalent experience6+ years of professional experience in systems, network or infrastructure software engineeringDeep command of Kubernetes networking architecture and CNI standards, with production experience operating Calico strongly preferredProficiency designing and maintaining modern mesh and VPN networking topologies - Tailscale, WireGuard or equivalentStrong Linux networking fundamentals: routing, netfilter and iptables/nftables, packet marking, network namespaces, and how these interact with container runtimesDemonstrated ability to debug distributed network problems at scale - packet capture, tracing, and correlating behaviour across many hosts to find a single root causeProficiency in Go, Python, C or a comparable systems languageClear written and verbal communication, and the ability to work effectively with engineers across multiple time zonesWays to stand out from the crowd:Direct experience architecting and operating massive-scale Kubernetes topologies across thousands of concurrent nodesExperience with high-performance fabrics in AI or HPC environments - InfiniBand, RoCE, or RDMA over converged networksUpstream contributions to Calico, Cilium, Tailscale, or Kubernetes networking SIGsExperience operating networking across multiple public clouds and on-premises environments simultaneouslyFamiliarity with Slurm or other HPC schedulers running on KubernetesYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until September 19, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, MA, Westford; US, TX, Austin; US, CO, Remote; US, NC, DurhamType: Full time

Vacancy posted 16 hours ago
Similar jobs that could be interesting for youBased on the Senior Software Engineer - Cluster Networking in Santa Clara, CA vacancy
  • $174k - $252k

     ...and the impact on hardware, network, or service operations and quality...  ...who hit issues in production clusters.Engage with the open source...  ....5 years of experience with software development in Go, C or...  ...Linux OSs.Google's software engineers develop the next-generation technologies... 
    Senior

    Google

    Sunnyvale, CA
    5 days ago
  •  ...your career.THE ROLE:We are seeking a Senior Network Engineer to join the AMD IT System Engineering...  ...networks supporting large-scale AMD GPU clusters. The engineer will own the network...  ...accelerators, ROCm, RCCL, and AMD GPU software environments.Experience with AMD Pensando... 
    Senior

    AMD

    San Jose, CA
    1 day ago
  • $152k - $241.5k

    NVIDIA is searching for a highly motivated, excellent Senior Software Engineer for design and verification to join the software tools group....  ...management, burning, configuration and debugging of all NVIDIA networking products.What you'll be doing:As a valued member of the... 
    Senior
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...work. Come join the team and see how you can make a lasting impact on the world.We are looking for a Senior Linux Kernel Software Engineer to join the Linux networking drivers R&D team. The work environment is versatile, informative, dynamic and challenging as our employees... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...work. Come join the team and see how you can make a lasting impact on the world.We are seeking an outstanding Software Engineer to join our US-based networking software team. As a technical leader, you will lead the transformation of AI networking systems. You will apply... 
    Senior
    Full time
    Shift work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $193.3k - $261.5k

     ...seeking an experienced engineer to work on distributed...  ...with high-speed networking or HPC interconnects is...  ...features for the largest clusters, with the largest customers...  ...develops hardware and software components that are...  ..., you can both expect senior mentorship and will be... 
    Senior
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  • $174k - $252k

     ...test, deploy, maintain and improve switch software.Manage individual project priorities,...  ...scalability and performance tuning of large-scale networks.Triage product or system issues and...  ...threading development.Google's software engineers develop the next-generation technologies... 
    Senior

    Google

    Sunnyvale, CA
    1 day ago
  • $224k - $356.5k

     ...make a lasting impact on the world.Join our outstanding team at NVIDIA to craft the future of computing! As a Senior Software Engineer - Traffic & Networking, you'll have the chance to define and deliver innovative solutions for our cloud infrastructure. Collaborate with... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $139k - $204k

     ...that drives innovation.  About the role As part of the Cluster Orchestration team, you will play a key role in advancing CoreWeave...  ...of what's possible with AI. What You'll Do As a Senior Software Engineer I (IC3), you will own multiple services within the orchestration... 
    Senior
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    15 days ago
  • $193.3k - $261.5k

    Amazon Leo is Amazon’s low Earth orbit satellite network. Our mission is to deliver fast, reliable internet connectivity to customers...  ...other public and private networks.A day in the lifeAs a senior software engineer you will be responsible for leading the design of embedded... 
    Senior
    Permanent employment
    Internship
    Local area
    Flexible hours

    Amazon

    Sunnyvale, CA
    3 days ago
  • $184k - $287.5k

    NVIDIA is seeking a Senior Software Engineer to help us develop distributed storage services for AI/ML. In this role you will work closely with...  ...and deployed a distributed service that runs on large-scale clusters, multi-petabyte to exabyte in size, with millions of... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...upcoming GB300 GPUs. NVIDIA seeks a Senior Software Engineer for our CSP (Cloud Service Provider)...  ...debugging of multi-rack, multi-tenant clusters: scheduler behavior, container runtime...  ...large-scale, cloud-native stacks across networking (RDMA/RoCE), storage, and control... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    5 days ago
  • $140k - $224.25k

     ...Cloud provides the infrastructure and software platform that enables enterprises to...  ...infrastructure.We are looking for a Senior Software Engineer to design and build the systems that...  ...capacity across cloud providers, regions, clusters, and products.Automate capacity-... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...AI revolution, building the software and systems that power the...  ...workloads. We are looking for a Senior Software Engineer to lead the bring-up,...  ...capabilities that keep large clusters productive. This is a hands...  ...across compute, memory, networking, and communication layers using... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    5 days ago
  • $224k - $356.5k

     ....NVIDIA's Local AI team is building the software stack that makes large language models and...  ..., and parallelism efficiency on edge cluster configurationsProduce performance analysis...  ..., or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience... 
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    5 days ago
  • $193.3k - $261.5k

     ...centers and all of the servers, storage, networking, power, and cooling equipment that...  .... You’ll join a diverse team of software, hardware, and network engineers, supply chain specialists, security...  ...an experienced, results-oriented, Senior Software Dev Engineer.What do we do... 
    Senior
    Internship
    Local area
    Flexible hours

    Amazon

    Santa Clara, CA
    1 day ago
  • $176k - $276k

     ...is part of NVIDIA’s Global Network Infrastructure (GNI) organization...  ...of this platform, including cluster provisioning and upgrades,...  ...enablement. We build software and automation to standardize...  ...are looking for a hands-on senior engineer to own the lifecycle and automation... 
    Senior
    Full time
    Remote work
    Weekend work

    Nvidia

    Santa Clara, CA
    3 days ago
  • $140k - $224.25k

    NVIDIA is looking for a top-tier Senior QA Software Test Engineer to join the NVIDIA-Cumulus system Ethernet group. This position will be part of our QA team while the main goal is testing of our Ethernet Switch/Router. You will be participating in requirements and design... 
    Senior
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    2 days ago
  • $193.3k - $261.5k

     ...power the world's largest machine learning clusters. Our team builds virtual platforms —...  ...models of these custom SoCs — that let software teams start development months before silicon...  ...silicon. We're looking for a software engineer to build and own the models and... 
    Senior
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    5 days ago
  • $186k - $388k

     ...event systems, shared memory), and collaborate across Product, Engineering, QA, and Ops to deliver resilient services spanning streaming,...  ...layers (e.g., queuing, event systems, shared memory clusters) and libraries that can be used across teamsReview technical specification... 
    Senior
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    1 day ago
  • $184k - $287.5k

    We are looking for a Senior Software Engineer to become part of our storage management plane team. The management plane is a web-based application...  ...solutions for improving our ability to handle large clusters of machines efficiently.What You Will Be Doing:Maintain and... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  • $184k - $287.5k

    At NVIDIA, we are redefining the future of technology, and our Senior Software Engineer, Networking role offers a uniquely ambitious opportunity to contribute to world-class innovations. If you thrive in a collaborative and inclusive environment, are driven to succeed,... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...operating system, server, RDMA, and network layers. Optimize low-...  ...cross-functionally, improve engineering practices, automate...  ...Strong debugging skills across software, operating system, server, and...  ...includes AI/ML, HPC, storage, GPU cluster, large-scale RoCE or InfiniBand... 
    Senior
    Full time

    Clockwork.io

    Palo Alto, CA
    7 days ago
  • $193.3k - $261.5k

     ...internship professional software development experience...  ...mentor, tech lead, or engineering manager. We require...  ...with high-speed networking or HPC interconnects such...  ...that run on our largest clusters and support our...  ...engineers and also learn from senior technical leaders.... 
    Senior
    Full time
    Internship
    Flexible hours

    Amazon.com Services LLC

    Cupertino, CA
    7 days ago
  • $170k - $205k

     ...team that believes in each other, come build with us at Crusoe. About This Role: We are seeking a Senior Software Engineer to build a platform for network tooling that supports edge, backbone, and data center operations and deployment tasks at Crusoe Cloud, a... 
    Senior
    Temporary work

    Crusoe

    Sunnyvale, CA
    23 days ago
  • $193.3k - $261.5k

     ...any single chip, the network between accelerators becomes...  ....We're looking for an engineer to work at the...  ...low-level data-movement software across accelerators, servers...  ...run on our largest clusters, for our largest...  ..., you can both expect senior mentorship and will be... 
    Senior
    Full time
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  • $153k - $204k

     ...company (Nasdaq: CRWV) in March 2025. Learn more at  What You'll Do: We are seeking a talented and experienced Senior Software Engineer to join our Network Datapath Team. As a Senior Software Engineer, you will play a critical role in designing, developing, and... 
    Senior
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    14 days ago
  • $176k - $276k

    NVIDIA is looking for an experienced HPC-AI Engineer to join the Networking Clusters Solutions Infrastructure team. we are focused on building supercomputers...  ...key player to the most exciting computing hardware and software to contribute to the latest breakthroughs in artificial... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $193.93k - $291.15k

     ...the Role The ability to monitor and assist our vehicles remotely plays a key role in our business strategy. As a Senior Software Engineer, Networking you will work on our in-house Teleoperations platform. You will work with a diverse team of engineers to build the... 
    Senior
    Full time
    Remote work

    Nuro

    Mountain View, CA
    1 day ago
  • $184k - $287.5k

    We are looking for a Senior System Software Engineer, Software Defined Networking to design, build, and operate highly performant and scalable SDN solutions for NVIDIA's AI Clouds hosting GPU-accelerated workloads — including hyperscale multi-node training, inference,... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Software Engineer - Cluster Networking. Be the first to apply!