AI Networking Software Engineer: NCCL & Multi-GPU Training
Jobleads-US
Meta is seeking a Software Engineer, SystemML - AI Networking to join the AI Networking Software team within the DC networking organization. You will help own the NCCL-based stack that enables multi-GPU and multi-node communication, integral to Meta's distributed ML workloads and tightly integrated with PyTorch.
You will provide technical leadership for the library, focus on GenAI/LLM scaling reliability and performance, optimize GPU interconnects, collaborate across teams, and translate complex
#J-18808-Ljbffr Jobleads-US$121.99k - $181k
...will be a member of the AI Networking Software team and part of the... ...stack around NCCL (NVIDIA Collective Communications... ...), which enables multi-GPU and multi-node data... ...multi-GPU distributed training. In other words,... .... And we are seeking engineers to work on the space...TrainingHourly payLocal area$193.93k - $291.15k
...profound opportunity for AI to drive positive... ...the systems that train the models at the... ...- from distributed GPU training and closed... ...infrastructure, spanning multi-generation... ...Science, Electrical Engineering, or a closely related... ...internals, including NCCL and collective...TrainingWork experience placementImmediate startFlexible hours$120k - $195k
...largest professional network, built to create... ...LinkedIn’s AI model training, feature engineering and serving with... ...infra, compute software, and hardware to... ...the power of our GPU fleet with thousands... ...cuTile, cuDNN, NCCL, RDMA,... ...millions of QPS, multi terabytes of data...TrainingFor contractorsWork experience placementWork at officeFlexible hours$180k - $300k
...large portion of training compute is wasted... ...Microsoft, Amazon, and AI visionaries like... ...research and data engineering necessary to solve... ...and maintain our multi-cloud infrastructure... ...-level debugging-networking issues, memory... ...inference clusters, GPU orchestration)...TrainingWork at officeWork from homeRelocation package- ...JPMorganChase is seeking a Software Engineer III to design and operate an end-to-end ML training platform on AWS and other clouds. You will run GPU workloads, optimize performance, and enable Gen AI workflows within a governed, secure environment. You will collaborate...Training
$61k - $101k
...year Requirements: We need formal training or certification in software engineering concepts, plus 3+ years of applied... ...using enterprise-authorized AI-assisted development tools in the work... ...exposure to deploying or operating GPU workloads in Kubernetes environments...TrainingFull time$174k - $252k
...improve switch software.Manage individual... ...of large-scale networks.Triage product... ...and multi-threading development... ...Google's software engineers develop the next... ...projects enabling AI networking and... ...for TPU and GPU workloads.Our Platforms... ...education or training. US: $174000 -...Training- ...opportunity for you to take your software engineering career to the next level. As... ...enterprise-authorized AI coding assist tools within the... ...capabilities, and skillsFormal training or certification on software... ...C++Exposure to deploying or operating GPU workloads in Kubernetes...Training
$182k - $242k
...Essential Cloud for AI™. Built for... ...-performance GPU infrastructure... .... Our stack is engineered for speed, scale... ...Inference (and Training) runs, including... ...team; decompose multi-service work... ...GPU/accelerator software, or performance... ...experience. NCCL and collective-...TrainingPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours- ...the world's largest AI chip, 56 times... ...industry-leading training and inference speeds... ...times faster than GPU-based hyperscale cloud... ...announced a multi-year partnership with... ...RoleWe're hiring a Software Engineer to help contribute... ..., security, networking, debugging, and productionization...Training
- ...for you to take your software engineering career to the next level... ...within the AI/ML data platform team... ...scalable, reliable ML training systems and pipelines... ...training workloads (often GPU-based), improve performance... ...memory, I/O throughput, networking, scheduling).Experience...Training
$135k - $155k
...enabling human life on Mars.SOFTWARE ENGINEER (PLATFORM TEAM) The Platform... ...every team at SpaceX to harness AI effectively. This team... ..., deployed applications, and trained models on managed, reliable computeCollaborate... ...-focused architecture, and multi-provider integrations (cloud...TrainingPermanent employmentTemporary work$188.5k - $282.7k
...govern, and remediate AI agents. Our team operates... ...and Governing Secure Multi-Cloud Foundations (30%... ...disparate environments.Engineering secure "Landing Zones"... ...configurations.Architecting network security perimeters... ...relevant education or training.US Pay Range$188,500—$2...Training- ...world's largest AI chip, 56 times... ...industry-leading training and inference... ...times faster than GPU-based... ...recently announced a multi-year partnership... ...RoleThe Host and Network IO Team... ...the WSE. As a software developer on the... ...or Electrical Engineering + 1 year industry...Training
$150k - $250k
...mission is to create AI systems that can... ..., and focused on engineering excellence. This... ...large-scale networks that underpin training and inference infrastructure... ...that connect GPU clusters, plus... ...not a network-software (telemetry/ZTP platform... ...specialist RoCE/NCCL ownership is a...TrainingTemporary workNight shift- ...world's largest AI chip, 56 times... ...industry-leading training and inference... ...times faster than GPU-based... ...recently announced a multi-year partnership... ...fleet expands, the software used to monitor... ...software engineer to build the platforms... ...their compute, networking, and hardware...Training
- ...mission is to create AI systems that can accurately... ..., and focused on engineering excellence. This organization... ...ROLE: As part of the Network Software and Services for AI (... ...the world's largest GPU supercomputing network fabrics used for AI training and serving customer...TrainingTemporary work
$345.04k - $399.42k
...for everyone.As a Principal Software Engineer on the Compute team, you... ...technical anchor for Roblox's GPU and AI accelerator capabilities.... ..., Machine Bootstrap, Networking, and Cloud to drive GPU strategy... ..., InfiniBand, RoCE), and multi-node training and inference patterns....TrainingFull timeWork experience placementH1bWork at officeLocal areaVisa sponsorshipMonday to Friday- ...world's largest AI chip, 56 times larger... ...industry-leading training and inference... ...times faster than GPU-based hyperscale... ...recently announced a multi-year partnership... ...Wafer-Scale Engine.We are hiring a Software Engineer to productionize... ...ROCm, GPU nodes, networking, and rack-scale...Training
$180k
...mission is to create AI systems that can... ...motivated, and focused on engineering excellence. This... ..., and optimize the network fabric that powers large-scale AI training and inference... ...AI clusters (100k+ GPU scale). Own vendor... ...emerging technologies (multi-core/hollow-core...TrainingTemporary work$156k - $190k
...vertically integrated AI infrastructure... ...Cloud Support Engineer , you are a... ..., SRE, Networking, Fleet, and Product... ...with SRE, Software teams (Storage... ...Troubleshoot NCCL, IB, GPU driver/firmware... ..., distributed training failures. Support... ...of resolving multi-layer,...TrainingTemporary work$182k - $242k
...Essential Cloud for AI™. Built for... ...looking for a Senior Engineer to be a driving... ...distributed training and inference workloads... ...time, whether a GPU fleet, a fabric,... ...that make network fabric and GPU-level... ...GPU fleets or multi-region clusters.... ...CUDA kernels, NCCL/SHARP, RDMA/NUMA...TrainingPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$182k - $242k
...Essential Cloud for AI™. Built for... ...looking for a Senior Engineer for CoreWeave's... ...-to-end MLPerf Training and Inference runs... ...understanding of networked systems and performance... ...-critical GPU systems (CUDA, NCCL, NVLink/PCIe, memory... ...GPU clusters or multi-region...TrainingPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours- Job Title: GPU Network Engineer Job Location: Sunnyvale, CA (hybrid)Job Salary: 200k-250k + BenefitsRequirements... ..., with direct experience in GPU or AI cluster environments.Hands-on experience... ...is the core of the role.Experience with multi-host networking for bare metal, KVM, and...Local areaRelocation
- ...the world's largest AI chip, 56 times... ...industry-leading training and inference speeds... ...times faster than GPU-based hyperscale cloud... ...announced a multi-year partnership with... ...with Security and Engineering teams to deliver secure... ...are serious about software make their own...Training
$140k - $300k
...Distinguished Engineer – Data Center Network Architecture & Engineering... ...Engineering, AI/ML, and... ...network architecture, software-defined networking... ...services supporting multi-site and multi-... ...supporting AI/ML, GPU, or high-performance... ..., education and training, the work...TrainingHourly payWork experience placementLocal area$20k
...goal of enabling human life on Mars.NETWORK ENGINEER, AI INFRASTRUCTURE (STARSHIELD)... ...next-gen communication and sensing software, and more.As a Network Engineer you... ...solutions for AI clusters (100k+ GPU scale)Collaborate with ML training teams to translate workload...TrainingPermanent employmentTemporary workInternshipImmediate startWeekend work$153.12k - $196.75k
...About the role:Foundation AI builds the platforms,... ...areas such as model training and inference, LLM services... ..., safety, creator, engine, discovery, and economy... ...distributed systems, and GPU infrastructure.Improve... ...years of experience in software engineering, distributed...TrainingFull timeWork experience placementInternshipH1bWork at officeLocal areaVisa sponsorshipMonday to Friday- ...world's largest AI chip, 56 times... ...industry-leading training and inference... ...times faster than GPU-based... ...recently announced a multi-year partnership... ...hiring a Staff Engineer to own major areas... ...experience in software engineering,... ...environments, including networking, compute...Training
- ...world's largest AI chip, 56 times... ...industry-leading training and inference... ...times faster than GPU-based... ...recently announced a multi-year partnership... ...hiring a Principal Engineer for our... ...experience in software engineering, with... ...environments, including networking, compute...Training
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Networking Software Engineer: NCCL & Multi-GPU Training. Be the first to apply!
- senior network engineer remote Menlo Park, CA
- network software engineer Menlo Park, CA
- cisco network engineer Menlo Park, CA
- network infrastructure engineer Menlo Park, CA
- network consulting engineer Menlo Park, CA
- network engineer Menlo Park, CA
- network engineer - transport Menlo Park, CA
- network developer Menlo Park, CA
- data center network engineer Menlo Park, CA
- software engineer full time Menlo Park, CA




