Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Networking Software Engineer: NCCL & Multi-GPU Training

Jobleads-US

Meta is seeking a Software Engineer, SystemML - AI Networking to join the AI Networking Software team within the DC networking organization. You will help own the NCCL-based stack that enables multi-GPU and multi-node communication, integral to Meta's distributed ML workloads and tightly integrated with PyTorch.

You will provide technical leadership for the library, focus on GenAI/LLM scaling reliability and performance, optimize GPU interconnects, collaborate across teams, and translate complex

#J-18808-Ljbffr Jobleads-US
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the AI Networking Software Engineer: NCCL & Multi-GPU Training in Menlo Park, CA vacancy
  • $121.99k - $181k

     ...will be a member of the AI Networking Software team and part of the...  ...stack around NCCL (NVIDIA Collective Communications...  ...), which enables multi-GPU and multi-node data...  ...multi-GPU distributed training. In other words,...  .... And we are seeking engineers to work on the space... 
    Training
    Hourly pay
    Local area

    Jobleads-US

    Menlo Park, CA
    2 days ago
  • $193.93k - $291.15k

     ...profound opportunity for AI to drive positive...  ...the systems that train the models at the...  ...- from distributed GPU training and closed...  ...infrastructure, spanning multi-generation...  ...Science, Electrical Engineering, or a closely related...  ...internals, including NCCL and collective... 
    Training
    Work experience placement
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    1 day ago
  • $120k - $195k

     ...largest professional network, built to create...  ...LinkedIn’s AI model training, feature engineering and serving with...  ...infra, compute software, and hardware to...  ...the power of our GPU fleet with thousands...  ...cuTile, cuDNN, NCCL, RDMA,...  ...millions of QPS, multi terabytes of data... 
    Training
    For contractors
    Work experience placement
    Work at office
    Flexible hours

    Linkedin

    Mountain View, CA
    1 day ago
  • $180k - $300k

     ...large portion of training compute is wasted...  ...Microsoft, Amazon, and AI visionaries like...  ...research and data engineering necessary to solve...  ...and maintain our multi-cloud infrastructure...  ...-level debugging-networking issues, memory...  ...inference clusters, GPU orchestration)... 
    Training
    Work at office
    Work from home
    Relocation package

    datologyai

    Redwood City, CA
    3 days ago
  •  ...JPMorganChase is seeking a Software Engineer III to design and operate an end-to-end ML training platform on AWS and other clouds. You will run GPU workloads, optimize performance, and enable Gen AI workflows within a governed, secure environment. You will collaborate... 
    Training

    Jobleads-US

    Palo Alto, CA
    4 days ago
  • $61k - $101k

     ...year Requirements: We need formal training or certification in software engineering concepts, plus 3+ years of applied...  ...using enterprise-authorized AI-assisted development tools in the work...  ...exposure to deploying or operating GPU workloads in Kubernetes environments... 
    Training
    Full time

    J.P. Morgan

    Palo Alto, CA
    7 days ago
  • $174k - $252k

     ...improve switch software.Manage individual...  ...of large-scale networks.Triage product...  ...and multi-threading development...  ...Google's software engineers develop the next...  ...projects enabling AI networking and...  ...for TPU and GPU workloads.Our Platforms...  ...education or training. US: $174000 -... 
    Training

    Google

    Sunnyvale, CA
    1 day ago
  •  ...opportunity for you to take your software engineering career to the next level. As...  ...enterprise-authorized AI coding assist tools within the...  ...capabilities, and skillsFormal training or certification on software...  ...C++Exposure to deploying or operating GPU workloads in Kubernetes... 
    Training

    JP Morgan Chase

    Palo Alto, CA
    2 days ago
  • $182k - $242k

     ...Essential Cloud for AI™. Built for...  ...-performance GPU infrastructure...  .... Our stack is engineered for speed, scale...  ...Inference (and Training) runs, including...  ...team; decompose multi-service work...  ...GPU/accelerator software, or performance...  ...experience. NCCL and collective-... 
    Training
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    25 days ago
  •  ...the world's largest AI chip, 56 times...  ...industry-leading training and inference speeds...  ...times faster than GPU-based hyperscale cloud...  ...announced a multi-year partnership with...  ...RoleWe're hiring a Software Engineer to help contribute...  ..., security, networking, debugging, and productionization... 
    Training

    Cerebras Systems

    Sunnyvale, CA
    2 days ago
  •  ...for you to take your software engineering career to the next level...  ...within the AI/ML data platform team...  ...scalable, reliable ML training systems and pipelines...  ...training workloads (often GPU-based), improve performance...  ...memory, I/O throughput, networking, scheduling).Experience... 
    Training

    JP Morgan Chase

    Palo Alto, CA
    2 days ago
  • $135k - $155k

     ...enabling human life on Mars.SOFTWARE ENGINEER (PLATFORM TEAM) The Platform...  ...every team at SpaceX to harness AI effectively. This team...  ..., deployed applications, and trained models on managed, reliable computeCollaborate...  ...-focused architecture, and multi-provider integrations (cloud... 
    Training
    Permanent employment
    Temporary work

    SpaceX

    Palo Alto, CA
    4 days ago
  • $188.5k - $282.7k

     ...govern, and remediate AI agents. Our team operates...  ...and Governing Secure Multi-Cloud Foundations (30%...  ...disparate environments.Engineering secure "Landing Zones"...  ...configurations.Architecting network security perimeters...  ...relevant education or training.US Pay Range$188,500—$2... 
    Training

    Rubrik

    Palo Alto, CA
    4 days ago
  •  ...world's largest AI chip, 56 times...  ...industry-leading training and inference...  ...times faster than GPU-based...  ...recently announced a multi-year partnership...  ...RoleThe Host and Network IO Team...  ...the WSE. As a software developer on the...  ...or Electrical Engineering + 1 year industry... 
    Training

    Cerebras Systems

    Sunnyvale, CA
    1 day ago
  • $150k - $250k

     ...mission is to create AI systems that can...  ..., and focused on engineering excellence. This...  ...large-scale networks that underpin training and inference infrastructure...  ...that connect GPU clusters, plus...  ...not a network-software (telemetry/ZTP platform...  ...specialist RoCE/NCCL ownership is a... 
    Training
    Temporary work
    Night shift

    SpaceXAI

    Palo Alto, CA
    5 days ago
  •  ...world's largest AI chip, 56 times...  ...industry-leading training and inference...  ...times faster than GPU-based...  ...recently announced a multi-year partnership...  ...fleet expands, the software used to monitor...  ...software engineer to build the platforms...  ...their compute, networking, and hardware... 
    Training

    Cerebras Systems

    Sunnyvale, CA
    3 days ago
  •  ...mission is to create AI systems that can accurately...  ..., and focused on engineering excellence. This organization...  ...ROLE: As part of the Network Software and Services for AI (...  ...the world's largest GPU supercomputing network fabrics used for AI training and serving customer... 
    Training
    Temporary work

    SpaceXAI

    Palo Alto, CA
    a month ago
  • $345.04k - $399.42k

     ...for everyone.As a Principal Software Engineer on the Compute team, you...  ...technical anchor for Roblox's GPU and AI accelerator capabilities....  ..., Machine Bootstrap, Networking, and Cloud to drive GPU strategy...  ..., InfiniBand, RoCE), and multi-node training and inference patterns.... 
    Training
    Full time
    Work experience placement
    H1b
    Work at office
    Local area
    Visa sponsorship
    Monday to Friday

    Roblox

    San Mateo, CA
    4 days ago
  •  ...world's largest AI chip, 56 times larger...  ...industry-leading training and inference...  ...times faster than GPU-based hyperscale...  ...recently announced a multi-year partnership...  ...Wafer-Scale Engine.We are hiring a Software Engineer to productionize...  ...ROCm, GPU nodes, networking, and rack-scale... 
    Training

    Cerebras Systems

    Sunnyvale, CA
    4 days ago
  • $180k

     ...mission is to create AI systems that can...  ...motivated, and focused on engineering excellence. This...  ..., and optimize the network fabric that powers large-scale AI training and inference...  ...AI clusters (100k+ GPU scale). Own vendor...  ...emerging technologies (multi-core/hollow-core... 
    Training
    Temporary work

    SpaceXAI

    Palo Alto, CA
    a month ago
  • $156k - $190k

     ...vertically integrated AI infrastructure...  ...Cloud Support Engineer , you are a...  ..., SRE, Networking, Fleet, and Product...  ...with SRE, Software teams (Storage...  ...Troubleshoot NCCL, IB, GPU driver/firmware...  ..., distributed training failures. Support...  ...of resolving multi-layer,... 
    Training
    Temporary work

    Crusoe

    Sunnyvale, CA
    a month ago
  • $182k - $242k

     ...Essential Cloud for AI™. Built for...  ...looking for a Senior Engineer to be a driving...  ...distributed training and inference workloads...  ...time, whether a GPU fleet, a fabric,...  ...that make network fabric and GPU-level...  ...GPU fleets or multi-region clusters....  ...CUDA kernels, NCCL/SHARP, RDMA/NUMA... 
    Training
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    20 days ago
  • $182k - $242k

     ...Essential Cloud for AI™. Built for...  ...looking for a Senior Engineer for CoreWeave's...  ...-to-end MLPerf Training and Inference runs...  ...understanding of networked systems and performance...  ...-critical GPU systems (CUDA, NCCL, NVLink/PCIe, memory...  ...GPU clusters or multi-region... 
    Training
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    a month ago
  • Job Title: GPU Network Engineer Job Location: Sunnyvale, CA (hybrid)Job Salary: 200k-250k + BenefitsRequirements...  ..., with direct experience in GPU or AI cluster environments.Hands-on experience...  ...is the core of the role.Experience with multi-host networking for bare metal, KVM, and... 
    Local area
    Relocation

    CyberCoders

    Sunnyvale, CA
    2 days ago
  •  ...the world's largest AI chip, 56 times...  ...industry-leading training and inference speeds...  ...times faster than GPU-based hyperscale cloud...  ...announced a multi-year partnership with...  ...with Security and Engineering teams to deliver secure...  ...are serious about software make their own... 
    Training

    Cerebras Systems

    Sunnyvale, CA
    3 days ago
  • $140k - $300k

     ...Distinguished Engineer – Data Center Network Architecture & Engineering...  ...Engineering, AI/ML, and...  ...network architecture, software-defined networking...  ...services supporting multi-site and multi-...  ...supporting AI/ML, GPU, or high-performance...  ..., education and training, the work... 
    Training
    Hourly pay
    Work experience placement
    Local area

    Jobleads-US

    Palo Alto, CA
    3 days ago
  • $20k

     ...goal of enabling human life on Mars.NETWORK ENGINEER, AI INFRASTRUCTURE (STARSHIELD)...  ...next-gen communication and sensing software, and more.As a Network Engineer you...  ...solutions for AI clusters (100k+ GPU scale)Collaborate with ML training teams to translate workload... 
    Training
    Permanent employment
    Temporary work
    Internship
    Immediate start
    Weekend work

    SpaceX

    Palo Alto, CA
    4 days ago
  • $153.12k - $196.75k

     ...About the role:Foundation AI builds the platforms,...  ...areas such as model training and inference, LLM services...  ..., safety, creator, engine, discovery, and economy...  ...distributed systems, and GPU infrastructure.Improve...  ...years of experience in software engineering, distributed... 
    Training
    Full time
    Work experience placement
    Internship
    H1b
    Work at office
    Local area
    Visa sponsorship
    Monday to Friday

    Roblox

    San Mateo, CA
    3 days ago
  •  ...world's largest AI chip, 56 times...  ...industry-leading training and inference...  ...times faster than GPU-based...  ...recently announced a multi-year partnership...  ...hiring a Staff Engineer to own major areas...  ...experience in software engineering,...  ...environments, including networking, compute... 
    Training

    Cerebras Systems

    Sunnyvale, CA
    4 days ago
  •  ...world's largest AI chip, 56 times...  ...industry-leading training and inference...  ...times faster than GPU-based...  ...recently announced a multi-year partnership...  ...hiring a Principal Engineer for our...  ...experience in software engineering, with...  ...environments, including networking, compute... 
    Training

    Cerebras Systems

    Sunnyvale, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Networking Software Engineer: NCCL & Multi-GPU Training. Be the first to apply!