Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

GPU Network Engineer: RDMA/NVLink at Scale

$350k

Thinking Machines Lab

Thinking Machines Lab is seeking a Network Engineer in San Francisco to manage and improve our GPU network fabric. The role requires in-depth knowledge of large-scale deployments and the ability to debug complex network issues. A collaborative environment is emphasized, where initiative and effective communication with cloud providers are key. The position offers a competitive salary ranging from $350,000 to $475,000 per year, depending on skills and experience. Benefits include unlimited PTO and health coverage, alongside visa sponsorship. #J-18808-Ljbffr Thinking Machines Lab

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the GPU Network Engineer: RDMA/NVLink at Scale in San Francisco, CA vacancy
  • Thinking Machines Lab Inc. is seeking a network engineer to own the lowest layers of the network stack for large-scale training and inference. You will ensure interconnect reliability across GPU fabrics, debugging NICs, and building instrumentation for faster troubleshooting... 
    Suggested

    Thinking Machines Lab Inc.

    San Francisco, CA
    3 days ago
  •  ...help build the platform engineers turn to to ship AI products...  ...multi-modal workloads scale, the network is the computer. We are...  ...foundational engineers to lead our GPU Networking efforts, making RDMA a first-class building...  ...communication across NVLink and InfiniBand for our... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    9 hours ago
  • $350k

     ...the Role We're looking for a network engineer to own the lowest layers of the...  ...network stack that our large-scale training and inference depend...  ...at scale, across large GPU fabrics - both the RDMA/RoCE fabric between nodes and the NVLink/NVSwitch domains within them.... 
    Suggested
    Visa sponsorship
    Work visa
    Relocation package

    Thinking Machines Lab Inc.

    San Francisco, CA
    2 days ago
  • $150k - $300k

    Prime Intellect in San Francisco seeks a Solutions Architect for GPU Infrastructure who will transform client requirements into robust systems capable of training advanced AI models. Responsibilities include designing GPU cluster architectures, deploying orchestration systems... 
    Suggested

    Prime Intellect

    San Francisco, CA
    4 days ago
  • $250k - $320k

     ...of AI infrastructure: large-scale AI datacenters and the orchestration...  ...Gimlet Labs is seeking a Network Engineer to design, build, and scale...  ...have Experience with AI/HPC, GPU, or large‑scale distributed infrastructure...  ...tooling. Familiarity with RDMA, RoCE, InfiniBand, or other... 
    Suggested

    Gimlet Labs, Inc.

    San Francisco, CA
    5 days ago
  • $190k - $280k

     ...Senior Network Engineer San Francisco About the Role Together AI...  ...boundaries. You will work on large-scale, multi-vendor data center...  ...~ Foundational knowledge of RDMA networking and technologies such...  .... Experience supporting GPU clusters, HPC environments, distributed... 
    Full time

    Together AI

    San Francisco, CA
    3 days ago
  •  ...superintelligence. One person, one GPU.If you'd like to build...  ...ll DoHelp to build and scale Lambda's high...  ...and configuring networking hardware for new and existing...  ...call rotation for Network Engineering teamYouHave 10+ years...  ...partitions), GPUDirect RDMA concepts.Experience... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda

    San Francisco, CA
    4 days ago
  •  ...Principal Network Engineer – AI InfrastructureThe Network Operations and Engineering...  ..., supporting tightly coupled GPU clusters where network...  ...of our Infiniband and RDMA-based network fabrics. You will...  ...reviewing, and evolving large-scale Infiniband and RoCE fabric architectures... 

    Nscale

    San Francisco, CA
    4 days ago
  •  ...About the Team The Core Network Engineering team owns the end-to-end networking...  ...xPU networking used for large-scale training and inference workloads...  ...across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-performance GPU interconnects Define and operationalize... 
    Full time

    OpenAI

    San Francisco, CA
    9 hours ago
  •  ...hardware and software. Speed and scale are our key differentiators....  ..., non-blocking backend networks for clusters of 100k+ accelerators...  ...lifecycle from customer requirements (GPU shape, workload, scale,...  ...lossless Ethernet fabrics for RDMA (RoCEv2): PFC, ECN tuning, traffic... 
    Local area

    Fluidstack

    San Francisco, CA
    3 days ago
  • $10k - $20k

     ...services. The company is looking for a Network Automation Engineer to design and build high-speed, low-latency network fabrics for large-scale GPU-accelerated computing. The role focuses...  ...specific experience with InfiniBand or RDMA (RoCEv2). CI/CD Pipeline Expertise:... 
    Full time
    San Francisco, CA
    14 days ago
  • $200k

     ...Ready to architect the high-speed networks powering the AI era? Join a trailblazing leader in GPU-accelerated computing, designing...  ..., or SaltStack. High-Scale Networking: A strong foundation...  ...specific experience with InfiniBand or RDMA (RoCEv2). CI/CD Pipeline Expertise... 
    Full time
    San Francisco, CA
    more than 2 months ago
  • $195k - $235k

     ...urgency, who believe in the scale of our ambition and thrive...  ...Crusoe Cloud is seeking a Staff Network Production Operations Engineer to help own production...  ...backbone, data center fabric, and GPU cluster interconnects. This...  ....Experience operating RDMA/RoCE lossless fabrics for... 
    Temporary work
    Worldwide

    Crusoe

    San Francisco, CA
    1 day ago
  • A technology solutions provider is looking for a Network Engineer to enhance and maintain a large-scale network. This role involves managing both wired and wireless infrastructures, conducting assessments, and ensuring network security. Candidates should have a degree... 

    CGS Federal (Contact Government Services)

    San Francisco, CA
    5 days ago
  • Sciforium is seeking a GPU Kernel Engineer to push performance on modern accelerators. You will design and optimize custom GPU kernels, from low...  ...to integrating ops in ML frameworks used for large-scale training and inference. Ideal candidates have 5+ years in GPU... 

    Sciforium

    San Francisco, CA
    5 days ago
  • $179k - $218k

     ...a sense of urgency, who believe in the scale of our ambition and thrive on a path not...  ...seeking a Senior Staff Data Center Operations Engineer, GPU Hardware Architecture to be the...  ...identifying "pre-failure" patterns in HBM or NVLink components before they impact customer... 
    Temporary work

    Crusoe

    San Francisco, CA
    5 days ago
  • $210k - $240k

     ...Build and help define the network foundation behind a...  ...AI platform supporting GPU infrastructure, distributed...  .... This is a network engineering role first . We are looking...  ...or HPC environments RDMA or RoCEv2 networking,...  ...in a startup or rapidly scaling technical environment Experience... 
    Immediate start

    Stratitech

    San Francisco, CA
    3 days ago
  • $250k - $300k

     ...who believe in the scale of our ambition and...  ...Deployment Automation Engineer for the Compute...  ...-scale, multi-node GPU clusters. You will...  ...multi-node context.Networking Knowledge: Strong understanding of RDMA, RoCE, and InfiniBand...  ...MNNVL (Multi-Node NVLink) or specialized AI... 
    Temporary work

    Crusoe

    San Francisco, CA
    3 days ago
  • $170k - $250k

     ...revenue within six months and is scaling rapidly with a small, high-...  ...This company is looking for an engineer to work directly with the CTO on complex GPU virtualisation challenges. The role...  ...complex issues across low-level networking, GPU drivers, and distributed systems... 
    Full time
    Visa sponsorship
    Flexible hours
    San Francisco, CA
    28 days ago
  • $250k

     ...opportunities? Join a rapidly scaling AI cloud infrastructure...  ...provider building a next-generation GPU platform designed for AI...  ...Senior / Staff Site Reliability Engineer to support and scale large-...  ...providers Strong understanding of networking fundamentals (DNS, TCP/IP,... 
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  •  ...largest AI infrastructure networks. The team owns day-to-day...  ...deliver highly available GPU infrastructure for AI...  ...Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support...  ...supporting RoCE v2 or RDMA‑based Ethernet fabrics, with... 
    Permanent employment

    OpenAI

    San Francisco, CA
    5 days ago
  • Linuxcareers in San Francisco is building AI research infrastructure. You will design, deploy, and operate large-scale GPU clusters powering training, evaluation, and serving for the research team. The role emphasizes extending orchestration with Kubernetes/Slurm, building... 

    Linuxcareers

    San Francisco, CA
    6 days ago
  •  ...serverless runtime that launches GPU-backed containers in less than 1 second and quickly scales out to thousands of GPUs....  ...Deploy and validate data center network infrastructure (front-end, back...  ...Operations, ICT, Hardware, and Network Engineering to identify blockers early,... 

    BEAM inc.

    San Francisco, CA
    5 days ago
  •  ...and help build the platform engineers turn to to ship AI...  ...Baseten is building its own GPU infrastructure for large-scale inference. As we move into...  ...workload symptoms that look like network problems, but are not....  ...networking and inference software. RDMA data paths, GPUDirect... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    9 hours ago
  • Hyperbolic is seeking a very senior Infrastructure Engineer to scale a GPU cloud marketplace. You will build a multi-tenant provisioning and virtualization layer, turning global GPU inventories into an orchestrated pool for AI developers and researchers. You will own the... 

    Hyperbolic

    San Francisco, CA
    2 days ago
  •  ...spanning hardware and software. Speed and scale are our key differentiators. Come be a...  ...matters to the world. The Production Engineering Team Examples of key exciting problems...  ...of 10s to 100s of GWs: at our scale, a GPU failure isn't a ticket. It's a throughput... 
    Full time
    Local area

    Fluidstack

    San Francisco, CA
    9 hours ago
  • $180k - $250k

     ...production, and do it at scale without compromise....  ...You are a hands-on engineer who builds the...  ...keep a large fleet of GPU servers healthy and...  ...errors, disk failures, network issues, thermals)...  ...health monitoring, DCGM, NVLink/NVSwitch diagnostics, RDMA, InfiniBand/RoCEv2... 
    Full time
    Local area
    Relocation package

    Falò

    San Francisco, CA
    9 hours ago
  •  ...Senior Network EngineerHouston; New York; San Francisco; SeattleAbout NscaleNscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI...  ...networking platformsExperience with large-scale DCI, long-haul optical transport, or... 

    Nscale

    San Francisco, CA
    4 days ago
  • $195k - $235k

     ...urgency, who believe in the scale of our ambition and thrive...  ...Crusoe Cloud is seeking a Staff Network Production Operations Engineer to help own production...  ...backbone, data center fabric, and GPU cluster interconnects. This...  .... ~ Experience operating RDMA/RoCE lossless fabrics for... 
    Temporary work
    Worldwide

    Crusoe

    San Francisco, CA
    8 days ago
  • $240k - $280k

     ...We're looking for a Software Engineer to build the systems that treat...  ...single API call to stand up, scale, or tear down a cluster, and the...  ..., inference bring-up to GPU driver/CUDA stack, health validation...  ..., Redfish/IPMI, BMC) and/or networking fundamentals (VLANs, BGP, fabric... 
    Full time

    Together Ai

    San Francisco, CA
    9 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to GPU Network Engineer: RDMA/NVLink at Scale. Be the first to apply!