Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

GPU Network Engineer: RDMA/NVLink at Scale

$350k

Thinking Machines Lab

Thinking Machines Lab is seeking a Network Engineer in San Francisco to manage and improve our GPU network fabric. The role requires in-depth knowledge of large-scale deployments and the ability to debug complex network issues. A collaborative environment is emphasized, where initiative and effective communication with cloud providers are key. The position offers a competitive salary ranging from $350,000 to $475,000 per year, depending on skills and experience. Benefits include unlimited PTO and health coverage, alongside visa sponsorship. #J-18808-Ljbffr Thinking Machines Lab

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the GPU Network Engineer: RDMA/NVLink at Scale in San Francisco, CA vacancy
  • $350k

    Thinkingmachines is seeking a network engineer to manage the lowest layers of the network stack pivotal for training and inference. You will ensure interconnect reliability for our large-scale GPU fabrics. The ideal candidate will have a degree in computer science or engineering... 
    Suggested

    Thinkingmachines

    San Francisco, CA
    4 days ago
  •  ...help build the platform engineers turn to to ship AI products...  ...multi-modal workloads scale, the network is the computer. We are...  ...foundational engineers to lead our GPU Networking efforts, making RDMA a first-class building...  ...communication across NVLink and InfiniBand for our... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    16 hours ago
  • $350k

     ...goals. We are scientists, engineers, and builders who’ve...  ...We're looking for a network engineer to own the lowest...  ...stack that our large‑scale training and inference...  ...scale, across large GPU fabrics — both the RDMA/RoCE fabric between nodes and the NVLink/NVSwitch domains... 
    Suggested
    Local area
    Visa sponsorship
    Relocation package

    Thinking Machines Lab

    San Francisco, CA
    4 days ago
  • Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance. This... 
    Suggested

    Vast.ai Inc.

    San Francisco, CA
    2 days ago
  • $150k - $300k

    Prime Intellect in San Francisco seeks a Solutions Architect for GPU Infrastructure who will transform client requirements into robust systems capable of training advanced AI models. Responsibilities include designing GPU cluster architectures, deploying orchestration systems... 
    Suggested

    Prime Intellect

    San Francisco, CA
    3 days ago
  •  ...About the Team The Core Network Engineering team owns the end-to-end networking...  ...xPU networking used for large-scale training and inference workloads...  ...across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-performance GPU interconnects Define and operationalize... 
    Full time

    OpenAI

    San Francisco, CA
    16 hours ago
  •  ...will design, deploy, and operate the network infrastructure underpinning Sesterce's GPU AI factories across Europe —...  ...physical cabling to BGP policies and RDMA fabric tuning. What you will do Design...  ...or 400G/800G Ethernet at scale Deep familiarity with RDMA, RoCE v... 

    Sesterce Group

    San Francisco, CA
    16 hours ago
  •  ...of AI infrastructure: large-scale AI datacenters and the orchestration...  ...Gimlet Labs is seeking a Network Engineer to design, build, and scale...  ...have Experience with AI/HPC, GPU, or large‑scale distributed infrastructure...  ...tooling. Familiarity with RDMA, RoCE, InfiniBand, or other... 

    Gimlet Labs

    San Francisco, CA
    3 days ago
  • Applied Compute is seeking a research scientist to design, implement, and optimize the large-scale training infrastructure powering our reinforcement learning stack in a San Francisco office. You’ll work with researchers to ensure the RL system is fast, reliable, and capable... 
    Work at office

    Applied Compute

    San Francisco, CA
    16 hours ago
  • $300 per month

     ...who believe in the scale of our ambition and...  ...Staff Hardware Systems Engineer to strengthen...  ...across Crusoe Cloud’s GPU- and CPU-based...  ..., memory, storage, networking, accelerators, and...  ...PCIe, InfiniBand, or NVLink.Hands-on experience...  ...Deep experience with RDMA, RoCE, CXL, NVLink... 
    Temporary work

    Crusoe

    San Francisco, CA
    2 days ago
  • $200k

     ...Ready to architect the high-speed networks powering the AI era? Join a trailblazing leader in GPU-accelerated computing, designing...  ..., or SaltStack. High-Scale Networking: A strong foundation...  ...specific experience with InfiniBand or RDMA (RoCEv2). CI/CD Pipeline Expertise... 
    Full time
    San Francisco, CA
    more than 2 months ago
  • $195k - $235k

     ...urgency, who believe in the scale of our ambition and thrive...  ...Crusoe Cloud is seeking a Staff Network Production Operations Engineer to help own production...  ...backbone, data center fabric, and GPU cluster interconnects. This...  ....Experience operating RDMA/RoCE lossless fabrics for... 
    Temporary work
    Worldwide

    Crusoe

    San Francisco, CA
    16 hours ago
  •  ...hardware and software. Speed and scale are our key differentiators....  ..., non-blocking backend networks for clusters of 100k+ accelerators...  ...lifecycle from customer requirements (GPU shape, workload, scale,...  ...lossless Ethernet fabrics for RDMA (RoCEv2): PFC, ECN tuning, traffic... 
    Local area

    Fluidstack

    San Francisco, CA
    2 days ago
  • $300 per month

     ...urgency, who believe in the scale of our ambition and...  ...highly skilled and motivated GPU Fleet Operations Engineer to join Crusoe’s Fleet...  ...systems, interconnects, and networking hardware.Conduct post-...  ...interconnects such as InfiniBand, NVLink, and RDMA over Converged Ethernet (... 
    Temporary work
    Work at office

    Crusoe

    San Francisco, CA
    3 days ago
  • $342k

     ...OpenAI is currently looking for an experienced Optical Network Engineer based in San Francisco, California. This role involves leading laser...  ...-related efforts in optical interconnect projects for large-scale compute systems, ensuring performance and manufacturability of... 

    OpenAI

    San Francisco, CA
    3 days ago
  • $179k - $218k

     ...a sense of urgency, who believe in the scale of our ambition and thrive on a path not...  ...seeking a Senior Staff Data Center Operations Engineer, GPU Hardware Architecture to be the...  ...identifying "pre-failure" patterns in HBM or NVLink components before they impact customer... 
    Temporary work

    Crusoe

    San Francisco, CA
    4 days ago
  • $250k - $300k

     ...who believe in the scale of our ambition and...  ...Deployment Automation Engineer for the Compute...  ...-scale, multi-node GPU clusters. You will...  ...multi-node context.Networking Knowledge: Strong understanding of RDMA, RoCE, and InfiniBand...  ...MNNVL (Multi-Node NVLink) or specialized AI... 
    Temporary work

    Crusoe

    San Francisco, CA
    1 day ago
  • A technology solutions provider is looking for a Network Engineer to enhance and maintain a large-scale network. This role involves managing both wired and wireless infrastructures, conducting assessments, and ensuring network security. Candidates should have a degree... 

    CGS Federal (Contact Government Services)

    San Francisco, CA
    4 days ago
  • $210k - $240k

     ...Build and help define the network foundation behind a...  ...AI platform supporting GPU infrastructure, distributed...  .... This is a network engineering role first . We are looking...  ...or HPC environments RDMA or RoCEv2 networking,...  ...in a startup or rapidly scaling technical environment Experience... 
    Immediate start

    Stratitech

    San Francisco, CA
    2 days ago
  • $250k

     ...opportunities? Join a rapidly scaling AI cloud infrastructure...  ...provider building a next-generation GPU platform designed for AI...  ...Senior / Staff Site Reliability Engineer to support and scale large-...  ...providers Strong understanding of networking fundamentals (DNS, TCP/IP,... 
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  • $170k - $250k

     ...revenue within six months and is scaling rapidly with a small, high-...  ...This company is looking for an engineer to work directly with the CTO on complex GPU virtualisation challenges. The role...  ...complex issues across low-level networking, GPU drivers, and distributed systems... 
    Full time
    Visa sponsorship
    Flexible hours
    San Francisco, CA
    7 days ago
  •  ...and help build the platform engineers turn to to ship AI...  ...Baseten is building its own GPU infrastructure for large-scale inference. As we move into...  ...workload symptoms that look like network problems, but are not....  ...networking and inference software. RDMA data paths, GPUDirect... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    16 hours ago
  • Fluidstack is seeking a senior network deployment engineer to lead end-to-end fabric turn-ups across data centers. You will own technical execution...  ...close. Candidates should have proven experience on large-scale networks, automation in Python or Go, and familiarity with... 

    Fluidstack

    San Francisco, CA
    16 hours ago
  • Linuxcareers in San Francisco is building AI research infrastructure. You will design, deploy, and operate large-scale GPU clusters powering training, evaluation, and serving for the research team. The role emphasizes extending orchestration with Kubernetes/Slurm, building... 

    Linuxcareers

    San Francisco, CA
    16 hours ago
  •  ...largest, most reliable, and frictionless GPU fleet to support OpenAI’s general purpose...  ...Much more! About the Role As an engineer within Fleet infrastructure, you will design...  ...one of the world’s largest GPU fleet. The scale is immense, the timelines are tight, and... 
    Full time
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    16 hours ago
  •  ...largest AI infrastructure networks. The team owns day-to-day...  ...deliver highly available GPU infrastructure for AI...  ...Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support...  ...supporting RoCE v2 or RDMA‑based Ethernet fabrics, with... 
    Permanent employment

    OpenAI

    San Francisco, CA
    4 days ago
  • $160k - $225k

     ...Cacheflow is seeking a Senior Software Engineer for AI Runtime at Databricks, located in San Francisco. You will be instrumental in building and scaling systems for large-scale GPU training, ensuring high throughput and resilience in training across expansive fleets of... 

    Cacheflow

    San Francisco, CA
    21 hours ago
  • $125k - $145k

     ...urgency, who believe in the scale of our ambition and...  ...offering sustainable, low-cost GPU compute power. As a Senior Cloud Support Engineer, you'll play a crucial...  ...Work closely with SRE, Networking, and Storage teams from...  ...such as Infiniband, RDMA, RoCE, and Software Defined... 
    Full time
    Temporary work

    Crusoe

    San Francisco, CA
    4 days ago
  • $224k - $284k

    Cssmerge is looking for an HPC Network Engineer to join our founding team in San Francisco, responsible...  ...-performance networking that connects our GPU compute. The ideal candidate has experience with network deployment and scaling in HPC or GPU environments. You will also... 

    Cssmerge

    San Francisco, CA
    4 days ago
  • Sesterce Group is seeking a skilled networking engineer to design, deploy, and operate the network infrastructure for their GPU AI factories. The role includes working with InfiniBand...  ...experience, including familiarity with RDMA and strong Linux networking skills.... 

    Sesterce Group

    San Francisco, CA
    16 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to GPU Network Engineer: RDMA/NVLink at Scale. Be the first to apply!