Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

GPU Network Engineer: RDMA/NVLink at Scale

$350k

Thinking Machines Lab

Thinking Machines Lab is seeking a Network Engineer in San Francisco to manage and improve our GPU network fabric. The role requires in-depth knowledge of large-scale deployments and the ability to debug complex network issues. A collaborative environment is emphasized, where initiative and effective communication with cloud providers are key. The position offers a competitive salary ranging from $350,000 to $475,000 per year, depending on skills and experience. Benefits include unlimited PTO and health coverage, alongside visa sponsorship. #J-18808-Ljbffr Thinking Machines Lab

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the GPU Network Engineer: RDMA/NVLink at Scale in San Francisco, CA vacancy
  •  ...help build the platform engineers turn to to ship AI products...  ...multi-modal workloads scale, the network is the computer. We are...  ...foundational engineers to lead our GPU Networking efforts, making RDMA a first-class building...  ...communication across NVLink and InfiniBand for our... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  • $350k

    Talent Network Engineer, Supercomputing Thinking Machines Lab IT San Francisco...  ...network stack that our large-scale training and inference depend...  ...at scale, across large GPU fabrics — both the RDMA/RoCE fabric between nodes and the NVLink/NVSwitch domains within them.... 
    Suggested
    Local area
    Immediate start
    Visa sponsorship
    Work visa
    Relocation package

    Socket

    San Francisco, CA
    1 day ago
  •  ...will design, deploy, and operate the network infrastructure underpinning Sesterce's GPU AI factories across Europe —...  ...physical cabling to BGP policies and RDMA fabric tuning. What you will do Design...  ...or 400G/800G Ethernet at scale Deep familiarity with RDMA, RoCE v... 
    Suggested

    Sesterce Group

    San Francisco, CA
    5 days ago
  • $250k - $320k

     ...of AI infrastructure: large-scale AI datacenters and the orchestration...  ...Gimlet Labs is seeking a Network Engineer to design, build, and scale...  ...have Experience with AI/HPC, GPU, or large‑scale distributed...  ...similar tooling. Familiarity with RDMA, RoCE, InfiniBand, or other... 
    Suggested

    Gimlet Labs, Inc.

    San Francisco, CA
    1 day ago
  •  ...About the Team The Core Network Engineering team owns the end-to-end networking...  ...xPU networking used for large-scale training and inference workloads...  ...across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-performance GPU interconnects Define and operationalize... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  • $342k

     ...OpenAI is currently looking for an experienced Optical Network Engineer based in San Francisco, California. This role involves leading laser...  ...-related efforts in optical interconnect projects for large-scale compute systems, ensuring performance and manufacturability of... 

    OpenAI

    San Francisco, CA
    3 days ago
  • $250k

     ...A Series A Funded start-up in California is seeking a Systems Engineer to design and optimize systems handling complex ML pipelines. The role involves building scalable infrastructure, developing CI/CD pipelines, and ensuring system performance. Key qualifications include... 
    Full time

    Acceler8 Talent

    San Francisco, CA
    2 hours ago
  •  ...and help build the platform engineers turn to to ship AI...  ...Baseten is building its own GPU infrastructure for large-scale inference. As we move into...  ...workload symptoms that look like network problems, but are not....  ...networking and inference software. RDMA data paths, GPUDirect... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  •  ...largest, most reliable, and frictionless GPU fleet to support OpenAI’s general purpose...  ...Much more! About the Role As an engineer within Fleet infrastructure, you will design...  ...one of the world’s largest GPU fleet. The scale is immense, the timelines are tight, and... 
    Full time
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...serverless runtime that launches GPU‑backed containers in less than 1 second and quickly scales out to thousands of GPUs....  ...Deploy and validate data center network infrastructure (front‑end, back...  ...Operations, ICT, Hardware, and Network Engineering to identify blockers early,... 

    BEAM inc.

    San Francisco, CA
    3 days ago
  • $200k

     ...Ready to architect the high-speed networks powering the AI era? Join a trailblazing leader in GPU-accelerated computing, designing...  ..., or SaltStack. High-Scale Networking: A strong foundation...  ...specific experience with InfiniBand or RDMA (RoCEv2). CI/CD Pipeline Expertise... 
    Permanent employment
    San Francisco, CA
    more than 2 months ago
  • $125k - $145k

     ...Senior Cloud Support Engineer Crusoe Cloud is revolutionizing...  ...sustainable, low-cost GPU compute power. As a...  ...hardware failures, and scaling tests using CLI and...  ...Work closely with SRE, Networking, and Storage teams from...  ...technologies such as Infiniband, RDMA, RoCE, and Software... 
    Full time
    Temporary work

    Crusoe

    San Francisco, CA
    3 days ago
  •  ...Senior Site Reliability Engineer Location: Global...  ...access to the kind of scaled AI infrastructure once...  ...building the systems, network, and orchestration layer...  ...and debug large-scale GPU infrastructure used for...  ...infrastructure (ECC errors, NVLink degradation, NCCL... 
    Full time
    Remote work

    Andromeda Cluster, Inc

    San Francisco, CA
    13 hours ago
  • A leading streaming platform is looking for a Staff DevOps Engineer in San Francisco, CA, to automate and scale systems supporting their streaming services. The role involves leading projects on infrastructure automation, best-practice adoption, and collaboration across... 
    Full time

    Crunchyroll

    San Francisco, CA
    2 hours ago
  • $300 per month

     ...intelligence. We’re crafting the engine that powers a world...  ...— without sacrificing scale, speed, or...  ...systems that deliver raw GPU throughput via zero-latency InfiniBand/RDMA fabrics, ensuring that massive...  ...ways of managing memory, networking, and compute that don't... 

    Crusoe

    San Francisco, CA
    1 day ago
  • $220k - $290k

     ...move from idea to production, and do it at scale without compromise. For developers and...  ...prospective customers. Collaborate with engineering and product development teams to tailor solutions...  ...optimization, MLOps practices, and GPU infrastructure. Professional proficiency... 
    Full time
    Work at office

    fal

    San Francisco, CA
    2 days ago
  • A leading identity solutions provider in San Francisco is looking for a Solutions Engineer to act as a technical advisor for customers, guiding them through the implementation of Persona's platform. This position involves solving complex business challenges through technical... 

    Persona

    San Francisco, CA
    3 days ago
  • $300 per month

     ...who believe in the scale of our ambition and...  ...Staff Hardware Systems Engineer to strengthen...  ...across Crusoe Cloud’s GPU- and CPU-based...  ..., memory, storage, networking, accelerators, and...  ...PCIe, InfiniBand, or NVLink. ~ Hands-on experience...  ...experience with RDMA, RoCE, CXL, NVLink... 
    Temporary work

    Crusoe

    San Francisco, CA
    a month ago
  • $250k

     ...opportunities? Join a rapidly scaling AI cloud infrastructure...  ...provider building a next-generation GPU platform designed for AI...  ...Senior / Staff Site Reliability Engineer to support and scale large-...  ...providers Strong understanding of networking fundamentals (DNS, TCP/IP,... 
    Permanent employment
    Remote work
    San Francisco, CA
    more than 2 months ago
  •  ...inference. About the Role We’re hiring engineers to scale and optimize OpenAI’s inference infrastructure across emerging GPU platforms. You’ll work across the stack - from...  ...inference workloads across memory, network, and compute layers. Validate correctness... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...About the Team The Scaling team is responsible for...  ...architectural and engineering backbone of OpenAI’s...  ...spans system software, networking, platform architecture...  ...system, including CPU, GPU, memory subsystem,...  ...including WAN traffic, NVlink and RDMA collectives), storage... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  • A leading identity solutions provider in San Francisco is looking for a Solutions Engineer to act as a technical advisor for customers, guiding them through the implementation of Persona's platform. This position involves solving complex business challenges through technical... 
    Full time

    Persona

    San Francisco, CA
    2 hours ago
  • $135.61k - $184.04k

     ...Network Engineer Employment Type: Full Time, Experienced level Department: Information Technology CGS is seeking an experienced Network Engineer...  ...on the evaluation, enhancement, and maintenance of a large scale network project that includes both wired and wireless network... 
    Full time
    Local area
    Monday to Friday
    Flexible hours

    Contact-Government-Services,-LL

    San Francisco, CA
    1 day ago
  • $225k - $275k

     ...urgency, who believe in the scale of our ambition and thrive on...  ...is seeking a Senior Staff Network Production Engineer to own production reliability...  ...backbone, data center fabric, and GPU cluster interconnects. You...  ...Arbor. ~ GPU Cluster and RDMA Networking: Hands-on... 
    Temporary work

    Crusoe

    San Francisco, CA
    29 days ago
  •  ...research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance,...  ...unchecked growth. About the role As a software engineer on the Fleet High Performance Computing (HPC) team,... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  • A growing AI startup in San Francisco is seeking a Staff Software Engineer to design, build, and scale backend systems for their AI agent platform. The ideal candidate has over 7 years of backend or full-stack engineering experience. You will lead complex features from... 
    Full time
    Immediate start

    Broccoli AI

    San Francisco, CA
    2 hours ago
  • $300 per month

     ...urgency, who believe in the scale of our ambition and...  ...highly skilled and motivated GPU Fleet Operations Engineer to join Crusoe’s Fleet...  ...systems, interconnects, and networking hardware. Conduct post...  ...such as InfiniBand, NVLink, and RDMA over Converged Ethernet (... 
    Temporary work
    Work at office

    Crusoe

    San Francisco, CA
    9 days ago
  • $179k - $218k

     ...a sense of urgency, who believe in the scale of our ambition and thrive on a path not...  ...seeking a Senior Staff Data Center Operations Engineer, GPU Hardware Architecture to be the...  ...identifying "pre-failure" patterns in HBM or NVLink components before they impact customer... 
    Temporary work

    Crusoe

    San Francisco, CA
    27 days ago
  •  ...Experienced level Department: Information Technology CGS is seeking an experienced Network Engineer to join a team focused on the evaluation, enhancement, and maintenance of a large scale network project that includes both wired and wireless network infrastructure and... 
    Full time
    Local area
    Monday to Friday
    Flexible hours

    CGS Federal (Contact Government Services)

    San Francisco, CA
    5 days ago
  •  ...advertisers. Our platform is purpose-built to deliver outcomes at scale, not just impressions. We are a team of builders, operators,...  ...what it becomes. Role Overview We are looking for a skilled Network Engineer to join our San Francisco engineering hub. In this role, you... 
    Immediate start

    Socket

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to GPU Network Engineer: RDMA/NVLink at Scale. Be the first to apply!