Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

GPU Network Engineer: RDMA/NVLink at Scale

$350k

Thinking Machines Lab

Thinking Machines Lab is seeking a Network Engineer in San Francisco to manage and improve our GPU network fabric. The role requires in-depth knowledge of large-scale deployments and the ability to debug complex network issues. A collaborative environment is emphasized, where initiative and effective communication with cloud providers are key. The position offers a competitive salary ranging from $350,000 to $475,000 per year, depending on skills and experience. Benefits include unlimited PTO and health coverage, alongside visa sponsorship. #J-18808-Ljbffr Thinking Machines Lab

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the GPU Network Engineer: RDMA/NVLink at Scale in San Francisco, CA vacancy
  • $350k

     ...goals. We are scientists, engineers, and builders who’ve...  ...We're looking for a network engineer to own the lowest...  ...stack that our large‑scale training and inference...  ...scale, across large GPU fabrics — both the RDMA/RoCE fabric between nodes and the NVLink/NVSwitch domains... 
    Suggested
    Local area
    Visa sponsorship
    Relocation package

    Thinking Machines Lab

    San Francisco, CA
    3 days ago
  •  ...Nscale, a GPU cloud company, seeks a Senior Network Engineer to own the design, automation, and in-service operation of AI-optimized network fabrics for large-scale training and inference workloads. You will lead end-to-end initiatives, mentor engineers, and collaborate... 
    Suggested

    Jobleads-US

    San Francisco, CA
    17 hours ago
  •  ...AI Cloud Our large-scale GPU cloud platform, Firmus AI...  ...As an NVIDIA Cloud and Engineering partner in Asia Pacific, you...  ...is seeking a skilled Senior Network Engineer to join our Engineering...  ...Ethernet Platform, and (iii) RDMA over Converged Ethernet (RoCE... 
    Suggested
    Full time

    Firmus Technologies

    San Francisco, CA
    1 day ago
  • $150k - $210k

     ...Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance...  .... About the Role The Network Engineering Team is...  ...Ethernet networks supporting large-scale training and inference...  ...Extensive hands-on experience with RDMA-aware networking (InfiniBand... 
    Suggested
    Flexible hours

    Jobleads-US

    San Francisco, CA
    17 hours ago
  • Network Engineer Gimlet is building the first multi-silicon neocloud designed...  ...inference. We combine large-scale compute infrastructure with...  ...networking technologies such as RDMA, RoCE, InfiniBand, or...  ...have: Experience with AI/HPC, GPU, or large-scale distributed infrastructure... 
    Suggested

    Gimlet Labs

    San Francisco, CA
    3 days ago
  •  ...of AI infrastructure: large-scale AI datacenters and the orchestration...  ...Gimlet Labs is seeking a Network Engineer to design, build, and scale...  ...have Experience with AI/HPC, GPU, or large‑scale distributed infrastructure...  ...tooling. Familiarity with RDMA, RoCE, InfiniBand, or other... 

    Gimlet Labs

    San Francisco, CA
    2 days ago
  •  ...superintelligence. One person, one GPU. If you'd like to build...  ...'ll Do Help to build and scale Lambda's high performance cloud network Work on deploying and...  ...rotation for Network Engineering team You Have 10+ years of...  ..., partitions), GPUDirect RDMA concepts. Experience with... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda

    San Francisco, CA
    4 days ago
  • $150k - $190k

     ...software with cost-efficient, large-scale compute. Teams get the tools...  ...AI is seeking a Senior Network Engineer with hands-on Cumulus Linux expertise...  ...Ethernet fabrics supporting GPU clusters and AI workloads...  ...operations Support RoCE/RDMA networking and low-latency transport... 
    Work at office
    Remote work
    Work from home
    Flexible hours
    2 days per week

    Lightning AI

    San Francisco, CA
    5 days ago
  • $165k - $200k

     ...urgency, who believe in the scale of our ambition and thrive...  ...Cloud is seeking a Senior Network Production Operations Engineer to support production reliability...  ..., data center fabric, and GPU cluster interconnects. This...  ...APIs.Experience operating RDMA/RoCE lossless fabrics for... 
    Temporary work
    Worldwide

    Crusoe

    San Francisco, CA
    2 days ago
  •  ...Hamilton Barnes is seeking a Senior Network Engineer to design, deploy, and operate ultra-low-latency, high-throughput networks for large GPU clusters. You will work with NVIDIA Spectrum and Cumulus Linux to build scalable 400G data center environments optimized for AI... 
    Remote job

    Jobleads-US

    San Francisco, CA
    17 hours ago
  • $200k

     ...Ready to architect the high-speed networks powering the AI era? Join a trailblazing leader in GPU-accelerated computing, designing...  ..., or SaltStack. High-Scale Networking: A strong foundation...  ...specific experience with InfiniBand or RDMA (RoCEv2). CI/CD Pipeline Expertise... 
    Full time
    San Francisco, CA
    more than 2 months ago
  •  ...Description: Job Title : Senior Network Engineer AI Data Center Job Type : W2/C2...  ...network infrastructure for large-scale AI data centers and GPU clusters.The ideal candidate will have...  ...InfiniBand and/or RoCE Ethernet fabrics, RDMA, high-speed data center networking,... 

    Artmac Soft LLC

    San Francisco, CA
    2 days ago
  • $250k

     ...opportunities? Join a rapidly scaling AI cloud infrastructure...  ...provider building a next-generation GPU platform designed for AI...  ...Senior / Staff Site Reliability Engineer to support and scale large-...  ...providers Strong understanding of networking fundamentals (DNS, TCP/IP,... 
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  •  ...largest AI infrastructure networks. The team owns day-to-day...  ...deliver highly available GPU infrastructure for AI...  ...Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support...  ...supporting RoCE v2 or RDMA-based Ethernet fabrics, with... 
    Permanent employment

    OpenAI

    San Francisco, CA
    19 days ago
  • $165k - $200k

     ...urgency, who believe in the scale of our ambition and thrive on...  ..., detail-oriented Senior Network Production Engineer to support the physical and...  ...performance compute (HPC) and GPU-based AI infrastructure, this...  ...GPU cluster networking (e.g., RDMA/RoCE, InfiniBand, or NVIDIA... 
    Full time
    Temporary work
    Remote work

    Crusoe

    San Francisco, CA
    2 days ago
  •  ...Job Description: Job Title : Staff Infrastructure Engineer - Cloud & GPU Job Type : W2/C2C Experience : 8+ Years Location...  ...Infrastructure Engineer to design, deploy, automate and maintain large-scale GPU computing infrastructure supporting AI/ML workloads.The... 

    Artmac Soft LLC

    San Francisco, CA
    3 days ago
  • $225k

     ...the next step in your career? Join a rapidly scaling AI cloud infrastructure provider building next-generation GPU platforms for large-scale AI training,...  ...exciting opportunity has arisen for a Senior Network Engineer to design, deploy, and operate ultra-low-latency... 
    Remote work

    Hamilton Barnes

    San Francisco, CA
    18 hours ago
  • $250k - $300k

     ...who believe in the scale of our ambition and...  ...Deployment Automation Engineer for the Compute...  ...-scale, multi-node GPU clusters. You will...  ...multi-node context. Networking Knowledge: Strong understanding of RDMA, RoCE, and...  ...MNNVL (Multi-Node NVLink) or specialized AI... 
    Full time
    Temporary work

    Crusoe

    San Francisco, CA
    6 days ago
  •  ...serverless runtime that launches GPU‑backed containers in less than 1 second and quickly scales out to thousands of GPUs....  ...Deploy and validate data center network infrastructure (front‑end, back...  ...Operations, ICT, Hardware, and Network Engineering to identify blockers early,... 

    Beam

    San Francisco, CA
    3 days ago
  • $300k

     ...generation of large-scale training, inference,...  ...Principal Software Engineer will take ownership...  ...GPUs, high-performance networking fabrics, storage...  ...workload scheduling, GPU resource management,...  ...Knowledge of CUDA, NCCL, NVLink, NVSwitch, GPUDirect RDMA, or GPU resource... 
    Full time
    Remote work
    Flexible hours
    San Francisco, CA
    more than 2 months ago
  • $225k - $275k

     ...sense of urgency, who believe in the scale of our ambition and thrive on a path not...  ...Cloud is seeking a Senior Staff Network Deployment Engineer to serve as the technical owner of how...  ...of high-performance compute (HPC) and GPU-based AI infrastructure, you will define... 
    Temporary work
    Remote work

    Crusoe

    San Francisco, CA
    2 days ago
  • Senior Network Engineer Houston; New York; San Francisco; Seattle About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure...  ...networking platforms Experience with large-scale DCI, long-haul optical transport, or... 

    Nscale

    San Francisco, CA
    4 days ago
  • $180k - $250k

     ...production, and do it at scale without compromise....  ...You are a hands-on engineer who builds the...  ...keep a large fleet of GPU servers healthy and...  ...errors, disk failures, network issues, thermals)...  ...health monitoring, DCGM, NVLink/NVSwitch diagnostics, RDMA, InfiniBand/RoCEv2... 
    Local area
    Relocation package

    features and labels

    San Francisco, CA
    2 days ago
  • A leading AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for designing high-performance GPU kernels and using advanced techniques to boost computation efficiency. Ideal... 

    Baseten

    San Francisco, CA
    3 days ago
  • $160k - $260k

     ...teamCorporate IT drives our IT support, IT engineering and business engineering functions at...  ...applications, identity and access management. Network Operations builds our corporate network,...  ...(AWS/Alibaba Cloud) at massive scale. Acting as a principal technical leader,... 
    Temporary work
    Work at office
    Local area
    Remote work
    Flexible hours
    Weekend work

    Airwallex

    San Francisco, CA
    4 hours ago
  •  ...Firmus AI Cloud Our large-scale GPU cloud platform, Firmus AI Cloud...  ...the founders, build a strong network, and see the impact of your work...  ...with the internal Firmus Engineering and Technology teams to develop...  ...Ethernet Platform, and RDMA over Converged Ethernet (RoCE... 
    Full time
    Work experience placement

    Firmus Technologies

    San Francisco, CA
    3 days ago
  •  ...seeking a Senior/Staff Rendering Systems Engineer to build and scale high-throughput, Unreal Engine–based...  ...will own the low-level rendering and GPU performance work required to run many-...  ...collective communications primitives (e.g., NVLink, NCCL); A strong plus, but not a... 
    Work at office
    Visa sponsorship
    Work visa
    Relocation package
    Flexible hours

    Engg

    San Francisco, CA
    18 hours ago
  •  ...Mastercard is seeking a visionary VP of Software Engineering to build and scale Decision Stream, Mastercard's AI-native decisioning platform. You will lead through technical excellence, shape architecture for distributed systems and cloud-native tech, and partner with... 

    Jobleads-US

    San Francisco, CA
    4 days ago
  • $260k - $340k

     ...urgency, who believe in the scale of our ambition and...  ...Principal Systems Software Engineer , you will serve as the...  ...that deliver raw GPU throughput via zero-latency InfiniBand/RDMA fabrics for massive-scale...  ...methods for managing memory, networking, and compute that don't... 
    Full time
    Temporary work

    Crusoe Energy Systems LLC

    San Francisco, CA
    3 days ago
  •  ...TRM Labs is seeking a Senior Software Engineer, Data Infrastructure, to own the petabyte-scale Postgres serving layer, optimize queries, and cut storage costs with AI-assisted analysis. You will automate database operations, validate changes, and maintain real-time data... 

    Jobleads-US

    San Francisco, CA
    17 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to GPU Network Engineer: RDMA/NVLink at Scale. Be the first to apply!