GPU Network Engineer: RDMA/NVLink at Scale
$350kThinking Machines Lab
Thinking Machines Lab is seeking a Network Engineer in San Francisco to manage and improve our GPU network fabric. The role requires in-depth knowledge of large-scale deployments and the ability to debug complex network issues. A collaborative environment is emphasized, where initiative and effective communication with cloud providers are key. The position offers a competitive salary ranging from $350,000 to $475,000 per year, depending on skills and experience. Benefits include unlimited PTO and health coverage, alongside visa sponsorship. #J-18808-Ljbffr Thinking Machines Lab
$350k
...goals. We are scientists, engineers, and builders who’ve... ...We're looking for a network engineer to own the lowest... ...stack that our large‑scale training and inference... ...scale, across large GPU fabrics — both the RDMA/RoCE fabric between nodes and the NVLink/NVSwitch domains...SuggestedLocal areaVisa sponsorshipRelocation package- ...Nscale, a GPU cloud company, seeks a Senior Network Engineer to own the design, automation, and in-service operation of AI-optimized network fabrics for large-scale training and inference workloads. You will lead end-to-end initiatives, mentor engineers, and collaborate...Suggested
- ...AI Cloud Our large-scale GPU cloud platform, Firmus AI... ...As an NVIDIA Cloud and Engineering partner in Asia Pacific, you... ...is seeking a skilled Senior Network Engineer to join our Engineering... ...Ethernet Platform, and (iii) RDMA over Converged Ethernet (RoCE...SuggestedFull time
$150k - $210k
...Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance... .... About the Role The Network Engineering Team is... ...Ethernet networks supporting large-scale training and inference... ...Extensive hands-on experience with RDMA-aware networking (InfiniBand...SuggestedFlexible hours- Network Engineer Gimlet is building the first multi-silicon neocloud designed... ...inference. We combine large-scale compute infrastructure with... ...networking technologies such as RDMA, RoCE, InfiniBand, or... ...have: Experience with AI/HPC, GPU, or large-scale distributed infrastructure...Suggested
- ...of AI infrastructure: large-scale AI datacenters and the orchestration... ...Gimlet Labs is seeking a Network Engineer to design, build, and scale... ...have Experience with AI/HPC, GPU, or large‑scale distributed infrastructure... ...tooling. Familiarity with RDMA, RoCE, InfiniBand, or other...
- ...superintelligence. One person, one GPU. If you'd like to build... ...'ll Do Help to build and scale Lambda's high performance cloud network Work on deploying and... ...rotation for Network Engineering team You Have 10+ years of... ..., partitions), GPUDirect RDMA concepts. Experience with...Work at officeLocal areaWork from homeFlexible hours
$150k - $190k
...software with cost-efficient, large-scale compute. Teams get the tools... ...AI is seeking a Senior Network Engineer with hands-on Cumulus Linux expertise... ...Ethernet fabrics supporting GPU clusters and AI workloads... ...operations Support RoCE/RDMA networking and low-latency transport...Work at officeRemote workWork from homeFlexible hours2 days per week$165k - $200k
...urgency, who believe in the scale of our ambition and thrive... ...Cloud is seeking a Senior Network Production Operations Engineer to support production reliability... ..., data center fabric, and GPU cluster interconnects. This... ...APIs.Experience operating RDMA/RoCE lossless fabrics for...Temporary workWorldwide- ...Hamilton Barnes is seeking a Senior Network Engineer to design, deploy, and operate ultra-low-latency, high-throughput networks for large GPU clusters. You will work with NVIDIA Spectrum and Cumulus Linux to build scalable 400G data center environments optimized for AI...Remote job
$200k
...Ready to architect the high-speed networks powering the AI era? Join a trailblazing leader in GPU-accelerated computing, designing... ..., or SaltStack. High-Scale Networking: A strong foundation... ...specific experience with InfiniBand or RDMA (RoCEv2). CI/CD Pipeline Expertise...Full time- ...Description: Job Title : Senior Network Engineer AI Data Center Job Type : W2/C2... ...network infrastructure for large-scale AI data centers and GPU clusters.The ideal candidate will have... ...InfiniBand and/or RoCE Ethernet fabrics, RDMA, high-speed data center networking,...
$250k
...opportunities? Join a rapidly scaling AI cloud infrastructure... ...provider building a next-generation GPU platform designed for AI... ...Senior / Staff Site Reliability Engineer to support and scale large-... ...providers Strong understanding of networking fundamentals (DNS, TCP/IP,...Full timeRemote work- ...largest AI infrastructure networks. The team owns day-to-day... ...deliver highly available GPU infrastructure for AI... ...Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support... ...supporting RoCE v2 or RDMA-based Ethernet fabrics, with...Permanent employment
$165k - $200k
...urgency, who believe in the scale of our ambition and thrive on... ..., detail-oriented Senior Network Production Engineer to support the physical and... ...performance compute (HPC) and GPU-based AI infrastructure, this... ...GPU cluster networking (e.g., RDMA/RoCE, InfiniBand, or NVIDIA...Full timeTemporary workRemote work- ...Job Description: Job Title : Staff Infrastructure Engineer - Cloud & GPU Job Type : W2/C2C Experience : 8+ Years Location... ...Infrastructure Engineer to design, deploy, automate and maintain large-scale GPU computing infrastructure supporting AI/ML workloads.The...
$225k
...the next step in your career? Join a rapidly scaling AI cloud infrastructure provider building next-generation GPU platforms for large-scale AI training,... ...exciting opportunity has arisen for a Senior Network Engineer to design, deploy, and operate ultra-low-latency...Remote work$250k - $300k
...who believe in the scale of our ambition and... ...Deployment Automation Engineer for the Compute... ...-scale, multi-node GPU clusters. You will... ...multi-node context. Networking Knowledge: Strong understanding of RDMA, RoCE, and... ...MNNVL (Multi-Node NVLink) or specialized AI...Full timeTemporary work- ...serverless runtime that launches GPU‑backed containers in less than 1 second and quickly scales out to thousands of GPUs.... ...Deploy and validate data center network infrastructure (front‑end, back... ...Operations, ICT, Hardware, and Network Engineering to identify blockers early,...
$300k
...generation of large-scale training, inference,... ...Principal Software Engineer will take ownership... ...GPUs, high-performance networking fabrics, storage... ...workload scheduling, GPU resource management,... ...Knowledge of CUDA, NCCL, NVLink, NVSwitch, GPUDirect RDMA, or GPU resource...Full timeRemote workFlexible hours$225k - $275k
...sense of urgency, who believe in the scale of our ambition and thrive on a path not... ...Cloud is seeking a Senior Staff Network Deployment Engineer to serve as the technical owner of how... ...of high-performance compute (HPC) and GPU-based AI infrastructure, you will define...Temporary workRemote work- Senior Network Engineer Houston; New York; San Francisco; Seattle About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure... ...networking platforms Experience with large-scale DCI, long-haul optical transport, or...
$180k - $250k
...production, and do it at scale without compromise.... ...You are a hands-on engineer who builds the... ...keep a large fleet of GPU servers healthy and... ...errors, disk failures, network issues, thermals)... ...health monitoring, DCGM, NVLink/NVSwitch diagnostics, RDMA, InfiniBand/RoCEv2...Local areaRelocation package- A leading AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for designing high-performance GPU kernels and using advanced techniques to boost computation efficiency. Ideal...
$160k - $260k
...teamCorporate IT drives our IT support, IT engineering and business engineering functions at... ...applications, identity and access management. Network Operations builds our corporate network,... ...(AWS/Alibaba Cloud) at massive scale. Acting as a principal technical leader,...Temporary workWork at officeLocal areaRemote workFlexible hoursWeekend work- ...Firmus AI Cloud Our large-scale GPU cloud platform, Firmus AI Cloud... ...the founders, build a strong network, and see the impact of your work... ...with the internal Firmus Engineering and Technology teams to develop... ...Ethernet Platform, and RDMA over Converged Ethernet (RoCE...Full timeWork experience placement
- ...seeking a Senior/Staff Rendering Systems Engineer to build and scale high-throughput, Unreal Engine–based... ...will own the low-level rendering and GPU performance work required to run many-... ...collective communications primitives (e.g., NVLink, NCCL); A strong plus, but not a...Work at officeVisa sponsorshipWork visaRelocation packageFlexible hours
- ...Mastercard is seeking a visionary VP of Software Engineering to build and scale Decision Stream, Mastercard's AI-native decisioning platform. You will lead through technical excellence, shape architecture for distributed systems and cloud-native tech, and partner with...
$260k - $340k
...urgency, who believe in the scale of our ambition and... ...Principal Systems Software Engineer , you will serve as the... ...that deliver raw GPU throughput via zero-latency InfiniBand/RDMA fabrics for massive-scale... ...methods for managing memory, networking, and compute that don't...Full timeTemporary work- ...TRM Labs is seeking a Senior Software Engineer, Data Infrastructure, to own the petabyte-scale Postgres serving layer, optimize queries, and cut storage costs with AI-assisted analysis. You will automate database operations, validate changes, and maintain real-time data...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to GPU Network Engineer: RDMA/NVLink at Scale. Be the first to apply!
- network applications engineer San Francisco, CA
- data center network engineer San Francisco, CA
- juniper network engineer San Francisco, CA
- production network engineer San Francisco, CA
- senior network engineer San Francisco, CA
- cisco ccnp network engineer San Francisco, CA
- wireless network engineer San Francisco, CA
- principal network engineer San Francisco, CA
- remote cisco network engineer San Francisco, CA
- core network engineer San Francisco, CA




