GPU Network Engineer: RDMA/NVLink at Scale
$350kThinking Machines Lab
Thinking Machines Lab is seeking a Network Engineer in San Francisco to manage and improve our GPU network fabric. The role requires in-depth knowledge of large-scale deployments and the ability to debug complex network issues. A collaborative environment is emphasized, where initiative and effective communication with cloud providers are key. The position offers a competitive salary ranging from $350,000 to $475,000 per year, depending on skills and experience. Benefits include unlimited PTO and health coverage, alongside visa sponsorship. #J-18808-Ljbffr Thinking Machines Lab
- Thinking Machines Lab Inc. is seeking a network engineer to own the lowest layers of the network stack for large-scale training and inference. You will ensure interconnect reliability across GPU fabrics, debugging NICs, and building instrumentation for faster troubleshooting...Suggested
- ...help build the platform engineers turn to to ship AI products... ...multi-modal workloads scale, the network is the computer. We are... ...foundational engineers to lead our GPU Networking efforts, making RDMA a first-class building... ...communication across NVLink and InfiniBand for our...SuggestedFull timeFlexible hours
$350k
...the Role We're looking for a network engineer to own the lowest layers of the... ...network stack that our large-scale training and inference depend... ...at scale, across large GPU fabrics - both the RDMA/RoCE fabric between nodes and the NVLink/NVSwitch domains within them....SuggestedVisa sponsorshipWork visaRelocation package$150k - $300k
Prime Intellect in San Francisco seeks a Solutions Architect for GPU Infrastructure who will transform client requirements into robust systems capable of training advanced AI models. Responsibilities include designing GPU cluster architectures, deploying orchestration systems...Suggested$250k - $320k
...of AI infrastructure: large-scale AI datacenters and the orchestration... ...Gimlet Labs is seeking a Network Engineer to design, build, and scale... ...have Experience with AI/HPC, GPU, or large‑scale distributed infrastructure... ...tooling. Familiarity with RDMA, RoCE, InfiniBand, or other...Suggested$190k - $280k
...Senior Network Engineer San Francisco About the Role Together AI... ...boundaries. You will work on large-scale, multi-vendor data center... ...~ Foundational knowledge of RDMA networking and technologies such... .... Experience supporting GPU clusters, HPC environments, distributed...Full time- ...superintelligence. One person, one GPU.If you'd like to build... ...ll DoHelp to build and scale Lambda's high... ...and configuring networking hardware for new and existing... ...call rotation for Network Engineering teamYouHave 10+ years... ...partitions), GPUDirect RDMA concepts.Experience...Work at officeLocal areaWork from homeFlexible hours
- ...Principal Network Engineer – AI InfrastructureThe Network Operations and Engineering... ..., supporting tightly coupled GPU clusters where network... ...of our Infiniband and RDMA-based network fabrics. You will... ...reviewing, and evolving large-scale Infiniband and RoCE fabric architectures...
- ...About the Team The Core Network Engineering team owns the end-to-end networking... ...xPU networking used for large-scale training and inference workloads... ...across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-performance GPU interconnects Define and operationalize...Full time
- ...hardware and software. Speed and scale are our key differentiators.... ..., non-blocking backend networks for clusters of 100k+ accelerators... ...lifecycle from customer requirements (GPU shape, workload, scale,... ...lossless Ethernet fabrics for RDMA (RoCEv2): PFC, ECN tuning, traffic...Local area
$10k - $20k
...services. The company is looking for a Network Automation Engineer to design and build high-speed, low-latency network fabrics for large-scale GPU-accelerated computing. The role focuses... ...specific experience with InfiniBand or RDMA (RoCEv2). CI/CD Pipeline Expertise:...Full time$200k
...Ready to architect the high-speed networks powering the AI era? Join a trailblazing leader in GPU-accelerated computing, designing... ..., or SaltStack. High-Scale Networking: A strong foundation... ...specific experience with InfiniBand or RDMA (RoCEv2). CI/CD Pipeline Expertise...Full time$195k - $235k
...urgency, who believe in the scale of our ambition and thrive... ...Crusoe Cloud is seeking a Staff Network Production Operations Engineer to help own production... ...backbone, data center fabric, and GPU cluster interconnects. This... ....Experience operating RDMA/RoCE lossless fabrics for...Temporary workWorldwide- A technology solutions provider is looking for a Network Engineer to enhance and maintain a large-scale network. This role involves managing both wired and wireless infrastructures, conducting assessments, and ensuring network security. Candidates should have a degree...
- Sciforium is seeking a GPU Kernel Engineer to push performance on modern accelerators. You will design and optimize custom GPU kernels, from low... ...to integrating ops in ML frameworks used for large-scale training and inference. Ideal candidates have 5+ years in GPU...
$179k - $218k
...a sense of urgency, who believe in the scale of our ambition and thrive on a path not... ...seeking a Senior Staff Data Center Operations Engineer, GPU Hardware Architecture to be the... ...identifying "pre-failure" patterns in HBM or NVLink components before they impact customer...Temporary work$210k - $240k
...Build and help define the network foundation behind a... ...AI platform supporting GPU infrastructure, distributed... .... This is a network engineering role first . We are looking... ...or HPC environments RDMA or RoCEv2 networking,... ...in a startup or rapidly scaling technical environment Experience...Immediate start$250k - $300k
...who believe in the scale of our ambition and... ...Deployment Automation Engineer for the Compute... ...-scale, multi-node GPU clusters. You will... ...multi-node context.Networking Knowledge: Strong understanding of RDMA, RoCE, and InfiniBand... ...MNNVL (Multi-Node NVLink) or specialized AI...Temporary work$170k - $250k
...revenue within six months and is scaling rapidly with a small, high-... ...This company is looking for an engineer to work directly with the CTO on complex GPU virtualisation challenges. The role... ...complex issues across low-level networking, GPU drivers, and distributed systems...Full timeVisa sponsorshipFlexible hours$250k
...opportunities? Join a rapidly scaling AI cloud infrastructure... ...provider building a next-generation GPU platform designed for AI... ...Senior / Staff Site Reliability Engineer to support and scale large-... ...providers Strong understanding of networking fundamentals (DNS, TCP/IP,...Full timeRemote work- ...largest AI infrastructure networks. The team owns day-to-day... ...deliver highly available GPU infrastructure for AI... ...Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support... ...supporting RoCE v2 or RDMA‑based Ethernet fabrics, with...Permanent employment
- Linuxcareers in San Francisco is building AI research infrastructure. You will design, deploy, and operate large-scale GPU clusters powering training, evaluation, and serving for the research team. The role emphasizes extending orchestration with Kubernetes/Slurm, building...
- ...serverless runtime that launches GPU-backed containers in less than 1 second and quickly scales out to thousands of GPUs.... ...Deploy and validate data center network infrastructure (front-end, back... ...Operations, ICT, Hardware, and Network Engineering to identify blockers early,...
- ...and help build the platform engineers turn to to ship AI... ...Baseten is building its own GPU infrastructure for large-scale inference. As we move into... ...workload symptoms that look like network problems, but are not.... ...networking and inference software. RDMA data paths, GPUDirect...Full timeFlexible hours
- Hyperbolic is seeking a very senior Infrastructure Engineer to scale a GPU cloud marketplace. You will build a multi-tenant provisioning and virtualization layer, turning global GPU inventories into an orchestrated pool for AI developers and researchers. You will own the...
- ...spanning hardware and software. Speed and scale are our key differentiators. Come be a... ...matters to the world. The Production Engineering Team Examples of key exciting problems... ...of 10s to 100s of GWs: at our scale, a GPU failure isn't a ticket. It's a throughput...Full timeLocal area
$180k - $250k
...production, and do it at scale without compromise.... ...You are a hands-on engineer who builds the... ...keep a large fleet of GPU servers healthy and... ...errors, disk failures, network issues, thermals)... ...health monitoring, DCGM, NVLink/NVSwitch diagnostics, RDMA, InfiniBand/RoCEv2...Full timeLocal areaRelocation package- ...Senior Network EngineerHouston; New York; San Francisco; SeattleAbout NscaleNscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI... ...networking platformsExperience with large-scale DCI, long-haul optical transport, or...
$195k - $235k
...urgency, who believe in the scale of our ambition and thrive... ...Crusoe Cloud is seeking a Staff Network Production Operations Engineer to help own production... ...backbone, data center fabric, and GPU cluster interconnects. This... .... ~ Experience operating RDMA/RoCE lossless fabrics for...Temporary workWorldwide$240k - $280k
...We're looking for a Software Engineer to build the systems that treat... ...single API call to stand up, scale, or tear down a cluster, and the... ..., inference bring-up to GPU driver/CUDA stack, health validation... ..., Redfish/IPMI, BMC) and/or networking fundamentals (VLANs, BGP, fabric...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to GPU Network Engineer: RDMA/NVLink at Scale. Be the first to apply!
- work from home network engineer San Francisco, CA
- cisco ccnp network engineer San Francisco, CA
- network engineer San Francisco, CA
- juniper network engineer San Francisco, CA
- network solutions engineer San Francisco, CA
- production network engineer San Francisco, CA
- network engineer level San Francisco, CA
- entry level network engineer San Francisco, CA
- senior network engineer San Francisco, CA
- IT network engineer San Francisco, CA



