GPU Network Engineer: RDMA/NVLink at Scale
$350kThinking Machines Lab
Thinking Machines Lab is seeking a Network Engineer in San Francisco to manage and improve our GPU network fabric. The role requires in-depth knowledge of large-scale deployments and the ability to debug complex network issues. A collaborative environment is emphasized, where initiative and effective communication with cloud providers are key. The position offers a competitive salary ranging from $350,000 to $475,000 per year, depending on skills and experience. Benefits include unlimited PTO and health coverage, alongside visa sponsorship. #J-18808-Ljbffr Thinking Machines Lab
- ...help build the platform engineers turn to to ship AI products... ...multi-modal workloads scale, the network is the computer. We are... ...foundational engineers to lead our GPU Networking efforts, making RDMA a first-class building... ...communication across NVLink and InfiniBand for our...SuggestedFull timeFlexible hours
$350k
Talent Network Engineer, Supercomputing Thinking Machines Lab IT San Francisco... ...network stack that our large-scale training and inference depend... ...at scale, across large GPU fabrics — both the RDMA/RoCE fabric between nodes and the NVLink/NVSwitch domains within them....SuggestedLocal areaImmediate startVisa sponsorshipWork visaRelocation package- ...will design, deploy, and operate the network infrastructure underpinning Sesterce's GPU AI factories across Europe —... ...physical cabling to BGP policies and RDMA fabric tuning. What you will do Design... ...or 400G/800G Ethernet at scale Deep familiarity with RDMA, RoCE v...Suggested
$250k - $320k
...of AI infrastructure: large-scale AI datacenters and the orchestration... ...Gimlet Labs is seeking a Network Engineer to design, build, and scale... ...have Experience with AI/HPC, GPU, or large‑scale distributed... ...similar tooling. Familiarity with RDMA, RoCE, InfiniBand, or other...Suggested- ...About the Team The Core Network Engineering team owns the end-to-end networking... ...xPU networking used for large-scale training and inference workloads... ...across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-performance GPU interconnects Define and operationalize...SuggestedFull time
$342k
...OpenAI is currently looking for an experienced Optical Network Engineer based in San Francisco, California. This role involves leading laser... ...-related efforts in optical interconnect projects for large-scale compute systems, ensuring performance and manufacturability of...$250k
...A Series A Funded start-up in California is seeking a Systems Engineer to design and optimize systems handling complex ML pipelines. The role involves building scalable infrastructure, developing CI/CD pipelines, and ensuring system performance. Key qualifications include...Full time- ...and help build the platform engineers turn to to ship AI... ...Baseten is building its own GPU infrastructure for large-scale inference. As we move into... ...workload symptoms that look like network problems, but are not.... ...networking and inference software. RDMA data paths, GPUDirect...Full timeFlexible hours
- ...largest, most reliable, and frictionless GPU fleet to support OpenAI’s general purpose... ...Much more! About the Role As an engineer within Fleet infrastructure, you will design... ...one of the world’s largest GPU fleet. The scale is immense, the timelines are tight, and...Full timeWork at officeRelocation package
- ...serverless runtime that launches GPU‑backed containers in less than 1 second and quickly scales out to thousands of GPUs.... ...Deploy and validate data center network infrastructure (front‑end, back... ...Operations, ICT, Hardware, and Network Engineering to identify blockers early,...
$200k
...Ready to architect the high-speed networks powering the AI era? Join a trailblazing leader in GPU-accelerated computing, designing... ..., or SaltStack. High-Scale Networking: A strong foundation... ...specific experience with InfiniBand or RDMA (RoCEv2). CI/CD Pipeline Expertise...Permanent employment$125k - $145k
...urgency, who believe in the scale of our ambition and... ...offering sustainable, low-cost GPU compute power. As a Senior Cloud Support Engineer, you'll play a crucial... ...Work closely with SRE, Networking, and Storage teams from... ...such as Infiniband, RDMA, RoCE, and Software Defined...Full timeTemporary work- A leading streaming platform is looking for a Staff DevOps Engineer in San Francisco, CA, to automate and scale systems supporting their streaming services. The role involves leading projects on infrastructure automation, best-practice adoption, and collaboration across...Full time
- ...Senior Site Reliability Engineer Location: Global... ...access to the kind of scaled AI infrastructure once... ...building the systems, network, and orchestration layer... ...and debug large-scale GPU infrastructure used for... ...infrastructure (ECC errors, NVLink degradation, NCCL...Full timeRemote work
$300 per month
...intelligence. We’re crafting the engine that powers a world... ...— without sacrificing scale, speed, or... ...systems that deliver raw GPU throughput via zero-latency InfiniBand/RDMA fabrics, ensuring that massive... ...ways of managing memory, networking, and compute that don't...$220k - $290k
...move from idea to production, and do it at scale without compromise. For developers and... ...prospective customers. Collaborate with engineering and product development teams to tailor solutions... ...optimization, MLOps practices, and GPU infrastructure. Professional proficiency...Full timeWork at office- A leading identity solutions provider in San Francisco is looking for a Solutions Engineer to act as a technical advisor for customers, guiding them through the implementation of Persona's platform. This position involves solving complex business challenges through technical...
$250k
...opportunities? Join a rapidly scaling AI cloud infrastructure... ...provider building a next-generation GPU platform designed for AI... ...Senior / Staff Site Reliability Engineer to support and scale large-... ...providers Strong understanding of networking fundamentals (DNS, TCP/IP,...Permanent employmentRemote work$300 per month
...who believe in the scale of our ambition and... ...Staff Hardware Systems Engineer to strengthen... ...across Crusoe Cloud’s GPU- and CPU-based... ..., memory, storage, networking, accelerators, and... ...PCIe, InfiniBand, or NVLink. ~ Hands-on experience... ...experience with RDMA, RoCE, CXL, NVLink...Temporary work- ...About the Team The Scaling team is responsible for... ...architectural and engineering backbone of OpenAI’s... ...spans system software, networking, platform architecture... ...system, including CPU, GPU, memory subsystem,... ...including WAN traffic, NVlink and RDMA collectives), storage...Full time
- ...efficiently, reliably, and at scale. We build and optimize the systems... ...the Role We’re hiring engineers to scale and optimize OpenAI’s... ...infrastructure across emerging GPU platforms. You’ll work across... ...inference workloads across memory, network, and compute layers....Full time
- A leading identity solutions provider in San Francisco is looking for a Solutions Engineer to act as a technical advisor for customers, guiding them through the implementation of Persona's platform. This position involves solving complex business challenges through technical...Full time
$135.61k - $184.04k
...Network Engineer Employment Type: Full Time, Experienced level Department: Information Technology CGS is seeking an experienced Network Engineer... ...on the evaluation, enhancement, and maintenance of a large scale network project that includes both wired and wireless network...Full timeLocal areaMonday to FridayFlexible hours$225k - $275k
...urgency, who believe in the scale of our ambition and thrive on... ...is seeking a Senior Staff Network Production Engineer to own production reliability... ...backbone, data center fabric, and GPU cluster interconnects. You... ...Arbor. ~ GPU Cluster and RDMA Networking: Hands-on...Temporary work- ...research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance,... ...unchecked growth. About the role As a software engineer on the Fleet High Performance Computing (HPC) team,...Full time
$300 per month
...urgency, who believe in the scale of our ambition and... ...highly skilled and motivated GPU Fleet Operations Engineer to join Crusoe’s Fleet... ...systems, interconnects, and networking hardware. Conduct post... ...such as InfiniBand, NVLink, and RDMA over Converged Ethernet (...Temporary workWork at office- A growing AI startup in San Francisco is seeking a Staff Software Engineer to design, build, and scale backend systems for their AI agent platform. The ideal candidate has over 7 years of backend or full-stack engineering experience. You will lead complex features from...Full timeImmediate start
$179k - $218k
...a sense of urgency, who believe in the scale of our ambition and thrive on a path not... ...seeking a Senior Staff Data Center Operations Engineer, GPU Hardware Architecture to be the... ...identifying "pre-failure" patterns in HBM or NVLink components before they impact customer...Temporary work- ...Experienced level Department: Information Technology CGS is seeking an experienced Network Engineer to join a team focused on the evaluation, enhancement, and maintenance of a large scale network project that includes both wired and wireless network infrastructure and...Full timeLocal areaMonday to FridayFlexible hours
- ...advertisers. Our platform is purpose-built to deliver outcomes at scale, not just impressions. We are a team of builders, operators,... ...what it becomes. Role Overview We are looking for a skilled Network Engineer to join our San Francisco engineering hub. In this role, you...Immediate start
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to GPU Network Engineer: RDMA/NVLink at Scale. Be the first to apply!
- network infrastructure engineer San Francisco, CA
- network software engineer San Francisco, CA
- network engineer full time San Francisco, CA
- juniper network engineer San Francisco, CA
- network applications engineer San Francisco, CA
- wireless network engineer San Francisco, CA
- network developer San Francisco, CA
- cisco network engineer San Francisco, CA
- network engineer - transport San Francisco, CA
- network consulting engineer San Francisco, CA







