Get new jobs by email
$220.92k - $311.89k
...this role, you will define how distributed systems move data efficiently, reliably, and at scale. You will drive the architecture of RDMA transports, congestion management frameworks, reliability mechanisms, and collective communication optimizations that directly...SuggestedFull timeLocal areaImmediate startShift work$90k - $130k
DescriptionKforce has a client in West Palm Beach, FL that is seeking a Network DevOps Engineer RDMA Fabric Automation.Key Responsibilities:* Automate deployment and operations of large-scale RDMA (RoCEv2) Ethernet fabrics across Vultr data centers.* Build Ansible and...Suggested$133.1k - $306.4k
Leads the design, architecture, engineering, and operational strategy for large-scale RDMA backend fabrics supporting AI, HPC, and cloud infrastructure. This role is responsible for building and scaling high-performance, low-latency network fabrics, driving end-to-end...SuggestedTemporary workFlexible hours$350k
Thinking Machines Lab is seeking a Network Engineer in San Francisco to manage and improve our GPU network fabric. The role requires in-depth knowledge of large-scale deployments and the ability to debug complex network issues. A collaborative environment is emphasized,...SuggestedVisa sponsorship- ...oversee the feeding path from Linux host networking through the EIC PCIe bridge, to NICs, and back via DMA into HBM. The role demands RDMA production experience, kernel networking fluency, and strong architectural writing skills. The ideal candidate has 10+ years in...Suggested
- Oracle seeks a senior Network Engineer to support global OCI networking, focusing on RDMA cluster architectures for AI/ML/HPC workloads. You’ll design, deploy, and operate scalable fabrics across a global footprint, with emphasis on protocol mastery, automation, and incident...Suggested
$123k - $163k
...modern data center networking, including EVPN-VXLAN, BGP, MLAG, QoS, and traffic engineering. We require deep knowledge of RoCEv2, RDMA transport tuning, ECN/PFC, and lossless Ethernet design. We value hands-on experience with automation tools such as Ansible and...SuggestedHourly payFull time- ByteDance is looking for a graduate or recent PhD graduate to design and build high-speed network technologies. This role, scheduled from January 5, 2026, to December 14, 2026, requires passion for networking and proficiency in relevant programming languages. Successful...Suggested
$148.2k - $300.96k
...speed IP networking, hardware‑software interaction, and hardware offloading technologies. Preferred Qualifications Familiarity with RDMA/RoCE network protocol. Experience developing software systems for large‑scale data center networks or distributed systems. Compensation...SuggestedTemporary workLocal areaWorldwide$177.3k - $265.9k
...experience (e.g., queues, RSS, offloads, AF_XDP, SmartNICs/BlueField). High-performance I/O: DPDK , SPDK , NVMe/NVMe-oF , RDMA/RoCE , iSCSI , storage stack tuning. Observability stacks (e.g., perf ETM, BPF ring buffers, Prometheus/Grafana, OpenTelemetry...SuggestedFull time- ...-performance storage and networking for AI workloads — parallel filesystems, object storage tiers, and high-throughput, low-latency RDMA fabrics. Operate Kubernetes clusters underpinning AI workloads — namespaces, RBAC, resource quotas, network policies, storage classes...SuggestedFull timeWork experience placementLocal area
- ...NVIDIA GPU Operator, device plugins, or Dynamic Resource Allocation (DRA). High-performance networking for distributed workloads: RDMA/RoCE, InfiniBand, EFA, SR-IOV, or NCCL tuning. Multi-cloud, hybrid-cloud, or bare-metal Kubernetes operations. Contributions...SuggestedPermanent employmentFull timeFlexible hours
- ...network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-performance GPU interconnects Define and operationalize networking protocols, readiness...SuggestedFull time
- ...of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts. Deep-dive performance on distributed training/inference: Collective...SuggestedFull time
- ...inform our future supercomputer network designs. You might thrive in this role if you: Have written distributed algorithms using RDMA in the past. Are comfortable writing low level performance sensitive CPU and/or GPU code. Are familiar with network simulation...SuggestedFull timeWork at officeRelocation package
$325k
...experience with one or more ML hardware accelerators (GPUs, TPUs, Trainium). Understand ML-specific networking optimizations like RDMA and InfiniBand. Have expertise in AI-specific observability tools and frameworks. Have experience with chaos engineering and...Full timeWork at officeVisa sponsorshipFlexible hours$166k - $201k
...storage solutions, such as Parallel Filesystems or petabyte+ scale Object Storage. Familiarity with networking technologies like RDMA and Infiniband. Familiarity with modern storage technologies (e.g GPU Direct Storage, F2FS, SPDK etc) Prior experience with Nvidia...Full timeTemporary work$139k - $204k
...compatible object storage and integrate dedicated storage clusters into diverse customer environments. Work with technologies such as RDMA, GPU Direct Storage, and distributed filesystems protocols such as NFS or FUSE to optimize storage performance and efficiency....Permanent employmentFull timeTemporary workCasual workWork at officeRemote workFlexible hours$188k - $275k
...CSI drivers, dynamic provisioning, snapshots, resize and RWO block volumes backed by public SLOs. Work with technologies such as RDMA, GPU Direct Storage, RoCE, InfiniBand, SPDK and distributed filesystems to raise storage performance and efficiency. Drive reliability...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours- ...creating metrics that will help prioritize the focus of the team and your own. Tech Stack Python Go TCP/IP BGP RDMA Annual Salary Range $150,000 - 250,000k Benefits Base salary is just one part of our total rewards package at xAI,...Full timeTemporary workRemote work
- ...control. Experience with multi-GPU and multi-node inference, including tensor parallelism, pipeline parallelism, expert parallelism, RDMA, collective communication, and failure handling. Experience optimizing Mixture-of-Experts or multimodal models. Knowledge of...Full time
$188k - $275k
...such as vLLM, Triton, TensorRT-LLM, Ray Serve, or TorchServe Experience with GPU systems and performance optimization (CUDA, NCCL, RDMA, NUMA, GPU interconnects) Experience leading multi-team or org-level technical initiatives Exposure to large-scale AI/ML...Permanent employmentFull timeTemporary workCasual workWork at officeRemote workFlexible hours- ...own the in-chassis IO subsystem consisting of i) several cluster-facing RoCE v2 network interfaces via a custom implementation of the RDMA protocol; ii) a large programmable switching fabric; and iii) Serial IO communication with the Cerebras WSE via a proprietary...
$90k - $110k
...application benchmarks based on enterprise computing requirements and generic server architecture, key examples include HammerDB, MySQL, RDMA, flashed based storage (NVDIMM, and NVMe)• Execute benchmark testing and build proof of concepts based on latest data center...Worldwide$215k - $285k
...Background building serverless compute, GPU clouds, virtualization platforms, or hyperscaler infrastructure. Experience with RDMA or other high-performance networking technologies. Familiarity with GPUDirect. Experience with NVLink, multi-GPU interconnects...Full timeWork at office$139k - $204k
...CD, and observability stacks (Prometheus, Grafana, OpenTelemetry). ~ Experience with performance-critical GPU systems (CUDA, NCCL, RDMA, NVLink/PCIe, memory bandwidth) and model-serving stacks (llm-d, vLLM, TensorRT-LLM, Megatron-LM). ~ Strong communicator...Permanent employmentFull timeTemporary workCasual workWork at officeRemote workFlexible hours$165k - $225k
...platforms with custom scheduling or resource management Knowledge of high-performance networking for GPU communication (InfiniBand, RDMA, NVLink, NVSwitch) Familiarity with AI/ML training frameworks (PyTorch, TensorFlow) and their infrastructure requirements...Full timeImmediate startRemote workFlexible hours- ...actionable diagnosis while an incident is still unfolding. This role sits at the boundary between networking and inference software. RDMA data paths, GPUDirect transfers, prefill/decode disaggregation, KV cache movement, request routing, and workload backpressure can...Full timeFlexible hours
$153k - $204k
...compatible object storage and integrate dedicated storage clusters into diverse customer environments.Work with technologies such as RDMA, GPU Direct Storage, and distributed filesystems protocols such as NFS or FUSE to optimize storage performance and efficiency.Lead...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours- ...enterprise or hyperscale compute deployments.Deep knowledge of InfiniBand, RoCE, NVIDIA networking technologies (NCCL, NVLink, GPU Direct RDMA), and large GPU cluster architectures.Experience with AI/ML training and inference infrastructure at scale.Strong background in...Permanent employmentWeekend work

