Get new jobs by email
$133.1k - $306.4k
Leads the design, architecture, engineering, and operational strategy for large-scale RDMA backend fabrics supporting AI, HPC, and cloud infrastructure. This role is responsible for building and scaling high-performance, low-latency network fabrics, driving end-to-end...SuggestedTemporary workFlexible hours$90k
Kforce has a client in West Palm Beach, FL that is seeking a Network DevOps Engineer RDMA Fabric Automation. Key Responsibilities: * Automate deployment and operations of large-scale RDMA (RoCEv2) Ethernet fabrics across Vultr data centers. * Build Ansible and Python-...Suggested$90k - $130k
...thinks programmatically and treats networks as software systems Key Responsibilities - Automate deployment and operations of large-scale RDMA (RoCEv2) Ethernet fabrics across Vultr data centers. - Build Ansible and Python-based frameworks to provision, validate, and...SuggestedFull timeWork at officeImmediate startRemote workFlexible hours- ...or hyperscale compute deployments. Deep knowledge of InfiniBand, RoCE, NVIDIA networking technologies (NCCL, NVLink, GPU Direct RDMA), and large GPU cluster architectures. Experience with AI/ML training and inference infrastructure at scale. Strong background...SuggestedPermanent employmentFull timeWeekend work
- ...segmentation, or zero-trust architectures Bonus Experience with any of the following is a plus: Kernel bypass / SmartNIC / SR-IOV / RDMA Custom BPF programs for observability or dataplane filtering Large-scale IPv6-only environments Advanced traffic simulation...SuggestedFull timeWorldwide
$180k - $250k
...BGP, tcpdump) Experience with NVIDIA GPU infrastructure: driver management, health monitoring, DCGM, NVLink/NVSwitch diagnostics, RDMA, InfiniBand/RoCEv2 Experience with AMD GPUs Experience with bare metal and VM provisioning (PXE/iPXE, Kickstart, libvirt, Qemu/...SuggestedFull timeLocal areaRelocation package$240k - $310k
...High-Throughput I/O: Experience with cutting-edge I/O architectures like DAOS or SPDK . Networking Foundations: Background in RDMA and high-performance networking, including SmartNICs and RoCEv2 . Distributed Systems Mastery: Experience with highly...SuggestedFull timeTemporary work$188k - $275k
...such as vLLM, Triton, TensorRT-LLM, Ray Serve, or TorchServe Experience with GPU systems and performance optimization (CUDA, NCCL, RDMA, NUMA, GPU interconnects) Experience leading multi-team or org-level technical initiatives Exposure to large-scale AI/ML...SuggestedPermanent employmentFull timeTemporary workCasual workWork at officeRemote workFlexible hours$188k - $275k
...CSI drivers, dynamic provisioning, snapshots, resize and RWO block volumes backed by public SLOs. Work with technologies such as RDMA, GPU Direct Storage, RoCE, InfiniBand, SPDK and distributed filesystems to raise storage performance and efficiency. Drive reliability...SuggestedPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours- ...-performance storage and networking for AI workloads — parallel filesystems, object storage tiers, and high-throughput, low-latency RDMA fabrics. Operate Kubernetes clusters underpinning AI workloads — namespaces, RBAC, resource quotas, network policies, storage classes...SuggestedFull timeWork experience placementLocal area
$166k - $201k
...storage solutions, such as Parallel Filesystems or petabyte+ scale Object Storage. Familiarity with networking technologies like RDMA and Infiniband. Familiarity with modern storage technologies (e.g GPU Direct Storage, F2FS, SPDK etc) Prior experience with Nvidia...SuggestedFull timeTemporary work- ...creating metrics that will help prioritize the focus of the team and your own. Tech Stack Python Go TCP/IP BGP RDMA Annual Salary Range $150,000 - 250,000k Benefits Base salary is just one part of our total rewards package at xAI,...SuggestedFull timeTemporary workRemote work
$139k - $204k
...compatible object storage and integrate dedicated storage clusters into diverse customer environments. Work with technologies such as RDMA, GPU Direct Storage, and distributed filesystems protocols such as NFS or FUSE to optimize storage performance and efficiency....SuggestedPermanent employmentFull timeTemporary workCasual workWork at officeRemote workFlexible hours$215k - $285k
...Background building serverless compute, GPU clouds, virtualization platforms, or hyperscaler infrastructure. Experience with RDMA or other high-performance networking technologies. Familiarity with GPUDirect. Experience with NVLink, multi-GPU interconnects...SuggestedFull timeWork at office$175k - $287k
...and updating internally maintained versions of deep learning frameworks and their companion libraries like CUDA, cuTile, cuDNN, NCCL, RDMA, Tensorflow, PyTorch, TorchRec, Flash Attention, PyTorch Lightning and more. Feature Engineering: this team shapes the future of...SuggestedFull timeFor contractorsWork experience placementWork at officeFlexible hours$110k - $135k
...and REST APIs.Foundational understanding of software architecture, debugging, troubleshooting, and performance analysis.Nice to have: RDMA, RoCE, VXLAN, K8s, Git, Redis, Grafana, AI-assisted development tools (Cursor).Strong analytical and problem-solving skills.Ability...Permanent employmentWorldwide- ...NVIDIA GPU Operator, device plugins, or Dynamic Resource Allocation (DRA). High-performance networking for distributed workloads: RDMA/RoCE, InfiniBand, EFA, SR-IOV, or NCCL tuning. Multi-cloud, hybrid-cloud, or bare-metal Kubernetes operations. Contributions...Permanent employmentFull timeFlexible hours
- ...network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-performance GPU interconnects Define and operationalize networking protocols, readiness...Full time
- ...of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts. Deep-dive performance on distributed training/inference: Collective...Full time
- ...inform our future supercomputer network designs. You might thrive in this role if you: Have written distributed algorithms using RDMA in the past. Are comfortable writing low level performance sensitive CPU and/or GPU code. Are familiar with network simulation...Full timeWork at officeRelocation package
- ...spreading, unfairness, etc. Strong skills in socket programming will enable you to deploy robust management operations, while deftness in RDMA Verbs will enable you to optimize CPU resources and shape network traffic to deliver real world impact on AI performance metrics, as...Full time
- ...Experience with AI frameworks such as PyTorch Distributed, DeepSpeed, and Megatron-LM. Familiarity with libfabric/OFI, UCX, and RDMA concepts. Experience with RoCEv2 and Ultra Ethernet. Experience building cluster-scale performance test infrastructure. Location...Remote jobFull timeFlexible hours
$165k - $225k
...clusters and distributed workloads. High-Performance Networking: Collaborate with infrastructure to optimize network paths for RDMA, RoCE, and GPU-to-GPU communication, ensuring minimal latency and maximum throughput for distributed training and large-scale computational...Full timeImmediate startRemote workFlexible hours$139k - $204k
...CD, and observability stacks (Prometheus, Grafana, OpenTelemetry). ~ Experience with performance-critical GPU systems (CUDA, NCCL, RDMA, NVLink/PCIe, memory bandwidth) and model-serving stacks (llm-d, vLLM, TensorRT-LLM, Megatron-LM). ~ Strong communicator...Permanent employmentFull timeTemporary workCasual workWork at officeRemote workFlexible hours$320k
...virtualization, device drivers, firmware, or hardware health/diagnostics daemons ~ Familiarity with high-performance networking (EFA, RDMA, InfiniBand) for distributed ML workloads. ~ Demonstrated ownership of production reliability for high-throughput, latency-...Full timeWork at officeVisa sponsorshipFlexible hours- ...control. Experience with multi-GPU and multi-node inference, including tensor parallelism, pipeline parallelism, expert parallelism, RDMA, collective communication, and failure handling. Experience optimizing Mixture-of-Experts or multimodal models. Knowledge of...Full time
$300 per month
...200, GB200, B200 and AMD 350X / 355X series platforms. Familiarity with high-speed interconnects such as InfiniBand, NVLink, and RDMA over Converged Ethernet (RoCE). Strong Linux experience (Ubuntu, Rocky Linux, CentOS) using the command line for diagnostics and...Full timeTemporary workWork at office$143k - $210k
...compatible object storage and integrate dedicated storage clusters into diverse customer environments.Work with technologies such as RDMA, GPU Direct Storage, and distributed filesystems protocols such as NFS or FUSE to optimize storage performance and efficiency.Lead...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours- ...object storageHigh Availability: Pacemaker/CorosyncProgramming: C/C++, plus Go/Python/Bash scriptingNetworking: TCP/IP, Ethernet, IB, RDMA/RoCEHardware: SAS, SCSI, NVMe, HBAsAPIs: REST/ gRPC, Redfish/SwordfishContainerization/Virtualization: Docker/Podman, KVM/...Full timeWork experience placementWork at officeLocal areaImmediate start
$224k - $356.5k
...activitiesPractical knowledge of sophisticated networking for AI data centers, including InfiniBand and Ethernet fabric topologies, RDMA protocols, host networking and switches.Experience with end-to-end network performance debugging - hosts, switches, and optics.Effective...Full time
