Get new jobs by email
  • $133.1k - $306.4k

    Leads the design, architecture, engineering, and operational strategy for large-scale RDMA backend fabrics supporting AI, HPC, and cloud infrastructure. This role is responsible for building and scaling high-performance, low-latency network fabrics, driving end-to-end... 
    Suggested
    Temporary work
    Flexible hours

    Oracle Corporation

    Austin, TX
    5 days ago
  • $90k

    Kforce has a client in West Palm Beach, FL that is seeking a Network DevOps Engineer RDMA Fabric Automation. Key Responsibilities: * Automate deployment and operations of large-scale RDMA (RoCEv2) Ethernet fabrics across Vultr data centers. * Build Ansible and Python-... 
    Suggested
    West Palm Beach, FL
    2 days ago
  • $90k - $130k

     ...thinks programmatically and treats networks as software systems Key Responsibilities - Automate deployment and operations of large-scale RDMA (RoCEv2) Ethernet fabrics across Vultr data centers. - Build Ansible and Python-based frameworks to provision, validate, and... 
    Suggested
    Full time
    Work at office
    Immediate start
    Remote work
    Flexible hours

    Vultr

    Remote
    a month ago
  •  ...or hyperscale compute deployments. Deep knowledge of InfiniBand, RoCE, NVIDIA networking technologies (NCCL, NVLink, GPU Direct RDMA), and large GPU cluster architectures. Experience with AI/ML training and inference infrastructure at scale. Strong background... 
    Suggested
    Permanent employment
    Full time
    Weekend work

    X.Ai

    Los Angeles, CA
    17 hours ago
  •  ...segmentation, or zero-trust architectures Bonus Experience with any of the following is a plus: Kernel bypass / SmartNIC / SR-IOV / RDMA Custom BPF programs for observability or dataplane filtering Large-scale IPv6-only environments Advanced traffic simulation... 
    Suggested
    Full time
    Worldwide

    Blaxel

    San Francisco, CA
    1 day ago
  • $180k - $250k

     ...BGP, tcpdump) Experience with NVIDIA GPU infrastructure: driver management, health monitoring, DCGM, NVLink/NVSwitch diagnostics, RDMA, InfiniBand/RoCEv2 Experience with AMD GPUs Experience with bare metal and VM provisioning (PXE/iPXE, Kickstart, libvirt, Qemu/... 
    Suggested
    Full time
    Local area
    Relocation package

    Falò

    San Francisco, CA
    1 day ago
  • $240k - $310k

     ...High-Throughput I/O: Experience with cutting-edge I/O architectures like DAOS or SPDK . Networking Foundations: Background in RDMA and high-performance networking, including SmartNICs and RoCEv2 . Distributed Systems Mastery: Experience with highly... 
    Suggested
    Full time
    Temporary work

    Crusoe

    Remote
    1 day ago
  • $188k - $275k

     ...such as vLLM, Triton, TensorRT-LLM, Ray Serve, or TorchServe Experience with GPU systems and performance optimization (CUDA, NCCL, RDMA, NUMA, GPU interconnects) Experience leading multi-team or org-level technical initiatives Exposure to large-scale AI/ML... 
    Suggested
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours

    Core Weave

    Sunnyvale, CA
    1 day ago
  • $188k - $275k

     ...CSI drivers, dynamic provisioning, snapshots, resize and RWO block volumes backed by public SLOs. Work with technologies such as RDMA, GPU Direct Storage, RoCE, InfiniBand, SPDK and distributed filesystems to raise storage performance and efficiency. Drive reliability... 
    Suggested
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    Core Weave

    Bellevue, WA
    1 day ago
  •  ...-performance storage and networking for AI workloads — parallel filesystems, object storage tiers, and high-throughput, low-latency RDMA fabrics. Operate Kubernetes clusters underpinning AI workloads — namespaces, RBAC, resource quotas, network policies, storage classes... 
    Suggested
    Full time
    Work experience placement
    Local area

    Blueally

    Atlanta, GA
    1 day ago
  • $166k - $201k

     ...storage solutions, such as Parallel Filesystems or petabyte+ scale Object Storage. Familiarity with networking technologies like RDMA and Infiniband. Familiarity with modern storage technologies (e.g GPU Direct Storage, F2FS, SPDK etc) Prior experience with Nvidia... 
    Suggested
    Full time
    Temporary work

    Crusoe

    San Francisco, CA
    17 hours ago
  •  ...creating metrics that will help prioritize the focus of the team and your own. Tech Stack Python Go TCP/IP BGP RDMA Annual Salary Range $150,000 - 250,000k Benefits Base salary is just one part of our total rewards package at xAI,... 
    Suggested
    Full time
    Temporary work
    Remote work

    X.Ai

    Remote
    1 day ago
  • $139k - $204k

     ...compatible object storage and integrate dedicated storage clusters into diverse customer environments. Work with technologies such as RDMA, GPU Direct Storage, and distributed filesystems protocols such as NFS or FUSE to optimize storage performance and efficiency.... 
    Suggested
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours

    Coreweave

    Northfield, MN
    1 day ago
  • $215k - $285k

     ...Background building serverless compute, GPU clouds, virtualization platforms, or hyperscaler infrastructure. Experience with RDMA or other high-performance networking technologies. Familiarity with GPUDirect. Experience with NVLink, multi-GPU interconnects... 
    Suggested
    Full time
    Work at office

    Recruiting From Scratch

    Remote
    1 day ago
  • $175k - $287k

     ...and updating internally maintained versions of deep learning frameworks and their companion libraries like CUDA, cuTile, cuDNN, NCCL, RDMA, Tensorflow, PyTorch, TorchRec, Flash Attention, PyTorch Lightning and more. Feature Engineering: this team shapes the future of... 
    Suggested
    Full time
    For contractors
    Work experience placement
    Work at office
    Flexible hours

    Linkedin

    Remote
    1 day ago
  • $110k - $135k

     ...and REST APIs.Foundational understanding of software architecture, debugging, troubleshooting, and performance analysis.Nice to have: RDMA, RoCE, VXLAN, K8s, Git, Redis, Grafana, AI-assisted development tools (Cursor).Strong analytical and problem-solving skills.Ability... 
    Permanent employment
    Worldwide

    Supermicro

    San Jose, CA
    1 day ago
  •  ...NVIDIA GPU Operator, device plugins, or Dynamic Resource Allocation (DRA). High-performance networking for distributed workloads: RDMA/RoCE, InfiniBand, EFA, SR-IOV, or NCCL tuning. Multi-cloud, hybrid-cloud, or bare-metal Kubernetes operations. Contributions... 
    Permanent employment
    Full time
    Flexible hours

    FriendliAI

    San Francisco, CA
    1 day ago
  •  ...network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-performance GPU interconnects Define and operationalize networking protocols, readiness... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts. Deep-dive performance on distributed training/inference: Collective... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...inform our future supercomputer network designs.  You might thrive in this role if you: Have written distributed algorithms using RDMA in the past. Are comfortable writing low level performance sensitive CPU and/or GPU code. Are familiar with network simulation... 
    Full time
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...spreading, unfairness, etc. Strong skills in socket programming will enable you to deploy robust management operations, while deftness in RDMA Verbs will enable you to optimize CPU resources and shape network traffic to deliver real world impact on AI performance metrics, as... 
    Full time

    Cerebras Systems

    Sunnyvale, CA
    1 day ago
  •  ...Experience with AI frameworks such as PyTorch Distributed, DeepSpeed, and Megatron-LM. Familiarity with libfabric/OFI, UCX, and RDMA concepts. Experience with RoCEv2 and Ultra Ethernet. Experience building cluster-scale performance test infrastructure. Location... 
    Remote job
    Full time
    Flexible hours

    Cornelis Networks, Inc.

    Austin, TX
    1 day ago
  • $165k - $225k

     ...clusters and distributed workloads. High-Performance Networking: Collaborate with infrastructure to optimize network paths for RDMA, RoCE, and GPU-to-GPU communication, ensuring minimal latency and maximum throughput for distributed training and large-scale computational... 
    Full time
    Immediate start
    Remote work
    Flexible hours

    Moon Lite, Inc.

    Chicago, IL
    17 hours ago
  • $139k - $204k

     ...CD, and observability stacks (Prometheus, Grafana, OpenTelemetry). ~ Experience with performance-critical GPU systems (CUDA, NCCL, RDMA, NVLink/PCIe, memory bandwidth) and model-serving stacks (llm-d, vLLM, TensorRT-LLM, Megatron-LM). ~ Strong communicator... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours

    Core Weave

    Remote
    1 day ago
  • $320k

     ...virtualization, device drivers, firmware, or hardware health/diagnostics daemons ~ Familiarity with high-performance networking (EFA, RDMA, InfiniBand) for distributed ML workloads. ~ Demonstrated ownership of production reliability for high-throughput, latency-... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    New York, NY
    1 day ago
  •  ...control. Experience with multi-GPU and multi-node inference, including tensor parallelism, pipeline parallelism, expert parallelism, RDMA, collective communication, and failure handling. Experience optimizing Mixture-of-Experts or multimodal models. Knowledge of... 
    Full time

    Cerebras Systems

    United States
    1 day ago
  • $300 per month

     ...200, GB200, B200 and AMD 350X / 355X series platforms. Familiarity with high-speed interconnects such as InfiniBand, NVLink, and RDMA over Converged Ethernet (RoCE). Strong Linux experience (Ubuntu, Rocky Linux, CentOS) using the command line for diagnostics and... 
    Full time
    Temporary work
    Work at office

    Crusoe

    San Francisco, CA
    1 day ago
  • $143k - $210k

     ...compatible object storage and integrate dedicated storage clusters into diverse customer environments.Work with technologies such as RDMA, GPU Direct Storage, and distributed filesystems protocols such as NFS or FUSE to optimize storage performance and efficiency.Lead... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    New York, NY
    1 day ago
  •  ...object storageHigh Availability: Pacemaker/CorosyncProgramming: C/C++, plus Go/Python/Bash scriptingNetworking: TCP/IP, Ethernet, IB, RDMA/RoCEHardware: SAS, SCSI, NVMe, HBAsAPIs: REST/ gRPC, Redfish/SwordfishContainerization/Virtualization: Docker/Podman, KVM/... 
    Full time
    Work experience placement
    Work at office
    Local area
    Immediate start

    Hewlett Packard Enterprise

    Bloomington, MN
    4 days ago
  • $224k - $356.5k

     ...activitiesPractical knowledge of sophisticated networking for AI data centers, including InfiniBand and Ethernet fabric topologies, RDMA protocols, host networking and switches.Experience with end-to-end network performance debugging - hosts, switches, and optics.Effective... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago