GPU DC East-West Network SRE Expert (SME) [Remote]
Bitdeer
- Remote job
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.
Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.
Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
To learn more, visit [ About the RoleYou keep the fabric that makes 10K GPUs act like one — and turn IB/RoCE telemetry into the ground truth for our congestion and link-failure predictors.
Bitdeer is building an AI-operated GPU cloud where East-West bandwidth is the difference between a healthy training job and a $50M training run stalled by a bad optic. In this role you operate the InfiniBand and RoCEv2 fabrics that carry NCCL traffic across NeoCloud's US DCs, and you feed the AIOps substrate with the fabric telemetry it needs to catch link degradation, congestion, and topology drift before they land on the pager.
What you'll own- InfiniBand fabrics: fat-tree, rail-optimized, and dragonfly topologies for GPU clusters of 100–10,000 GPUs.
- RoCEv2 networks across Nvidia, Arista, and Cisco platforms for RDMA workloads.
- UFM (Unified Fabric Manager) for IB fabric monitoring, diagnostics, and subnet management.
- IB and RoCE performance monitoring and tuning: adaptive routing, congestion control (DCQCN/ECN), traffic isolation.
- NCCL communication tuning: topology detection, ring/tree algorithm selection, GDR configuration.
- Firmware lifecycle across IB switches and HCAs.
- Fault diagnosis: link flaps, symbol errors, packet drops, routing anomalies, credit stalls.
- Coordination with Nvidia/Mellanox support for escalations, bugs, and RMA.
- Wire IB/RoCE telemetry (ibdiagnet, perfquery, ibstat, PortRcvErrors, PortXmitDiscards, adaptive-routing state) into the platform's collection pipeline.
- Partner with the platform team to define the Link and Straggler predictors: what a "bad optic 30 minutes from failure" looks like in the counters.
- Convert every incident into a labeled example the fault-prediction engine can learn from — and every routine mitigation into a workflow the remediation actuator can run.
- 5+ years in data center networking, with at least 3 years focused on InfiniBand or RoCE fabrics
- Hands-on experience deploying and operating Nvidia/Mellanox InfiniBand switches at scale
- Strong understanding of IB subnet management, partitioning, and QoS
- Experience with RoCEv2 deployment including PFC, ECN, DCQCN configuration
- Proficiency with UFM or equivalent IB fabric management tools
- Knowledge of 400G/800G optics, cabling standards, and structured cabling best practices
- Experience diagnosing IB/RoCE network issues using ibdiagnet, perfquery, ibstat, and similar tools
- Understanding of NCCL and how GPU communication maps to network topology
- Instinct for telemetry-driven ops — you've either built dashboards/alerts on RDMA counters at scale, or you can articulate the feature set a fabric-health model would need.
- Runbook-as-code mindset — the diagnostics you run today should become automation next quarter.
$180k - $320k
...policy. NeoCloud is building an AI-operated GPU cloud spanning 4 US DCs, APAC sites, and... ...— DCI, WAN, internet edge, tenant-facing networks — and you build the automation that lets... ...traffic engineering, and capacity planning for DC networks ~ AIOps aptitude — you've...NetworkRemote jobFull timeLocal area- ...execute. NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all of that... ...: GPU locality, NVLink domain awareness, network rail affinity. Custom Resource... ...GitOps workflows (ArgoCD/Flux) ~ Strong SRE background: SLI/SLO frameworks, incident management...NetworkRemote jobFull timeLocal area
- ...stalls a $50M training run. Bitdeer is building an AI-operated GPU cloud. Storage is where AI workloads either fly or fall over: a... ...direct GPU-to-storage data paths. Deploy and manage storage networking (NFS over RDMA, NVMe-oF, high-speed storage fabrics) and Nvidia...NetworkRemote jobFull timeLocal area
- ...your career.THE TEAM:AMD's Data Center GPU organization is transforming the... ...for its Tier 1 CSP Business Group (DC GPU CBG), focused on driving our business... ..., media, analysts, technical experts and senior executivesPossess a network of industry relationships with potential...NetworkWork experience placement
- Job Title: GPU Network Engineer Job Location: Sunnyvale, CA (hybrid)Job Salary: 200k-250k + BenefitsRequirements: Data center, HPC networking... ..., please read on!What You Will Be Doing:Design and deploy east-west network fabrics for GPU-to-GPU, rack-to-rack, and cluster-to-...NetworkLocal areaRelocation
- Prolific is seeking fluent Hindi speakers for its Expert Network in San Jose, California. The role involves completing AI training tasks, judging AI performance, and improving AI models using legal expertise. Candidates must have strong English skills, attention to detail...NetworkRemote jobSelf employmentWork from homeFlexible hours
- ...supports critical infrastructure at world-class data center facilities, performing install, configure, and maintenance of server and network equipment. You'll handle cabling, environmental monitoring, incident response, and documentation, with on-call rotation and night/...NetworkFull timeNight shiftRotating shiftWeekend work
- ...beyond. Together, we advance your career. THE TEAMAMD's Data Center GPU organization is transforming the AI and HPC landscape. Our... ...model architectures (e.g., LLMs, transformer variants, graph neural networks), datatypes, and scaling methodologies to anticipate future...NetworkRemote work
$167.7k - $245.2k
...looking to excel in a fast-paced environment creating groundbreaking network management solutions and want to make a significant long-... ...with the ability to translate complex technical concepts for non-experts while fostering collaboration and driving excellence within...NetworkFull timeTemporary workLocal areaFlexible hours- ...NeoCloud is building an AI-operated GPU cloud. That doesn't mean fewer... ...Monitor GPU cluster health, network status, storage systems, and... ...Collect diagnostic data for L2/SME escalation: logs, DCGM output,... ...ServiceNow/Jira). Perform physical DC tasks: cable installation,...NetworkFull timeLocal areaShift workNight shift
- ...for an AI Devops Infrastructure Engineer/GPU Infrastructure Engineer. Job Title:... ...foundation in infrastructure engineering, DevOps/SRE, platform engineering, or similar... ...containerized environments, including compute, networking, storage, and accelerator scheduling....NetworkContract work
- ...seeking a visionary and hands-on Cloud SRE Architect to lead the design,... ...the end-to-end architecture across CPU, GPU, RDS, storage, networking, serverless, and AI services, ensuring... ...frameworks, and three operational tiers (Edge DC → Regional Controller → Global Hub)....NetworkRemote jobFull timeContract workLocal areaShift work
- ...Site Reliability Engineer (SRE) Location: Santa Clara Valley (Cupertino), California, Hybrid. Duration: 6+ Months Job Description... ...: Puppet, Chef, Ansible, or Salt. Understanding of standard networking protocols and components such as: DNS, ECMP, TCP/IP, ICMP, the...Network
- ...bounded contexts of the NeoCloud SRE platform — the multi-region... ...observes, protects, and operates a GPU rental fleet across self-built... ...probe. Hardware Lifecycle & DC Ops: hardware-lifecycle, dc-... ...optimizer, gpu-efficiency-dashboard, network-stability-dashboard, patching-...NetworkRemote jobFull timeContract workLocal area
- .... Together, we advance your career. THE TEAM:AMD's Data Center GPU organization is transforming the AI and HPC landscape. Our mission... ...Familiarity with data center infrastructure including networking, cooling, and platform design Experience in product operations,...NetworkRemote workFlexible hours
- ...SRE Role - SRE Location, RTP/NC and San Jose, CA Duration - Fulltime Job Description: Must Have Technical/Functional... ...in K8s Mandatory and good knowledge with K8s storage and networking. Should have deployed applications in Kubernetes. Good knowledge...NetworkFull time
- ...Head of GPU Cloud About the Company Building software foundation for next-generation accelerated computing platform for large... ...experience with GPU memory behavior and high-bandwidth networking is essential. The role also requires a leader with strong technical...Network
- ...Job Description Job Description GPU Programmer / Software Engineer Job Type: Contractor Location: Remote Job Overview We are seeking experienced GPU Programmers / Software Engineers to design and optimize GPU-based tasks for AI and LLM applications. You...Remote jobFor contractors
- ...Position: Site Reliability Engineering (SRE) Location: Santa Clara, CA (Onsite) Duration: W2 / C2C Contract Experience: 10... ...infrastructure experience including EC2, SSM, vulnerability management ,VPC networking, Web Application Firewalls, ECS/Fargate, IAM • WS Terraform...NetworkContract workImmediate start
- ...gWMS Functional Expert San Jose, CA onsite 1) WH Automation 2) gWMS Feature requests Understanding logistics business... ...Logistics Operations, Systems Engineering, Project Management, IT Network/Infrastructure teams, and third-party logistics (3PL) providers...Network
- ...optical transceivers in advanced CMOS PDKs. We are seeking an expert in analog/mixed-signal circuit design and architecture.THE PERSON... ...and correction), and analog front-end circuits (pad matching networks, ESD, CTLE, linear amplifier)To define circuit-driven micro-architectures...Network
- ...Senior Systems Software Engineer – GPU Performance at Scale We are looking for a dedicated engineer for the Senior Systems Software... ...next‑generation datacenter builds for NVIDIA GPUs, CPUs, and networking hardware. Engage early with HW/FW/SW/platform internal and...Network
$184k - $287.5k
...of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self‑driving cars... ...with next-generation datacenter builds for NVIDIA GPUs, CPUs, and networking hardware. Engage early with HW/FW/SW/platform internal and...Network$138k - $184k
...Enterprise Account Executive - West Eightfold was founded with a vision to solve for employment in our society. For decades, the... ...been based on who the individuals are and the strength of their network, not their potential. Eightfold leverages artificial intelligence...NetworkWork at officeRemote workFlexible hours$115k - $130k
...provider of advanced server, storage, and networking solutions for Data Center, Cloud... ...support, and troubleshooting connectivity for GPU clusters and storage systems. Essential... ...architectures, including spine-leaf topologies and east-west traffic patterns. Basic proficiency with...NetworkWork at officeWorldwideNight shift$196k - $310.5k
...next era of computing. An era in which our GPU acts as the brains of computers, robots,... ...on the world. NVIDIA Enterprise Network Architecture team is seeking experienced... ...implicitly-trusted model toward explicit east-west segmentation — internal/collapsed firewalls...NetworkWork at office$27 - $41 per hour
...dedicated year-round TurboTax Local Service Experts in one of our new TurboTax locations... ...grass-roots marketing initiatives, referral networks, and community partnerships. Serve as the... ...Washington: $27.00 - $41.00 Washington, DC: $26.00 - $39.00 This position will be eligible...NetworkWork at officeLocal areaMonday to Friday$86.02k - $117.71k
...Pilot (First Officer) - West Region Company: NetJets Aviation, Inc. Area of Interest: Crewmembers Location: US Los Angeles, CA, US... ...including Medical, Dental, and Vision benefits, with access to robust networks of nationwide providers. NetJets offers benefits so you can LIVE...NetworkTemporary workFlexible hoursNight shiftWeekend workWeekday work- ...Job Description Job Description Developer & Infrastructure Expert Role Type: Contractor Location: Remote Job Overview We... ...workflows across software development, cloud infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,...Remote jobFor contractors
$154.9k - $225.3k
...protect enterprise, data center, and cloud networks from evolving cyber threats. This team... ...environment. You will act as a subject matter expert, mentoring junior engineers and... ...Proficiency in selecting and implementing DC-DC converters, LDOs, and managing power sequencing...NetworkFull timeTemporary workLocal areaFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to GPU DC East-West Network SRE Expert (SME) [Remote]. Be the first to apply!
- guest service support expert San Jose, CA
- fulfillment expert San Jose, CA
- technology expert San Jose, CA
- rn network San Jose, CA
- director of network operations San Jose, CA
- networking San Jose, CA
- network cabling San Jose, CA
- compass health network San Jose, CA
- staffing network San Jose, CA
- data network cabling San Jose, CA




