GPU DC East-West Network SRE Expert (SME) [Remote]
Bitdeer Technologies Group
- Remote job
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.
Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.
Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
To learn more, visit (About the Role
You keep the fabric that makes 10K GPUs act like one — and turn IB/RoCE telemetry into the ground truth for our congestion and link-failure predictors.
Bitdeer is building an AI-operated GPU cloud where East-West bandwidth is the difference between a healthy training job and a $50M training run stalled by a bad optic. In this role you operate the InfiniBand and RoCEv2 fabrics that carry NCCL traffic across NeoCloud's US DCs, and you feed the AIOps substrate with the fabric telemetry it needs to catch link degradation, congestion, and topology drift before they land on the pager.
What you'll own
- InfiniBand fabrics: fat-tree, rail-optimized, and dragonfly topologies for GPU clusters of 100–10,000 GPUs.
- RoCEv2 networks across Nvidia, Arista, and Cisco platforms for RDMA workloads.
- UFM (Unified Fabric Manager) for IB fabric monitoring, diagnostics, and subnet management.
- IB and RoCE performance monitoring and tuning: adaptive routing, congestion control (DCQCN/ECN), traffic isolation.
- NCCL communication tuning: topology detection, ring/tree algorithm selection, GDR configuration.
- Firmware lifecycle across IB switches and HCAs.
- Fault diagnosis: link flaps, symbol errors, packet drops, routing anomalies, credit stalls.
- Coordination with Nvidia/Mellanox support for escalations, bugs, and RMA.
Feed the AIOps substrate
- Wire IB/RoCE telemetry (ibdiagnet, perfquery, ibstat, PortRcvErrors, PortXmitDiscards, adaptive-routing state) into the platform's collection pipeline.
- Partner with the platform team to define the Link and Straggler predictors: what a "bad optic 30 minutes from failure" looks like in the counters.
- Convert every incident into a labeled example the fault-prediction engine can learn from — and every routine mitigation into a workflow the remediation actuator can run.
Job Requirement:
- 5+ years in data center networking, with at least 3 years focused on InfiniBand or RoCE fabrics
- Hands-on experience deploying and operating Nvidia/Mellanox InfiniBand switches at scale
- Strong understanding of IB subnet management, partitioning, and QoS
- Experience with RoCEv2 deployment including PFC, ECN, DCQCN configuration
- Proficiency with UFM or equivalent IB fabric management tools
- Knowledge of 400G/800G optics, cabling standards, and structured cabling best practices
- Experience diagnosing IB/RoCE network issues using ibdiagnet, perfquery, ibstat, and similar tools
- Understanding of NCCL and how GPU communication maps to network topology
- Instinct for telemetry-driven ops — you've either built dashboards/alerts on RDMA counters at scale, or you can articulate the feature set a fabric-health model would need.
- Runbook-as-code mindset — the diagnostics you run today should become automation next quarter.
--------------------------------------------------------------------
Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.
- ...failure predictors. Bitdeer is building an AI-operated GPU cloud where East-West bandwidth is the difference between a healthy training job... ...topologies for GPU clusters of 100–10,000 GPUs. RoCEv2 networks across Nvidia, Arista, and Cisco platforms for RDMA workloads...NetworkRemote jobFull timeLocal area
- ...policy. NeoCloud is building an AI-operated GPU cloud spanning 4 US DCs, APAC sites, and... ...— DCI, WAN, internet edge, tenant-facing networks — and you build the automation that lets... ...traffic engineering, and capacity planning for DC networks ~ AIOps aptitude — you've...NetworkRemote jobFull timeLocal area
$286.2k - $364.4k
...future of work across our Americas West region. As the architect of... ...effectively across a highly networked and matrixed organization to translate... ...Leverage and hold responsible SME functions within WPR to... ...network of doers and experts, and you’ll see that the opportunities...NetworkFull timeTemporary workLocal areaFlexible hours- ...beyond. Together, we advance your career. THE TEAMAMD's Data Center GPU organization is transforming the AI and HPC landscape. Our... ...model architectures (e.g., LLMs, transformer variants, graph neural networks), datatypes, and scaling methodologies to anticipate future...NetworkRemote work
- ...organization at the forefront of AI infrastructure is seeking a GPU Network Engineer to help design and operate the high-performance... ...network architectures that deliver low latency, high bandwidth east west traffic across distributed GPU environments. Implement and support...Network
- .... Together, we advance your career. THE TEAM:AMD's Data Center GPU organization is transforming the AI and HPC landscape. Our mission... ...Familiarity with data center infrastructure including networking, cooling, and platform design Experience in product operations,...NetworkRemote workFlexible hours
$186.9k - $267.7k
...unified approach empowers our customers to expertly deploy and manage AI-powered applications... ....As a Staff Site Reliability Engineer (SRE), you will provide technical leadership for... ...infrastructure, databases, and networking.Drive automation to eliminate operational...NetworkFull timeTemporary workLocal areaFlexible hours2 days per week$167.7k - $245.2k
...control.As a Senior Site Reliability Engineer (SRE), you will build, operate, and... ...complex production issues spanning Kubernetes, networking, storage, and application layers• Manage... ...that our worldwide network of doers and experts, and you’ll see that the opportunities to...NetworkFull timeTemporary workLocal areaFlexible hours2 days per week- ...NeoCloud is building an AI-operated GPU cloud. That doesn't mean fewer... ...Monitor GPU cluster health, network status, storage systems, and... ...Collect diagnostic data for L2/SME escalation: logs, DCGM output,... ...ServiceNow/Jira). Perform physical DC tasks: cable installation,...NetworkFull timeLocal areaShift workNight shift
- ...bounded contexts of the NeoCloud SRE platform — the multi-region... ...observes, protects, and operates a GPU rental fleet across self-built... ...probe. Hardware Lifecycle & DC Ops: hardware-lifecycle, dc-... ...optimizer, gpu-efficiency-dashboard, network-stability-dashboard, patching-...NetworkFull timeContract workLocal area
- Job Title: Network SME / Network Architect (Cisco ISE & Zero Trust) | Fourways Consulting | Santa Clara, CA. Recruiting Company: Fourways Consulting. Job Location: Santa Clara, California, USA (Onsite - Face‑to‑Face Final Interview). Job Type: Full-Time / Contract. Application...NetworkFull timeContract workRemote work
- ...optical transceivers in advanced CMOS PDKs. We are seeking an expert in analog/mixed-signal circuit design and architecture.THE PERSON... ...and correction), and analog front-end circuits (pad matching networks, ESD, CTLE, linear amplifier)To define circuit-driven micro-architectures...Network
$151.2k - $227.6k
SummaryImagine what you could do here. At Apple, new ideas have a way of becoming extraordinary products, services, and customer experiences very quickly. Bring passion and dedication to your job and there's no telling what you could accomplish. The people here at Apple...RelocationOverseas$184k - $287.5k
...computing. We are constantly looking for ways to improve our GPU architecture and maintain our leadership. NVIDIA is seeking a... ...tradeoffs, including performance bottlenecks, TCO, Power Delivery Network (PDN), DC Networking, etcStrong communication and interpersonal skills,...NetworkFull timeWork experience placementRemote workNight shift$256k - $414k
...lead the design, scaling, and operations of high-performance networking for GPU-based cloud infrastructure. This role is critical to enabling... ....Work directly with AI platform teams, hardware vendors, and SRE groups to influence technology direction and vendor selection...NetworkFull timeLocal area- ...Practice. HDR is a hydropower and dam safety industry leader. We are expert FERC Part 12 Dam Safety Program practitioners, and we are at... ...our authentic selves to work every day.Our eight Employee Network Groups (Asian Pacific, Black, Hispanic/Latino(a), LGBTQ+, People...NetworkContract workWork at officeFlexible hours
- Prolific is seeking fluent Hindi speakers for its Expert Network in San Jose, California. The role involves completing AI training tasks, judging AI performance, and improving AI models using legal expertise. Candidates must have strong English skills, attention to detail...NetworkRemote jobSelf employmentWork from homeFlexible hours
- ...stalls a $50M training run. Bitdeer is building an AI-operated GPU cloud. Storage is where AI workloads either fly or fall over: a... ...direct GPU-to-storage data paths. Deploy and manage storage networking (NFS over RDMA, NVMe-oF, high-speed storage fabrics) and Nvidia...NetworkRemote jobFull timeLocal area
- NVIDIA Networking is seeking a Senior Networking Architect to advance the next generation of AI-focused networks for accelerating data centers. You will contribute across ASIC design, algorithms, and system-level networking to push innovation and performance. The role...Network
- ...United States. This full-time, on-site role involves supporting critical infrastructure, including installation and maintenance of network equipment. Candidates should have over 3 years of data center operations experience, along with strong cabling skills and the ability...NetworkFull time
$145k - $160k
...and labelling is accurate. The Eurofins network of companies is the global leader in food,... ...STRUCTURE This position reports to the West Coast Director of Operations/General... ...maintain an electro-mechanical system SME – Environmental/Dynamics/Electronics Developing...NetworkFull timeContract workMonday to Friday- ...leading technology company is seeking a Product Manager in Santa Clara, CA to lead the strategy and development of networking infrastructure products, focusing on GPU clusters. The ideal candidate will have over 7 years of experience in network product management and possess...Network
- ...seeking a visionary and hands-on Cloud SRE Architect to lead the design,... ...the end-to-end architecture across CPU, GPU, RDS, storage, networking, serverless, and AI services, ensuring... ...frameworks, and three operational tiers (Edge DC → Regional Controller → Global Hub)....NetworkRemote jobFull timeContract workLocal areaShift work
- ...and project initiatives. The role requires a high school diploma and at least 2 years in PSR or related roles; EPIC experience is preferred. Located within the Stanford Medicine Partners network with travel up to 20%. #J-18808-Ljbffr LE0050 University HealthCare AllianceNetwork
- ...through innovative and sustainable solutions. Job Summary: The Expert Field Service Technician will be responsible for initiating and... ...buildings, with a commitment to providing smart metering, network technologies, and advanced analytics for water, electric, and...NetworkWork experience placementWork at officeLocal areaNight shift
- El Camino Health Medical Network is seeking an experienced Endocrinologist to join our outpatient specialty team. You will evaluate and treat a broad range of endocrine disorders, including diabetes, thyroid, osteoporosis, and pituitary conditions, in a patient-centered...Network
$184k - $287.5k
As a Senior Software Architect in the GPU Networking Architecture team, you will define Software Defined Networking (SDN) architectural solutions and be part of a team of specialists who span across numerous technological fields related to the modern data center, such...NetworkFull time$272k - $431.25k
...NVIDIA's strength is to innovate how we architect and develop our GPU for the changing AI and accelerated workloads. We are... ...delivering multiple high-volume SoC systems (such as CPU, GPU, modem, networking or similar).Proficiency in evaluating power/performance and architectural...NetworkFull timeRemote work- NVIDIA Corporation is seeking a hands-on Manager of Solutions Architecture to lead a team delivering large-scale GPU and AI networking deployments for emerging AI labs and strategic customers. You will recruit, mentor, and ensure high-quality delivery while actively participating...Network
- ...company based in Santa Clara is looking for a skilled Product Manager to drive the development of networking infrastructure products, focusing on high-performance systems for GPU clusters. The ideal candidate should have extensive experience in network product management,...Network
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to GPU DC East-West Network SRE Expert (SME) [Remote]. Be the first to apply!
- guest service support expert San Jose, CA
- technology expert San Jose, CA
- fulfillment expert San Jose, CA
- networking San Jose, CA
- food network test kitchen San Jose, CA
- provider network consultant San Jose, CA
- director of network operations San Jose, CA
- network operations center technician San Jose, CA
- network operations center engineer San Jose, CA
- food network San Jose, CA


