GPU DC East-West Network SRE Expert (SME)
Bitdeer
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.
Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.
Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
To learn more, visit [ About the RoleYou keep the fabric that makes 10K GPUs act like one — and turn IB/RoCE telemetry into the ground truth for our congestion and link-failure predictors.
Bitdeer is building an AI-operated GPU cloud where East-West bandwidth is the difference between a healthy training job and a $50M training run stalled by a bad optic. In this role you operate the InfiniBand and RoCEv2 fabrics that carry NCCL traffic across NeoCloud's US DCs, and you feed the AIOps substrate with the fabric telemetry it needs to catch link degradation, congestion, and topology drift before they land on the pager.
What you'll own- InfiniBand fabrics: fat-tree, rail-optimized, and dragonfly topologies for GPU clusters of 100–10,000 GPUs.
- RoCEv2 networks across Nvidia, Arista, and Cisco platforms for RDMA workloads.
- UFM (Unified Fabric Manager) for IB fabric monitoring, diagnostics, and subnet management.
- IB and RoCE performance monitoring and tuning: adaptive routing, congestion control (DCQCN/ECN), traffic isolation.
- NCCL communication tuning: topology detection, ring/tree algorithm selection, GDR configuration.
- Firmware lifecycle across IB switches and HCAs.
- Fault diagnosis: link flaps, symbol errors, packet drops, routing anomalies, credit stalls.
- Coordination with Nvidia/Mellanox support for escalations, bugs, and RMA.
- Wire IB/RoCE telemetry (ibdiagnet, perfquery, ibstat, PortRcvErrors, PortXmitDiscards, adaptive-routing state) into the platform's collection pipeline.
- Partner with the platform team to define the Link and Straggler predictors: what a "bad optic 30 minutes from failure" looks like in the counters.
- Convert every incident into a labeled example the fault-prediction engine can learn from — and every routine mitigation into a workflow the remediation actuator can run.
- 5+ years in data center networking, with at least 3 years focused on InfiniBand or RoCE fabrics
- Hands-on experience deploying and operating Nvidia/Mellanox InfiniBand switches at scale
- Strong understanding of IB subnet management, partitioning, and QoS
- Experience with RoCEv2 deployment including PFC, ECN, DCQCN configuration
- Proficiency with UFM or equivalent IB fabric management tools
- Knowledge of 400G/800G optics, cabling standards, and structured cabling best practices
- Experience diagnosing IB/RoCE network issues using ibdiagnet, perfquery, ibstat, and similar tools
- Understanding of NCCL and how GPU communication maps to network topology
- Instinct for telemetry-driven ops — you've either built dashboards/alerts on RDMA counters at scale, or you can articulate the feature set a fabric-health model would need.
- Runbook-as-code mindset — the diagnostics you run today should become automation next quarter.
- ...execute. NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all of that... ...: GPU locality, NVLink domain awareness, network rail affinity. Custom Resource... ...GitOps workflows (ArgoCD/Flux) ~ Strong SRE background: SLI/SLO frameworks, incident management...NetworkFull timeLocal area
$180k - $320k
...policy. NeoCloud is building an AI-operated GPU cloud spanning 4 US DCs, APAC sites, and... ...— DCI, WAN, internet edge, tenant-facing networks — and you build the automation that lets... ...traffic engineering, and capacity planning for DC networks ~ AIOps aptitude — you've...NetworkRemote jobFull timeLocal area$180k - $320k
...stalls a $50M training run. Bitdeer is building an AI-operated GPU cloud. Storage is where AI workloads either fly or fall over: a... ...direct GPU-to-storage data paths. Deploy and manage storage networking (NFS over RDMA, NVMe-oF, high-speed storage fabrics) and Nvidia...NetworkRemote jobFull timeLocal area- ...your career.THE TEAM:AMD's Data Center GPU organization is transforming the... ...for its Tier 1 CSP Business Group (DC GPU CBG), focused on driving our business... ..., media, analysts, technical experts and senior executivesPossess a network of industry relationships with potential...NetworkWork experience placement
- Job Title: GPU Network Engineer Job Location: Sunnyvale, CA (hybrid)Job Salary: 200k-250k + BenefitsRequirements: Data center, HPC networking... ..., please read on!What You Will Be Doing:Design and deploy east-west network fabrics for GPU-to-GPU, rack-to-rack, and cluster-to-...NetworkLocal areaRelocation
- ...Together, we advance your career. THE TEAM AMD's Data Center GPU organization is transforming the AI and HPC landscape. Our... ...environments Familiarity with data center infrastructure including networking, cooling, and platform design Experience in product...NetworkRemote workFlexible hours
- ...bounded contexts of the NeoCloud SRE platform — the multi-region... ...observes, protects, and operates a GPU rental fleet across self-built... ...probe. Hardware Lifecycle & DC Ops: hardware-lifecycle, dc-... ...optimizer, gpu-efficiency-dashboard, network-stability-dashboard, patching-...NetworkFull timeContract workLocal area
- ...NeoCloud is building an AI-operated GPU cloud. That doesn't mean fewer... ...Monitor GPU cluster health, network status, storage systems, and... ...Collect diagnostic data for L2/SME escalation: logs, DCGM output,... ...ServiceNow/Jira). Perform physical DC tasks: cable installation,...NetworkFull timeLocal areaShift workNight shift
- ...seeking a visionary and hands-on Cloud SRE Architect to lead the design,... ...the end-to-end architecture across CPU, GPU, RDS, storage, networking, serverless, and AI services, ensuring... ...frameworks, and three operational tiers (Edge DC → Regional Controller → Global Hub)....NetworkFull timeContract workLocal areaShift work
$98.04k - $227.82k
...one of the worlds top tax firms. Enjoy a collaborative, future-forward culture that empowers your success. Work with KPMGs extensive network of specialists; enjoy access to our Ignition Centers, where deep industry knowledge merges with cutting-edge technologies to create...NetworkLocal area$152k - $241.5k
...DGX Cloud builds and operates large-scale GPU infrastructure for AI workloads. We are looking for Software Engineers with SRE or Production Engineering experience who have... ...servers, DPUs, GPU systems, CPU systems, networking, Linux, and Kubernetes; turn recurring issues...NetworkPermanent employmentFull time- Prolific is seeking fluent Hindi speakers for its Expert Network in San Jose, California. The role involves completing AI training tasks, judging AI performance, and improving AI models using legal expertise. Candidates must have strong English skills, attention to detail...NetworkRemote jobSelf employmentWork from homeFlexible hours
- ...gWMS Functional Expert San Jose, CA onsite 1) WH Automation 2) gWMS Feature requests Understanding logistics business operations... ...Operations, Systems Engineering, Project Management, IT Network/Infrastructure teams, and third-party logistics (3PL) providers...Network
- ...Practice. HDR is a hydropower and dam safety industry leader. We are expert FERC Part 12 Dam Safety Program practitioners, and we are at... ...our authentic selves to work every day.Our eight Employee Network Groups (Asian Pacific, Black, Hispanic/Latino(a), LGBTQ+, People...NetworkContract workWork at officeFlexible hours
- ...for an AI Devops Infrastructure Engineer/GPU Infrastructure Engineer. Job Title:... ...foundation in infrastructure engineering, DevOps/SRE, platform engineering, or similar... ...containerized environments, including compute, networking, storage, and accelerator scheduling....NetworkContract work
- ...Corporation in San Jose, CA is seeking an Onsite AIOps Observability/SRE Lead to build and guide a reliability engineering team across... ...will partner with IT, security, and cloud teams to modernize networks, enable hybrid cloud, and ensure secure, reliable connectivity....Network
$208k - $333.5k
...Site Reliability Engineering (SRE) at NVIDIA is an engineering field focused on designing... ...expertise across distributed systems, networking, Kubernetes, public cloud, observability,... ...Performance Computing, and Visualization. The GPU, our invention, serves as the visual...NetworkFull time- ...Palo Alto Networks is seeking a Senior Principal Engineer/Architect to serve as the technical authority for global SRE and Platform Engineering initiatives across the US and India. You will architect AI-driven, self-healing platform capabilities and partner with product...Network
- ...Job Description Job Description Java SRE Engineer Onsite San Francisco Bay Area Infrastructure Engineer (2 Positions)... ...migration projects Strong understanding of distributed systems and networking Preferred Qualifications: Experience with Akamai CDN and...Network
$145k - $160k
...and labelling is accurate. The Eurofins network of companies is the global leader in food,... ...STRUCTURE This position reports to the West Coast Director of Operations/General... ...maintain an electro-mechanical system SME – Environmental/Dynamics/Electronics Developing...NetworkFull timeContract workMonday to Friday- ...Head of GPU Cloud About the Company Building software foundation for next-generation accelerated computing platform for large... ...experience with GPU memory behavior and high-bandwidth networking is essential. The role also requires a leader with strong technical...Network
$124k - $195.5k
...intelligence. Our data-center platforms bring together GPUs, CPUs, networking, systems, and software to solve some of the world’s most... ...We are looking for a System Software Engineer to join NVIDIA’s GPU Performance and Power Management Software team. You will help design...NetworkFull timeInternship- Network Architect Location - Santa Clara. This is a hands-on architecture... ...and scalable interconnects for GPU-accelerated data centers and... ...interconnects, intra and inter DC routing, and dark fiber deployments... ...in infrastructure automation. SME in networking technologies:...Network
$224k - $356.5k
We are now looking for a GPU System Performance Architect:The NVIDIA Architecture group is looking for extraordinary computer architects... ...it drives, and the objects it must avoid. In healthcare, neural networks trained with millions of medical images can find clues in MRIs...NetworkFull timeWork experience placement- ...Job Description Job Description GPU Programmer / Software Engineer Job Type: Contractor Location: Remote Job Overview We are seeking experienced GPU Programmers / Software Engineers to design and optimize GPU-based tasks for AI and LLM applications. You...Remote jobFor contractors
$184k - $287.5k
...of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self‑driving cars... ...with next-generation datacenter builds for NVIDIA GPUs, CPUs, and networking hardware. Engage early with HW/FW/SW/platform internal and...Network- ...Senior Systems Software Engineer – GPU Performance at Scale We are looking for a dedicated engineer for the Senior Systems Software... ...next‑generation datacenter builds for NVIDIA GPUs, CPUs, and networking hardware. Engage early with HW/FW/SW/platform internal and...Network
$187k - $270.7k
## AIOPs Observability/SRE LeadApply: San Jose, California, United States: Full time: Posted Yesterday: R03209# **Job Details:**###... ...teams, security teams, and external service providers to modernize network infrastructure, enable hybrid cloud connectivity, and ensure a...NetworkFull timeWork at officeFlexible hours$250k - $330k
...Network Infrastructure Architect San Jose, California, United States The era of pervasive AI has arrived. In this era, organizations... ...2, Ethernet-based AI fabrics, 400G/800G interconnects, and GPU-to-GPU east-west traffic patterns ~ Experience with lossless Ethernet...NetworkLocal area$27 - $41 per hour
...dedicated year-round TurboTax Local Service Experts in one of our new TurboTax locations... ...grass-roots marketing initiatives, referral networks, and community partnerships. Serve as the... ...: $27.00 - $41.00 Washington, DC: $26.00 - $39.00 This position will be...NetworkWork at officeLocal areaMonday to Friday
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to GPU DC East-West Network SRE Expert (SME). Be the first to apply!
- guest service support expert San Jose, CA
- fulfillment expert San Jose, CA
- technology expert San Jose, CA
- rn network San Jose, CA
- director of network operations San Jose, CA
- networking San Jose, CA
- network cabling San Jose, CA
- compass health network San Jose, CA
- staffing network San Jose, CA
- data network cabling San Jose, CA





