GPU DC East-West Network SRE Expert (SME)
Bitdeer
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.
Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.
Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
To learn more, visit [ About the RoleYou keep the fabric that makes 10K GPUs act like one — and turn IB/RoCE telemetry into the ground truth for our congestion and link-failure predictors.
Bitdeer is building an AI-operated GPU cloud where East-West bandwidth is the difference between a healthy training job and a $50M training run stalled by a bad optic. In this role you operate the InfiniBand and RoCEv2 fabrics that carry NCCL traffic across NeoCloud's US DCs, and you feed the AIOps substrate with the fabric telemetry it needs to catch link degradation, congestion, and topology drift before they land on the pager.
What you'll own- InfiniBand fabrics: fat-tree, rail-optimized, and dragonfly topologies for GPU clusters of 100–10,000 GPUs.
- RoCEv2 networks across Nvidia, Arista, and Cisco platforms for RDMA workloads.
- UFM (Unified Fabric Manager) for IB fabric monitoring, diagnostics, and subnet management.
- IB and RoCE performance monitoring and tuning: adaptive routing, congestion control (DCQCN/ECN), traffic isolation.
- NCCL communication tuning: topology detection, ring/tree algorithm selection, GDR configuration.
- Firmware lifecycle across IB switches and HCAs.
- Fault diagnosis: link flaps, symbol errors, packet drops, routing anomalies, credit stalls.
- Coordination with Nvidia/Mellanox support for escalations, bugs, and RMA.
- Wire IB/RoCE telemetry (ibdiagnet, perfquery, ibstat, PortRcvErrors, PortXmitDiscards, adaptive-routing state) into the platform's collection pipeline.
- Partner with the platform team to define the Link and Straggler predictors: what a "bad optic 30 minutes from failure" looks like in the counters.
- Convert every incident into a labeled example the fault-prediction engine can learn from — and every routine mitigation into a workflow the remediation actuator can run.
- 5+ years in data center networking, with at least 3 years focused on InfiniBand or RoCE fabrics
- Hands-on experience deploying and operating Nvidia/Mellanox InfiniBand switches at scale
- Strong understanding of IB subnet management, partitioning, and QoS
- Experience with RoCEv2 deployment including PFC, ECN, DCQCN configuration
- Proficiency with UFM or equivalent IB fabric management tools
- Knowledge of 400G/800G optics, cabling standards, and structured cabling best practices
- Experience diagnosing IB/RoCE network issues using ibdiagnet, perfquery, ibstat, and similar tools
- Understanding of NCCL and how GPU communication maps to network topology
- Instinct for telemetry-driven ops — you've either built dashboards/alerts on RDMA counters at scale, or you can articulate the feature set a fabric-health model would need.
- Runbook-as-code mindset — the diagnostics you run today should become automation next quarter.
- ...policy. NeoCloud is building an AI-operated GPU cloud spanning 4 US DCs, APAC sites, and... ...— DCI, WAN, internet edge, tenant-facing networks — and you build the automation that lets... ...traffic engineering, and capacity planning for DC networks ~ AIOps aptitude — you've...NetworkFull timeLocal area
$286.2k - $364.4k
...future of work across our Americas West region. As the architect of... ...effectively across a highly networked and matrixed organization to translate... ...Leverage and hold responsible SME functions within WPR to... ...network of doers and experts, and you’ll see that the opportunities...NetworkFull timeTemporary workLocal areaFlexible hours- Job Title: GPU Network Engineer Job Location: Sunnyvale, CA (hybrid)Job Salary: 200k-250k + BenefitsRequirements: Data center, HPC networking... ..., please read on!What You Will Be Doing:Design and deploy east-west network fabrics for GPU-to-GPU, rack-to-rack, and cluster-to-...NetworkLocal areaRelocation
- ...beyond. Together, we advance your career. THE TEAMAMD's Data Center GPU organization is transforming the AI and HPC landscape. Our... ...model architectures (e.g., LLMs, transformer variants, graph neural networks), datatypes, and scaling methodologies to anticipate future...NetworkRemote work
- ...Cloud Sre Architect Bitdeer is seeking a visionary and hands-on Cloud SRE Architect... ...end-to-end architecture across CPU, GPU, RDS, storage, networking, serverless, and AI services, ensuring... ..., and three operational tiers (Edge DC → Regional Controller → Global Hub)....NetworkContract workShift work
- ...organization at the forefront of AI infrastructure is seeking a GPU Network Engineer to help design and operate the high-performance... ...network architectures that deliver low latency, high bandwidth east west traffic across distributed GPU environments. Implement and support...Network
- ...stalls a $50M training run. Bitdeer is building an AI-operated GPU cloud. Storage is where AI workloads either fly or fall over: a... ...direct GPU-to-storage data paths. Deploy and manage storage networking (NFS over RDMA, NVMe-oF, high-speed storage fabrics) and Nvidia...NetworkFull timeLocal area
- .... Together, we advance your career. THE TEAM:AMD's Data Center GPU organization is transforming the AI and HPC landscape. Our mission... ...Familiarity with data center infrastructure including networking, cooling, and platform design Experience in product operations,...NetworkRemote workFlexible hours
- Job Title: Network SME / Network Architect (Cisco ISE & Zero Trust) | Fourways Consulting | Santa Clara, CA. Recruiting Company: Fourways Consulting. Job Location: Santa Clara, California, USA (Onsite - Face‑to‑Face Final Interview). Job Type: Full-Time / Contract. Application...NetworkFull timeContract workRemote work
- ...NeoCloud is building an AI-operated GPU cloud. That doesn't mean fewer... ...Monitor GPU cluster health, network status, storage systems, and... ...Collect diagnostic data for L2/SME escalation: logs, DCGM output,... ...ServiceNow/Jira). Perform physical DC tasks: cable installation,...NetworkFull timeLocal areaShift workNight shift
- ...bounded contexts of the NeoCloud SRE platform — the multi-region... ...observes, protects, and operates a GPU rental fleet across self-built... ...probe. Hardware Lifecycle & DC Ops: hardware-lifecycle, dc-... ...optimizer, gpu-efficiency-dashboard, network-stability-dashboard, patching-...NetworkFull timeContract workLocal area
- ...optical transceivers in advanced CMOS PDKs. We are seeking an expert in analog/mixed-signal circuit design and architecture.THE PERSON... ...and correction), and analog front-end circuits (pad matching networks, ESD, CTLE, linear amplifier)To define circuit-driven micro-architectures...Network
$256k - $414k
...lead the design, scaling, and operations of high-performance networking for GPU-based cloud infrastructure. This role is critical to enabling... ....Work directly with AI platform teams, hardware vendors, and SRE groups to influence technology direction and vendor selection...NetworkFull timeLocal area- ...Practice. HDR is a hydropower and dam safety industry leader. We are expert FERC Part 12 Dam Safety Program practitioners, and we are at... ...our authentic selves to work every day.Our eight Employee Network Groups (Asian Pacific, Black, Hispanic/Latino(a), LGBTQ+, People...NetworkContract workWork at officeFlexible hours
- Prolific is seeking fluent Hindi speakers for its Expert Network in San Jose, California. The role involves completing AI training tasks, judging AI performance, and improving AI models using legal expertise. Candidates must have strong English skills, attention to detail...NetworkRemote jobSelf employmentWork from homeFlexible hours
- ...Support Technician to provide on-site IT support in Santa Clara, CA. You will manage end-user devices, Zebra printers, VIP escalations, network issues, and PC logistics with a hands-on, customer‑focused approach. The role requires a bachelor’s degree or higher and strong...NetworkRemote work
- ...United States. This full-time, on-site role involves supporting critical infrastructure, including installation and maintenance of network equipment. Candidates should have over 3 years of data center operations experience, along with strong cabling skills and the ability...NetworkFull time
- ...leading technology company is seeking a Product Manager in Santa Clara, CA to lead the strategy and development of networking infrastructure products, focusing on GPU clusters. The ideal candidate will have over 7 years of experience in network product management and possess...Network
- ...and project initiatives. The role requires a high school diploma and at least 2 years in PSR or related roles; EPIC experience is preferred. Located within the Stanford Medicine Partners network with travel up to 20%. #J-18808-Ljbffr LE0050 University HealthCare AllianceNetwork
$145k - $160k
...and labelling is accurate. The Eurofins network of companies is the global leader in food,... ...STRUCTURE This position reports to the West Coast Director of Operations/General... ...maintain an electro-mechanical system SME – Environmental/Dynamics/Electronics Developing...NetworkFull timeContract workMonday to Friday- ...processing, collaborating with product, design, and engineering teams. The ideal candidate has CUDA C/C++ expertise, deep knowledge of GPU architectures, and production profiling experience, with 6+ years in GPU programming and performance engineering. #J-18808-Ljbffr...Remote job
- ...A leading tech company seeks a Senior Software Engineer for its Fabric Networking team in Santa Clara, CA. You will design and maintain software enabling GPU communication and participate in architecture definition. Ideal candidates have a B.S/M.S/Ph.D. in a related field...NetworkRemote work
- El Camino Health Medical Network is seeking an experienced Endocrinologist to join our outpatient specialty team. You will evaluate and treat a broad range of endocrine disorders, including diabetes, thyroid, osteoporosis, and pituitary conditions, in a patient-centered...Network
- NVIDIA Corporation is seeking a hands-on Manager of Solutions Architecture to lead a team delivering large-scale GPU and AI networking deployments for emerging AI labs and strategic customers. You will recruit, mentor, and ensure high-quality delivery while actively participating...Network
- ...NVIDIA in Santa Clara, CA is seeking a Senior Software Architect for the GPU Networking Architecture team. You will define SDN architectural solutions spanning data center AI networks, distributed AI workloads, Networking Operating Systems, virtualization and storage....Network
- ...company based in Santa Clara is looking for a skilled Product Manager to drive the development of networking infrastructure products, focusing on high-performance systems for GPU clusters. The ideal candidate should have extensive experience in network product management,...Network
- ...Palo Alto Networks in Santa Clara seeks a visionary Senior Principal Engineer/Architect to serve as the technical authority for global SRE and Platform Engineering initiatives. You will be the primary architect and driver of our AI-driven Autonomous SRE transformation...Network
- ...advance the CUDA driver, collaborating with hardware architects and DL experts to unlock GPU compute performance. You will design, implement, and maintain kernel-mode components, compilers, and networking software, with system-level work spanning OS interfaces and memory...Network
- Advanced Micro Devices is seeking a Senior GPU Inference Performance Engineer to own end-to-end profiling of GPU-accelerated AI inference... .... You will collaborate across teams on multi-server inference networking and Kubernetes-related performance concerns. #J-18808-Ljbffr...Network
- NVIDIA AI is seeking a hands-on Solutions Architect Manager to lead a team of GPU, networking and software architects. You will drive end-to-end deployments in customer data centers, mentor engineers, and shape solution design through direct technical reviews and critical...Network
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to GPU DC East-West Network SRE Expert (SME). Be the first to apply!
- fulfillment expert San Jose, CA
- guest service support expert San Jose, CA
- technology expert San Jose, CA
- network operations center technician San Jose, CA
- network operations center engineer San Jose, CA
- compass health network San Jose, CA
- computer network San Jose, CA
- food network San Jose, CA
- network operations center manager San Jose, CA
- learning network San Jose, CA



