Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

GPU DC East-West Network SRE Expert (SME) [Remote]

Full-time

Bitdeer Technologies Group

San Jose, CA
  • Remote job

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

To learn more, visit  (

About the Role  

You keep the fabric that makes 10K GPUs act like one — and turn IB/RoCE telemetry into the ground truth for our congestion and link-failure predictors.

Bitdeer is building an AI-operated GPU cloud where East-West bandwidth is the difference between a healthy training job and a $50M training run stalled by a bad optic. In this role you operate the InfiniBand and RoCEv2 fabrics that carry NCCL traffic across NeoCloud's US DCs, and you feed the AIOps substrate with the fabric telemetry it needs to catch link degradation, congestion, and topology drift before they land on the pager.

What you'll own

  • InfiniBand fabrics: fat-tree, rail-optimized, and dragonfly topologies for GPU clusters of 100–10,000 GPUs.
  • RoCEv2 networks across Nvidia, Arista, and Cisco platforms for RDMA workloads.
  • UFM (Unified Fabric Manager) for IB fabric monitoring, diagnostics, and subnet management.
  • IB and RoCE performance monitoring and tuning: adaptive routing, congestion control (DCQCN/ECN), traffic isolation.
  • NCCL communication tuning: topology detection, ring/tree algorithm selection, GDR configuration.
  • Firmware lifecycle across IB switches and HCAs.
  • Fault diagnosis: link flaps, symbol errors, packet drops, routing anomalies, credit stalls.
  • Coordination with Nvidia/Mellanox support for escalations, bugs, and RMA.

Feed the AIOps substrate

  • Wire IB/RoCE telemetry (ibdiagnet, perfquery, ibstat, PortRcvErrors, PortXmitDiscards, adaptive-routing state) into the platform's collection pipeline.
  • Partner with the platform team to define the Link and Straggler predictors: what a "bad optic 30 minutes from failure" looks like in the counters.
  • Convert every incident into a labeled example the fault-prediction engine can learn from — and every routine mitigation into a workflow the remediation actuator can run.

Job Requirement:

  • 5+ years in data center networking, with at least 3 years focused on InfiniBand or RoCE fabrics
  • Hands-on experience deploying and operating Nvidia/Mellanox InfiniBand switches at scale
  • Strong understanding of IB subnet management, partitioning, and QoS
  • Experience with RoCEv2 deployment including PFC, ECN, DCQCN configuration
  • Proficiency with UFM or equivalent IB fabric management tools
  • Knowledge of 400G/800G optics, cabling standards, and structured cabling best practices
  • Experience diagnosing IB/RoCE network issues using ibdiagnet, perfquery, ibstat, and similar tools
  • Understanding of NCCL and how GPU communication maps to network topology
  • Instinct for telemetry-driven ops — you've either built dashboards/alerts on RDMA counters at scale, or you can articulate the feature set a fabric-health model would need.
  • Runbook-as-code mindset — the diagnostics you run today should become automation next quarter.

--------------------------------------------------------------------

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Vacancy posted 14 days ago
Similar jobs that could be interesting for youBased on the GPU DC East-West Network SRE Expert (SME) [Remote] in San Jose, CA vacancy
  •  ...failure predictors. Bitdeer is building an AI-operated GPU cloud where East-West bandwidth is the difference between a healthy training job...  ...topologies for GPU clusters of 100–10,000 GPUs. RoCEv2 networks across Nvidia, Arista, and Cisco platforms for RDMA workloads... 
    Network
    Remote job
    Full time
    Local area

    Bitdeer

    San Jose, CA
    1 day ago
  •  ...policy. NeoCloud is building an AI-operated GPU cloud spanning 4 US DCs, APAC sites, and...  ...— DCI, WAN, internet edge, tenant-facing networks — and you build the automation that lets...  ...traffic engineering, and capacity planning for DC networks ~ AIOps aptitude — you've... 
    Network
    Remote job
    Full time
    Local area

    Bitdeer

    San Jose, CA
    2 days ago
  • $286.2k - $364.4k

     ...future of work across our Americas West region. As the architect of...  ...effectively across a highly networked and matrixed organization to translate...  ...Leverage and hold responsible SME functions within WPR to...  ...network of doers and experts, and you’ll see that the opportunities... 
    Network
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    3 days ago
  •  ...beyond. Together, we advance your career. THE TEAMAMD's Data Center GPU organization is transforming the AI and HPC landscape. Our...  ...model architectures (e.g., LLMs, transformer variants, graph neural networks), datatypes, and scaling methodologies to anticipate future... 
    Network
    Remote work

    AMD

    San Jose, CA
    3 days ago
  •  ...organization at the forefront of AI infrastructure is seeking a GPU Network Engineer to help design and operate the high-performance...  ...network architectures that deliver low latency, high bandwidth east west traffic across distributed GPU environments. Implement and support... 
    Network

    Blue Signal Search

    Santa Clara, CA
    3 days ago
  •  .... Together, we advance your career. THE TEAM:AMD's Data Center GPU organization is transforming the AI and HPC landscape. Our mission...  ...Familiarity with data center infrastructure including networking, cooling, and platform design Experience in product operations,... 
    Network
    Remote work
    Flexible hours

    AMD

    Santa Clara, CA
    16 hours ago
  • $186.9k - $267.7k

     ...unified approach empowers our customers to expertly deploy and manage AI-powered applications...  ....As a Staff Site Reliability Engineer (SRE), you will provide technical leadership for...  ...infrastructure, databases, and networking.Drive automation to eliminate operational... 
    Network
    Full time
    Temporary work
    Local area
    Flexible hours
    2 days per week

    CISCO Systems

    Milpitas, CA
    1 day ago
  • $167.7k - $245.2k

     ...control.As a Senior Site Reliability Engineer (SRE), you will build, operate, and...  ...complex production issues spanning Kubernetes, networking, storage, and application layers• Manage...  ...that our worldwide network of doers and experts, and you’ll see that the opportunities to... 
    Network
    Full time
    Temporary work
    Local area
    Flexible hours
    2 days per week

    CISCO Systems

    Milpitas, CA
    16 hours ago
  •  ...NeoCloud is building an AI-operated GPU cloud. That doesn't mean fewer...  ...Monitor GPU cluster health, network status, storage systems, and...  ...Collect diagnostic data for L2/SME escalation: logs, DCGM output,...  ...ServiceNow/Jira). Perform physical DC tasks: cable installation,... 
    Network
    Full time
    Local area
    Shift work
    Night shift

    Bitdeer

    San Jose, CA
    1 day ago
  •  ...bounded contexts of the NeoCloud SRE platform — the multi-region...  ...observes, protects, and operates a GPU rental fleet across self-built...  ...probe. Hardware Lifecycle & DC Ops: hardware-lifecycle, dc-...  ...optimizer, gpu-efficiency-dashboard, network-stability-dashboard, patching-... 
    Network
    Full time
    Contract work
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    a month ago
  • Job Title: Network SME / Network Architect (Cisco ISE & Zero Trust) | Fourways Consulting | Santa Clara, CA. Recruiting Company: Fourways Consulting. Job Location: Santa Clara, California, USA (Onsite - Face‑to‑Face Final Interview). Job Type: Full-Time / Contract. Application... 
    Network
    Full time
    Contract work
    Remote work

    Fourways Consulting

    Santa Clara, CA
    1 day ago
  •  ...optical transceivers in advanced CMOS PDKs. We are seeking an expert in analog/mixed-signal circuit design and architecture.THE PERSON...  ...and correction), and analog front-end circuits (pad matching networks, ESD, CTLE, linear amplifier)To define circuit-driven micro-architectures... 
    Network

    AMD

    San Jose, CA
    2 days ago
  • $151.2k - $227.6k

    SummaryImagine what you could do here. At Apple, new ideas have a way of becoming extraordinary products, services, and customer experiences very quickly. Bring passion and dedication to your job and there's no telling what you could accomplish. The people here at Apple...
    Relocation
    Overseas

    Apple

    Cupertino, CA
    1 day ago
  • $184k - $287.5k

     ...computing. We are constantly looking for ways to improve our GPU architecture and maintain our leadership. NVIDIA is seeking a...  ...tradeoffs, including performance bottlenecks, TCO, Power Delivery Network (PDN), DC Networking, etcStrong communication and interpersonal skills,... 
    Network
    Full time
    Work experience placement
    Remote work
    Night shift

    Nvidia

    Santa Clara, CA
    1 day ago
  • $256k - $414k

     ...lead the design, scaling, and operations of high-performance networking for GPU-based cloud infrastructure. This role is critical to enabling...  ....Work directly with AI platform teams, hardware vendors, and SRE groups to influence technology direction and vendor selection... 
    Network
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...Practice. HDR is a hydropower and dam safety industry leader. We are expert FERC Part 12 Dam Safety Program practitioners, and we are at...  ...our authentic selves to work every day.Our eight Employee Network Groups (Asian Pacific, Black, Hispanic/Latino(a), LGBTQ+, People... 
    Network
    Contract work
    Work at office
    Flexible hours

    HDR

    San Jose, CA
    1 day ago
  • Prolific is seeking fluent Hindi speakers for its Expert Network in San Jose, California. The role involves completing AI training tasks, judging AI performance, and improving AI models using legal expertise. Candidates must have strong English skills, attention to detail... 
    Network
    Remote job
    Self employment
    Work from home
    Flexible hours

    Prolific

    San Jose, CA
    2 days ago
  •  ...stalls a $50M training run. Bitdeer is building an AI-operated GPU cloud. Storage is where AI workloads either fly or fall over: a...  ...direct GPU-to-storage data paths. Deploy and manage storage networking (NFS over RDMA, NVMe-oF, high-speed storage fabrics) and Nvidia... 
    Network
    Remote job
    Full time
    Local area

    Bitdeer

    San Jose, CA
    1 day ago
  • NVIDIA Networking is seeking a Senior Networking Architect to advance the next generation of AI-focused networks for accelerating data centers. You will contribute across ASIC design, algorithms, and system-level networking to push innovation and performance. The role... 
    Network

    NVIDIA

    Santa Clara, CA
    2 days ago
  •  ...United States. This full-time, on-site role involves supporting critical infrastructure, including installation and maintenance of network equipment. Candidates should have over 3 years of data center operations experience, along with strong cabling skills and the ability... 
    Network
    Full time

    Rebootmonkey

    Milpitas, CA
    2 days ago
  • $145k - $160k

     ...and labelling is accurate. The Eurofins network of companies is the global leader in food,...  ...STRUCTURE This position reports to the West Coast Director of Operations/General...  ...maintain an electro-mechanical system SME – Environmental/Dynamics/Electronics Developing... 
    Network
    Full time
    Contract work
    Monday to Friday

    Eurofins USA Consumer Product Testing

    Santa Clara, CA
    26 days ago
  •  ...leading technology company is seeking a Product Manager in Santa Clara, CA to lead the strategy and development of networking infrastructure products, focusing on GPU clusters. The ideal candidate will have over 7 years of experience in network product management and possess... 
    Network

    Advanced Micro Devices, Inc.

    Santa Clara, CA
    4 days ago
  •  ...seeking a visionary and hands-on Cloud SRE Architect to lead the design,...  ...the end-to-end architecture across CPU, GPU, RDS, storage, networking, serverless, and AI services, ensuring...  ...frameworks, and three operational tiers (Edge DC → Regional Controller → Global Hub).... 
    Network
    Remote job
    Full time
    Contract work
    Local area
    Shift work

    Bitdeer

    San Jose, CA
    5 days ago
  •  ...and project initiatives. The role requires a high school diploma and at least 2 years in PSR or related roles; EPIC experience is preferred. Located within the Stanford Medicine Partners network with travel up to 20%. #J-18808-Ljbffr LE0050 University HealthCare Alliance
    Network

    LE0050 University HealthCare Alliance

    Los Gatos, CA
    6 hours ago
  •  ...through innovative and sustainable solutions. Job Summary: The Expert Field Service Technician will be responsible for initiating and...  ...buildings, with a commitment to providing smart metering, network technologies, and advanced analytics for water, electric, and... 
    Network
    Work experience placement
    Work at office
    Local area
    Night shift

    Xylem

    Milpitas, CA
    16 hours ago
  • El Camino Health Medical Network is seeking an experienced Endocrinologist to join our outpatient specialty team. You will evaluate and treat a broad range of endocrine disorders, including diabetes, thyroid, osteoporosis, and pituitary conditions, in a patient-centered... 
    Network

    El Camino Health Medical Network

    Cupertino, CA
    4 days ago
  • $184k - $287.5k

    As a Senior Software Architect in the GPU Networking Architecture team, you will define Software Defined Networking (SDN) architectural solutions and be part of a team of specialists who span across numerous technological fields related to the modern data center, such... 
    Network
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $272k - $431.25k

     ...NVIDIA's strength is to innovate how we architect and develop our GPU for the changing AI and accelerated workloads. We are...  ...delivering multiple high-volume SoC systems (such as CPU, GPU, modem, networking or similar).Proficiency in evaluating power/performance and architectural... 
    Network
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • NVIDIA Corporation is seeking a hands-on Manager of Solutions Architecture to lead a team delivering large-scale GPU and AI networking deployments for emerging AI labs and strategic customers. You will recruit, mentor, and ensure high-quality delivery while actively participating... 
    Network

    NVIDIA

    Santa Clara, CA
    2 days ago
  •  ...company based in Santa Clara is looking for a skilled Product Manager to drive the development of networking infrastructure products, focusing on high-performance systems for GPU clusters. The ideal candidate should have extensive experience in network product management,... 
    Network

    AMD

    Santa Clara, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to GPU DC East-West Network SRE Expert (SME) [Remote]. Be the first to apply!