Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

GPU DC East-West Network SRE Expert (SME)

Full-time

Bitdeer

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

To learn more, visit [ About the Role

You keep the fabric that makes 10K GPUs act like one — and turn IB/RoCE telemetry into the ground truth for our congestion and link-failure predictors.

Bitdeer is building an AI-operated GPU cloud where East-West bandwidth is the difference between a healthy training job and a $50M training run stalled by a bad optic. In this role you operate the InfiniBand and RoCEv2 fabrics that carry NCCL traffic across NeoCloud's US DCs, and you feed the AIOps substrate with the fabric telemetry it needs to catch link degradation, congestion, and topology drift before they land on the pager.

What you'll own
  • InfiniBand fabrics: fat-tree, rail-optimized, and dragonfly topologies for GPU clusters of 100–10,000 GPUs.
  • RoCEv2 networks across Nvidia, Arista, and Cisco platforms for RDMA workloads.
  • UFM (Unified Fabric Manager) for IB fabric monitoring, diagnostics, and subnet management.
  • IB and RoCE performance monitoring and tuning: adaptive routing, congestion control (DCQCN/ECN), traffic isolation.
  • NCCL communication tuning: topology detection, ring/tree algorithm selection, GDR configuration.
  • Firmware lifecycle across IB switches and HCAs.
  • Fault diagnosis: link flaps, symbol errors, packet drops, routing anomalies, credit stalls.
  • Coordination with Nvidia/Mellanox support for escalations, bugs, and RMA.
Feed the AIOps substrate
  • Wire IB/RoCE telemetry (ibdiagnet, perfquery, ibstat, PortRcvErrors, PortXmitDiscards, adaptive-routing state) into the platform's collection pipeline.
  • Partner with the platform team to define the Link and Straggler predictors: what a "bad optic 30 minutes from failure" looks like in the counters.
  • Convert every incident into a labeled example the fault-prediction engine can learn from — and every routine mitigation into a workflow the remediation actuator can run.
Job Requirement:
  • 5+ years in data center networking, with at least 3 years focused on InfiniBand or RoCE fabrics
  • Hands-on experience deploying and operating Nvidia/Mellanox InfiniBand switches at scale
  • Strong understanding of IB subnet management, partitioning, and QoS
  • Experience with RoCEv2 deployment including PFC, ECN, DCQCN configuration
  • Proficiency with UFM or equivalent IB fabric management tools
  • Knowledge of 400G/800G optics, cabling standards, and structured cabling best practices
  • Experience diagnosing IB/RoCE network issues using ibdiagnet, perfquery, ibstat, and similar tools
  • Understanding of NCCL and how GPU communication maps to network topology
  • Instinct for telemetry-driven ops — you've either built dashboards/alerts on RDMA counters at scale, or you can articulate the feature set a fabric-health model would need.
  • Runbook-as-code mindset — the diagnostics you run today should become automation next quarter.
-------------------------------------------------------------------- Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.
Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the GPU DC East-West Network SRE Expert (SME) in San Jose, CA vacancy
  •  ...execute. NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all of that...  ...: GPU locality, NVLink domain awareness, network rail affinity. Custom Resource...  ...GitOps workflows (ArgoCD/Flux) ~ Strong SRE background: SLI/SLO frameworks, incident management... 
    Network
    Full time
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    21 days ago
  • $180k - $320k

     ...policy. NeoCloud is building an AI-operated GPU cloud spanning 4 US DCs, APAC sites, and...  ...— DCI, WAN, internet edge, tenant-facing networks — and you build the automation that lets...  ...traffic engineering, and capacity planning for DC networks ~ AIOps aptitude — you've... 
    Network
    Remote job
    Full time
    Local area

    Bitdeer

    San Jose, CA
    21 days ago
  • $180k - $320k

     ...stalls a $50M training run. Bitdeer is building an AI-operated GPU cloud. Storage is where AI workloads either fly or fall over: a...  ...direct GPU-to-storage data paths. Deploy and manage storage networking (NFS over RDMA, NVMe-oF, high-speed storage fabrics) and Nvidia... 
    Network
    Remote job
    Full time
    Local area

    Bitdeer

    San Jose, CA
    21 days ago
  •  ...your career.THE TEAM:AMD's Data Center GPU organization is transforming the...  ...for its Tier 1 CSP Business Group (DC GPU CBG), focused on driving our business...  ..., media, analysts, technical experts and senior executivesPossess a network of industry relationships with potential... 
    Network
    Work experience placement

    AMD

    Santa Clara, CA
    3 days ago
  • Job Title: GPU Network Engineer Job Location: Sunnyvale, CA (hybrid)Job Salary: 200k-250k + BenefitsRequirements: Data center, HPC networking...  ..., please read on!What You Will Be Doing:Design and deploy east-west network fabrics for GPU-to-GPU, rack-to-rack, and cluster-to-... 
    Network
    Local area
    Relocation

    CyberCoders

    San Jose, CA
    3 days ago
  •  ...Together, we advance your career. THE TEAM AMD's Data Center GPU organization is transforming the AI and HPC landscape. Our...  ...environments Familiarity with data center infrastructure including networking, cooling, and platform design Experience in product... 
    Network
    Remote work
    Flexible hours

    Advanced Micro Devices

    Santa Clara, CA
    3 days ago
  •  ...bounded contexts of the NeoCloud SRE platform — the multi-region...  ...observes, protects, and operates a GPU rental fleet across self-built...  ...probe. Hardware Lifecycle & DC Ops: hardware-lifecycle, dc-...  ...optimizer, gpu-efficiency-dashboard, network-stability-dashboard, patching-... 
    Network
    Full time
    Contract work
    Local area

    Bitdeer

    San Jose, CA
    5 days ago
  •  ...NeoCloud is building an AI-operated GPU cloud. That doesn't mean fewer...  ...Monitor GPU cluster health, network status, storage systems, and...  ...Collect diagnostic data for L2/SME escalation: logs, DCGM output,...  ...ServiceNow/Jira). Perform physical DC tasks: cable installation,... 
    Network
    Full time
    Local area
    Shift work
    Night shift

    Bitdeer

    San Jose, CA
    5 days ago
  •  ...seeking a visionary and hands-on Cloud SRE Architect to lead the design,...  ...the end-to-end architecture across CPU, GPU, RDS, storage, networking, serverless, and AI services, ensuring...  ...frameworks, and three operational tiers (Edge DC → Regional Controller → Global Hub).... 
    Network
    Full time
    Contract work
    Local area
    Shift work

    Bitdeer

    San Jose, CA
    5 days ago
  • $98.04k - $227.82k

     ...one of the worlds top tax firms. Enjoy a collaborative, future-forward culture that empowers your success. Work with KPMGs extensive network of specialists; enjoy access to our Ignition Centers, where deep industry knowledge merges with cutting-edge technologies to create... 
    Network
    Local area

    KPMG

    Santa Clara, CA
    10 hours ago
  • $152k - $241.5k

     ...DGX Cloud builds and operates large-scale GPU infrastructure for AI workloads. We are looking for Software Engineers with SRE or Production Engineering experience who have...  ...servers, DPUs, GPU systems, CPU systems, networking, Linux, and Kubernetes; turn recurring issues... 
    Network
    Permanent employment
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • Prolific is seeking fluent Hindi speakers for its Expert Network in San Jose, California. The role involves completing AI training tasks, judging AI performance, and improving AI models using legal expertise. Candidates must have strong English skills, attention to detail... 
    Network
    Remote job
    Self employment
    Work from home
    Flexible hours

    Prolific

    San Jose, CA
    1 day ago
  •  ...gWMS Functional Expert San Jose, CA onsite 1) WH Automation 2) gWMS Feature requests Understanding logistics business operations...  ...Operations, Systems Engineering, Project Management, IT Network/Infrastructure teams, and third-party logistics (3PL) providers... 
    Network

    Shrive Technologies LLC

    San Jose, CA
    2 days ago
  •  ...Practice. HDR is a hydropower and dam safety industry leader. We are expert FERC Part 12 Dam Safety Program practitioners, and we are at...  ...our authentic selves to work every day.Our eight Employee Network Groups (Asian Pacific, Black, Hispanic/Latino(a), LGBTQ+, People... 
    Network
    Contract work
    Work at office
    Flexible hours

    HDR

    San Jose, CA
    17 hours ago
  •  ...for an  AI Devops Infrastructure Engineer/GPU Infrastructure Engineer. Job Title:...  ...foundation in infrastructure engineering, DevOps/SRE, platform engineering, or similar...  ...containerized environments, including compute, networking, storage, and accelerator scheduling.... 
    Network
    Contract work

    Maxonic

    San Jose, CA
    9 days ago
  •  ...Corporation in San Jose, CA is seeking an Onsite AIOps Observability/SRE Lead to build and guide a reliability engineering team across...  ...will partner with IT, security, and cloud teams to modernize networks, enable hybrid cloud, and ensure secure, reliable connectivity.... 
    Network

    Jobleads-US

    San Jose, CA
    1 day ago
  • $208k - $333.5k

     ...Site Reliability Engineering (SRE) at NVIDIA is an engineering field focused on designing...  ...expertise across distributed systems, networking, Kubernetes, public cloud, observability,...  ...Performance Computing, and Visualization. The GPU, our invention, serves as the visual... 
    Network
    Full time

    NVIDIA

    Santa Clara, CA
    2 days ago
  •  ...Palo Alto Networks is seeking a Senior Principal Engineer/Architect to serve as the technical authority for global SRE and Platform Engineering initiatives across the US and India. You will architect AI-driven, self-healing platform capabilities and partner with product... 
    Network

    Jobleads-US

    Santa Clara, CA
    4 days ago
  •  ...Job Description Job Description Java SRE Engineer Onsite San Francisco Bay Area Infrastructure Engineer (2 Positions)...  ...migration projects Strong understanding of distributed systems and networking Preferred Qualifications: Experience with Akamai CDN and... 
    Network

    Eitacies Inc

    Santa Clara, CA
    a month ago
  • $145k - $160k

     ...and labelling is accurate. The Eurofins network of companies is the global leader in food,...  ...STRUCTURE This position reports to the West Coast Director of Operations/General...  ...maintain an electro-mechanical system SME – Environmental/Dynamics/Electronics Developing... 
    Network
    Full time
    Contract work
    Monday to Friday

    Eurofins USA Consumer Product Testing

    Santa Clara, CA
    17 days ago
  •  ...Head of GPU Cloud About the Company Building software foundation for next-generation accelerated computing platform for large...  ...experience with GPU memory behavior and high-bandwidth networking is essential. The role also requires a leader with strong technical... 
    Network

    Confidential

    San Jose, CA
    17 hours ago
  • $124k - $195.5k

     ...intelligence. Our data-center platforms bring together GPUs, CPUs, networking, systems, and software to solve some of the world’s most...  ...We are looking for a System Software Engineer to join NVIDIA’s GPU Performance and Power Management Software team. You will help design... 
    Network
    Full time
    Internship

    Nvidia

    Santa Clara, CA
    2 days ago
  • Network Architect Location - Santa Clara. This is a hands-on architecture...  ...and scalable interconnects for GPU-accelerated data centers and...  ...interconnects, intra and inter DC routing, and dark fiber deployments...  ...in infrastructure automation. SME in networking technologies:... 
    Network

    Tranzeal

    Santa Clara, CA
    4 days ago
  • $224k - $356.5k

    We are now looking for a GPU System Performance Architect:The NVIDIA Architecture group is looking for extraordinary computer architects...  ...it drives, and the objects it must avoid. In healthcare, neural networks trained with millions of medical images can find clues in MRIs... 
    Network
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...Job Description Job Description GPU Programmer / Software Engineer Job Type: Contractor Location: Remote Job Overview We are seeking experienced GPU Programmers / Software Engineers to design and optimize GPU-based tasks for AI and LLM applications. You... 
    Remote job
    For contractors

    YO AI Labs

    San Jose, CA
    15 days ago
  • $184k - $287.5k

     ...of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self‑driving cars...  ...with next-generation datacenter builds for NVIDIA GPUs, CPUs, and networking hardware. Engage early with HW/FW/SW/platform internal and... 
    Network

    NVIDIA Gruppe

    Santa Clara, CA
    1 day ago
  •  ...Senior Systems Software Engineer – GPU Performance at Scale We are looking for a dedicated engineer for the Senior Systems Software...  ...next‑generation datacenter builds for NVIDIA GPUs, CPUs, and networking hardware. Engage early with HW/FW/SW/platform internal and... 
    Network

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $187k - $270.7k

    ## AIOPs Observability/SRE LeadApply: San Jose, California, United States: Full time: Posted Yesterday: R03209# **Job Details:**###...  ...teams, security teams, and external service providers to modernize network infrastructure, enable hybrid cloud connectivity, and ensure a... 
    Network
    Full time
    Work at office
    Flexible hours

    Altera Corporation

    San Jose, CA
    1 day ago
  • $250k - $330k

     ...Network Infrastructure Architect San Jose, California, United States The era of pervasive AI has arrived. In this era, organizations...  ...2, Ethernet-based AI fabrics, 400G/800G interconnects, and GPU-to-GPU east-west traffic patterns ~ Experience with lossless Ethernet... 
    Network
    Local area

    SambaNova Systems

    San Jose, CA
    1 day ago
  • $27 - $41 per hour

     ...dedicated year-round TurboTax Local Service Experts in one of our new TurboTax locations...  ...grass-roots marketing initiatives, referral networks, and community partnerships. Serve as the...  ...: $27.00 - $41.00  Washington, DC: $26.00 - $39.00  This position will be... 
    Network
    Work at office
    Local area
    Monday to Friday

    Intuit

    Los Gatos, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to GPU DC East-West Network SRE Expert (SME). Be the first to apply!