Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

GPU DC East-West Network SRE Expert (SME) [Remote]

Full-time

Bitdeer

San Jose, CA
  • Remote job

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

To learn more, visit [ About the Role

You keep the fabric that makes 10K GPUs act like one — and turn IB/RoCE telemetry into the ground truth for our congestion and link-failure predictors.

Bitdeer is building an AI-operated GPU cloud where East-West bandwidth is the difference between a healthy training job and a $50M training run stalled by a bad optic. In this role you operate the InfiniBand and RoCEv2 fabrics that carry NCCL traffic across NeoCloud's US DCs, and you feed the AIOps substrate with the fabric telemetry it needs to catch link degradation, congestion, and topology drift before they land on the pager.

What you'll own
  • InfiniBand fabrics: fat-tree, rail-optimized, and dragonfly topologies for GPU clusters of 100–10,000 GPUs.
  • RoCEv2 networks across Nvidia, Arista, and Cisco platforms for RDMA workloads.
  • UFM (Unified Fabric Manager) for IB fabric monitoring, diagnostics, and subnet management.
  • IB and RoCE performance monitoring and tuning: adaptive routing, congestion control (DCQCN/ECN), traffic isolation.
  • NCCL communication tuning: topology detection, ring/tree algorithm selection, GDR configuration.
  • Firmware lifecycle across IB switches and HCAs.
  • Fault diagnosis: link flaps, symbol errors, packet drops, routing anomalies, credit stalls.
  • Coordination with Nvidia/Mellanox support for escalations, bugs, and RMA.
Feed the AIOps substrate
  • Wire IB/RoCE telemetry (ibdiagnet, perfquery, ibstat, PortRcvErrors, PortXmitDiscards, adaptive-routing state) into the platform's collection pipeline.
  • Partner with the platform team to define the Link and Straggler predictors: what a "bad optic 30 minutes from failure" looks like in the counters.
  • Convert every incident into a labeled example the fault-prediction engine can learn from — and every routine mitigation into a workflow the remediation actuator can run.
Job Requirement:
  • 5+ years in data center networking, with at least 3 years focused on InfiniBand or RoCE fabrics
  • Hands-on experience deploying and operating Nvidia/Mellanox InfiniBand switches at scale
  • Strong understanding of IB subnet management, partitioning, and QoS
  • Experience with RoCEv2 deployment including PFC, ECN, DCQCN configuration
  • Proficiency with UFM or equivalent IB fabric management tools
  • Knowledge of 400G/800G optics, cabling standards, and structured cabling best practices
  • Experience diagnosing IB/RoCE network issues using ibdiagnet, perfquery, ibstat, and similar tools
  • Understanding of NCCL and how GPU communication maps to network topology
  • Instinct for telemetry-driven ops — you've either built dashboards/alerts on RDMA counters at scale, or you can articulate the feature set a fabric-health model would need.
  • Runbook-as-code mindset — the diagnostics you run today should become automation next quarter.
-------------------------------------------------------------------- Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.
Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the GPU DC East-West Network SRE Expert (SME) [Remote] in San Jose, CA vacancy
  • $180k - $320k

     ...policy. NeoCloud is building an AI-operated GPU cloud spanning 4 US DCs, APAC sites, and...  ...— DCI, WAN, internet edge, tenant-facing networks — and you build the automation that lets...  ...traffic engineering, and capacity planning for DC networks ~ AIOps aptitude — you've... 
    Network
    Remote job
    Full time
    Local area

    Bitdeer

    San Jose, CA
    18 days ago
  •  ...execute. NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all of that...  ...: GPU locality, NVLink domain awareness, network rail affinity. Custom Resource...  ...GitOps workflows (ArgoCD/Flux) ~ Strong SRE background: SLI/SLO frameworks, incident management... 
    Network
    Remote job
    Full time
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    29 days ago
  •  ...stalls a $50M training run. Bitdeer is building an AI-operated GPU cloud. Storage is where AI workloads either fly or fall over: a...  ...direct GPU-to-storage data paths. Deploy and manage storage networking (NFS over RDMA, NVMe-oF, high-speed storage fabrics) and Nvidia... 
    Network
    Remote job
    Full time
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    29 days ago
  •  ...your career.THE TEAM:AMD's Data Center GPU organization is transforming the...  ...for its Tier 1 CSP Business Group (DC GPU CBG), focused on driving our business...  ..., media, analysts, technical experts and senior executivesPossess a network of industry relationships with potential... 
    Network
    Work experience placement

    AMD

    Santa Clara, CA
    14 days ago
  • Job Title: GPU Network Engineer Job Location: Sunnyvale, CA (hybrid)Job Salary: 200k-250k + BenefitsRequirements: Data center, HPC networking...  ..., please read on!What You Will Be Doing:Design and deploy east-west network fabrics for GPU-to-GPU, rack-to-rack, and cluster-to-... 
    Network
    Local area
    Relocation

    CyberCoders

    San Jose, CA
    4 days ago
  • Prolific is seeking fluent Hindi speakers for its Expert Network in San Jose, California. The role involves completing AI training tasks, judging AI performance, and improving AI models using legal expertise. Candidates must have strong English skills, attention to detail... 
    Network
    Remote job
    Self employment
    Work from home
    Flexible hours

    Prolific

    San Jose, CA
    3 days ago
  •  ...supports critical infrastructure at world-class data center facilities, performing install, configure, and maintenance of server and network equipment. You'll handle cabling, environmental monitoring, incident response, and documentation, with on-call rotation and night/... 
    Network
    Full time
    Night shift
    Rotating shift
    Weekend work

    Reboot Monkey

    Milpitas, CA
    2 days ago
  •  ...beyond. Together, we advance your career. THE TEAMAMD's Data Center GPU organization is transforming the AI and HPC landscape. Our...  ...model architectures (e.g., LLMs, transformer variants, graph neural networks), datatypes, and scaling methodologies to anticipate future... 
    Network
    Remote work

    AMD

    San Jose, CA
    a month ago
  • $167.7k - $245.2k

     ...looking to excel in a fast-paced environment creating groundbreaking network management solutions and want to make a significant long-...  ...with the ability to translate complex technical concepts for non-experts while fostering collaboration and driving excellence within... 
    Network
    Full time
    Temporary work
    Local area
    Flexible hours

    020 Cisco Systems, Inc.

    Milpitas, CA
    4 days ago
  •  ...NeoCloud is building an AI-operated GPU cloud. That doesn't mean fewer...  ...Monitor GPU cluster health, network status, storage systems, and...  ...Collect diagnostic data for L2/SME escalation: logs, DCGM output,...  ...ServiceNow/Jira). Perform physical DC tasks: cable installation,... 
    Network
    Full time
    Local area
    Shift work
    Night shift

    Bitdeer Technologies Group

    San Jose, CA
    a month ago
  •  ...for an  AI Devops Infrastructure Engineer/GPU Infrastructure Engineer. Job Title:...  ...foundation in infrastructure engineering, DevOps/SRE, platform engineering, or similar...  ...containerized environments, including compute, networking, storage, and accelerator scheduling.... 
    Network
    Contract work

    Maxonic

    San Jose, CA
    7 days ago
  •  ...seeking a visionary and hands-on Cloud SRE Architect to lead the design,...  ...the end-to-end architecture across CPU, GPU, RDS, storage, networking, serverless, and AI services, ensuring...  ...frameworks, and three operational tiers (Edge DC → Regional Controller → Global Hub).... 
    Network
    Remote job
    Full time
    Contract work
    Local area
    Shift work

    Bitdeer

    San Jose, CA
    a month ago
  •  ...Site Reliability Engineer (SRE) Location: Santa Clara Valley (Cupertino), California, Hybrid. Duration: 6+ Months Job Description...  ...: Puppet, Chef, Ansible, or Salt. Understanding of standard networking protocols and components such as: DNS, ECMP, TCP/IP, ICMP, the... 
    Network

    Zortech Solutions

    Cupertino, CA
    2 days ago
  •  ...bounded contexts of the NeoCloud SRE platform — the multi-region...  ...observes, protects, and operates a GPU rental fleet across self-built...  ...probe. Hardware Lifecycle & DC Ops: hardware-lifecycle, dc-...  ...optimizer, gpu-efficiency-dashboard, network-stability-dashboard, patching-... 
    Network
    Remote job
    Full time
    Contract work
    Local area

    Bitdeer

    San Jose, CA
    a month ago
  •  .... Together, we advance your career. THE TEAM:AMD's Data Center GPU organization is transforming the AI and HPC landscape. Our mission...  ...Familiarity with data center infrastructure including networking, cooling, and platform design Experience in product operations,... 
    Network
    Remote work
    Flexible hours

    AMD

    Santa Clara, CA
    a month ago
  •  ...SRE Role - SRE Location, RTP/NC and San Jose, CA Duration - Fulltime Job Description: Must Have Technical/Functional...  ...in K8s Mandatory and good knowledge with K8s storage and networking. Should have deployed applications in Kubernetes. Good knowledge... 
    Network
    Full time

    RGH - Global Ltd

    San Jose, CA
    1 day ago
  •  ...Head of GPU Cloud About the Company Building software foundation for next-generation accelerated computing platform for large...  ...experience with GPU memory behavior and high-bandwidth networking is essential. The role also requires a leader with strong technical... 
    Network

    Confidential

    San Jose, CA
    2 days ago
  •  ...Job Description Job Description GPU Programmer / Software Engineer Job Type: Contractor Location: Remote Job Overview We are seeking experienced GPU Programmers / Software Engineers to design and optimize GPU-based tasks for AI and LLM applications. You... 
    Remote job
    For contractors

    YO AI Labs

    San Jose, CA
    13 days ago
  •  ...Position: Site Reliability Engineering (SRE) Location: Santa Clara, CA (Onsite) Duration: W2 / C2C Contract Experience: 10...  ...infrastructure experience including EC2, SSM, vulnerability management ,VPC networking, Web Application Firewalls, ECS/Fargate, IAM • WS Terraform... 
    Network
    Contract work
    Immediate start

    Syntricate Technologies

    Santa Clara, CA
    12 hours ago
  •  ...gWMS Functional Expert San Jose, CA onsite 1) WH Automation 2) gWMS Feature requests Understanding logistics business...  ...Logistics Operations, Systems Engineering, Project Management, IT Network/Infrastructure teams, and third-party logistics (3PL) providers... 
    Network

    Shrive Technologies LLC

    San Jose, CA
    3 days ago
  •  ...optical transceivers in advanced CMOS PDKs. We are seeking an expert in analog/mixed-signal circuit design and architecture.THE PERSON...  ...and correction), and analog front-end circuits (pad matching networks, ESD, CTLE, linear amplifier)To define circuit-driven micro-architectures... 
    Network

    AMD

    San Jose, CA
    a month ago
  •  ...Senior Systems Software Engineer – GPU Performance at Scale We are looking for a dedicated engineer for the Senior Systems Software...  ...next‑generation datacenter builds for NVIDIA GPUs, CPUs, and networking hardware. Engage early with HW/FW/SW/platform internal and... 
    Network

    NVIDIA

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self‑driving cars...  ...with next-generation datacenter builds for NVIDIA GPUs, CPUs, and networking hardware. Engage early with HW/FW/SW/platform internal and... 
    Network

    NVIDIA Gruppe

    Santa Clara, CA
    3 days ago
  • $138k - $184k

     ...Enterprise Account Executive - West Eightfold was founded with a vision to solve for employment in our society. For decades, the...  ...been based on who the individuals are and the strength of their network, not their potential. Eightfold leverages artificial intelligence... 
    Network
    Work at office
    Remote work
    Flexible hours

    Eightfold LLC

    Santa Clara, CA
    4 days ago
  • $115k - $130k

     ...provider of advanced server, storage, and networking solutions for Data Center, Cloud...  ...support, and troubleshooting connectivity for GPU clusters and storage systems. Essential...  ...architectures, including spine-leaf topologies and east-west traffic patterns. Basic proficiency with... 
    Network
    Work at office
    Worldwide
    Night shift

    Supermicro

    San Jose, CA
    5 days ago
  • $196k - $310.5k

     ...next era of computing. An era in which our GPU acts as the brains of computers, robots,...  ...on the world. NVIDIA Enterprise Network Architecture team is seeking experienced...  ...implicitly-trusted model toward explicit east-west segmentation — internal/collapsed firewalls... 
    Network
    Work at office

    NVIDIA Gruppe

    Santa Clara, CA
    1 day ago
  • $27 - $41 per hour

     ...dedicated year-round TurboTax Local Service Experts in one of our new TurboTax locations...  ...grass-roots marketing initiatives, referral networks, and community partnerships. Serve as the...  ...Washington: $27.00 - $41.00 Washington, DC: $26.00 - $39.00 This position will be eligible... 
    Network
    Work at office
    Local area
    Monday to Friday

    Intuit

    Los Gatos, CA
    2 days ago
  • $86.02k - $117.71k

     ...Pilot (First Officer) - West Region Company: NetJets Aviation, Inc. Area of Interest: Crewmembers Location: US Los Angeles, CA, US...  ...including Medical, Dental, and Vision benefits, with access to robust networks of nationwide providers. NetJets offers benefits so you can LIVE... 
    Network
    Temporary work
    Flexible hours
    Night shift
    Weekend work
    Weekday work

    NetJets

    San Jose, CA
    3 days ago
  •  ...Job Description Job Description Developer & Infrastructure Expert Role Type: Contractor Location: Remote Job Overview We...  ...workflows across software development, cloud infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,... 
    Remote job
    For contractors

    YO AI Labs

    San Jose, CA
    3 days ago
  • $154.9k - $225.3k

     ...protect enterprise, data center, and cloud networks from evolving cyber threats. This team...  ...environment. You will act as a subject matter expert, mentoring junior engineers and...  ...Proficiency in selecting and implementing DC-DC converters, LDOs, and managing power sequencing... 
    Network
    Full time
    Temporary work
    Local area
    Flexible hours

    Webex Events (formerly Socio)

    San Jose, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to GPU DC East-West Network SRE Expert (SME) [Remote]. Be the first to apply!