Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Sr GPU Cloud South-North Network SRE Expert (SRE SME) [Remote]

Full-time

Bitdeer Technologies Group

San Jose, CA
  • Remote job

About Bitdeer Technologies Group

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

To learn more, visit  (

Job Description

You own the north-south fabric that connects our AI cloud to the world — and drive the automation that turns EVPN/BGP ops from tickets into policy.

NeoCloud is building an AI-operated GPU cloud spanning 4 US DCs, APAC sites, and Iceland. In this role you own the IP/Ethernet infrastructure that connects that fleet — DCI, WAN, internet edge, tenant-facing networks — and you build the automation that lets the AIOps substrate express network intent as code, not as change tickets.

What you'll own

  • EVPN-VxLAN data center fabrics for multi-tenant network isolation and workload mobility.
  • Multi-site DCI overlay/underlay connecting 4 US DCs with consistent L2/L3 services.
  • BGP (eBGP/iBGP), OSPF, ECMP load balancing, and VxLAN across spine-leaf architectures.
  • IP transit, peering relationships, and internet edge infrastructure.
  • WAN/IP backbone connecting US DCs to APAC and Iceland sites.
  • Network equipment: Arista, Cisco, and Palo Alto switches, routers, and firewalls.
  • Network automation with Ansible, Terraform, Nautobot/NetBox for IPAM/DCIM.
  • Network health, capacity, and SLA monitoring across all sites.
  • Network architecture docs, standard configs, and change procedures.

Feed the AIOps substrate

  • Every route policy, ACL, VRF, and peering session lives as code in the platform's config-management substrate — no config drift, no snowflakes.
  • Every BGP flap, every DCI degradation feeds the network-stability predictor.
  • Every routine change becomes a workflow the remediation actuator can execute under a maintenance window.

Job Requirement:

  • 5+ years in data center network engineering, with hands-on EVPN-VxLAN deployment experience
  • Strong BGP expertise (eBGP/iBGP, route policies, communities, traffic engineering)
  • Experience designing and operating multi-site DCI with EVPN multi-homing
  • Proficiency with Arista EOS and/or Cisco NX-OS in spine-leaf data center environments
  • Experience with firewall platforms (Palo Alto, Fortinet) for perimeter and inter-tenant security
  • Hands-on experience with IP transit provisioning and internet peering
  • Familiarity with network automation tools (Ansible, Python/Netmiko, Nautobot/NetBox)
  • Strong understanding of QoS, traffic engineering, and capacity planning for DC networks
  • AIOps aptitude — you've either driven a network-automation program from ticket-driven ops to intent-driven policy, or you have a clear thesis on how you would.
  • Runbook-as-code mindset — configs, changes, and rollbacks should be executable by an agent.

--------------------------------------------------------------------

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Vacancy posted 13 days ago
Similar jobs that could be interesting for youBased on the Sr GPU Cloud South-North Network SRE Expert (SRE SME) [Remote] in San Jose, CA vacancy
  •  .... Bitdeer also offers advanced cloud capabilities to customers with...  ...NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all...  ..., NVLink domain awareness, network rail affinity. Custom Resource...  ...(ArgoCD/Flux) ~ Strong SRE background: SLI/SLO frameworks,... 
    Cloud
    Senior
    Network
    Remote job
    Full time
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    13 days ago
  •  ...operations. Bitdeer also offers advanced cloud capabilities to customers with high demand...  .... Bitdeer is building an AI-operated GPU cloud. Storage is where AI workloads either...  ...data paths. Deploy and manage storage networking (NFS over RDMA, NVMe-oF, high-speed... 
    Cloud
    Senior
    Network
    Full time
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    2 days ago
  • $170k - $277k

     ...Our Mission At Palo Alto Networks®, we’re united by a shared mission—to protect our digital...  ...the technical authority for our global SRE and Platform Engineering initiatives...  ...observability pipelines scale seamlessly across cloud-native environments. Collaborate Cross... 
    Cloud
    Senior
    Network
    Full time
    Work at office
    Visa sponsorship
    Work visa
    Flexible hours

    Palo Alto Networks, Inc.

    Santa Clara, CA
    8 hours ago
  •  ...Palo Alto Networks is seeking a Senior Principal Engineer/Architect to serve as the technical authority for global SRE and Platform Engineering initiatives across the US and India. You...  ...ensure reliability and cost efficiency in cloud-native #J-18808-Ljbffr Jobleads-US
    Cloud
    Senior
    Network

    Jobleads-US

    Santa Clara, CA
    5 days ago
  • $101k - $161k

    Company DescriptionArista Networks is an industry leader in data-driven, client-to-cloud networking for large data center, campus and routing environments. What sets...  ...Arista’s CloudVision-as-a-Service (CVaaS) global SRE team. SREs at Arista combine strong software... 
    Cloud
    Senior
    Network

    Arista Networks

    Santa Clara, CA
    4 days ago
  • $256k - $414k

     ...NOW is the global leader in cloud gaming, dedicated to making high...  ...of high-performance networking for GPU-based cloud infrastructure. This...  ...teams, hardware vendors, and SRE groups to influence technology...  ...large-scale configurations using SR-IOV, Xen virtualization, or... 
    Cloud
    Senior
    Network
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    4 days ago
  • $364k

     ...Google’s technical Infrastructure and Google Cloud products.Architect observability...  ..., Google’s Site Reliability Engineering (SRE) is an engineering discipline for building...  ...apart so we can rebuild them. We keep our networks up and running, ensuring our users have the... 
    Cloud
    Senior
    Network

    Google

    Sunnyvale, CA
    1 day ago
  •  ...inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of...  ...RoleWe are seeking an experienced IT SRE Team Lead to build and run the reliability...  ..., collaboration, SaaS, and internal networking. The right candidate will bring a... 
    Cloud
    Network

    Cerebras Systems

    Sunnyvale, CA
    1 day ago
  • $148.75k - $361k

     ...infrastructure. The ideal candidate has a strong background in DevOps/SRE practices, cloud infrastructure management, and MLOps tooling — with a passion...  ...for ML workloads across AWS and GCP, including GPU/TPU-based training and inference environmentsArchitect and improve... 
    Cloud
    Senior
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    1 day ago
  • $207k - $300k

     ...stakeholders.Site Reliability Engineering (SRE) combines software and systems...  ...tolerant systems. SRE ensures that Google Cloud's services—both our internally critical and...  ...apart so we can rebuild them. We keep our networks up and running, ensuring our users have the... 
    Cloud
    Network

    Google

    San Jose, CA
    2 days ago
  •  ...Job Description Job Description Java SRE Engineer Onsite San Francisco Bay Area...  ...Platform Engineer to support large-scale cloud migrations and production systems on AWS...  ...understanding of distributed systems and networking Preferred Qualifications: Experience... 
    Cloud
    Network

    Eitacies Inc

    Santa Clara, CA
    19 days ago
  • $155.8k - $224.2k

     ...rapidly becoming the new norm.  Sr. Staff Cloud EngineerWe are looking for a...  ...response.Drive adoption of SRE practices including service level...  ..., cloud architecture, cloud networking, security, resiliency, and...  ...Experience with AI/ML infrastructure, GPU workloads, data platforms, or... 
    Cloud
    Senior
    Network
    Full time
    Worldwide

    Bloom Energy

    San Jose, CA
    3 days ago
  • $100k - $170k

     ...Infrastructure as Code using Terraform for AWS cloud resources Develop and optimize CI/CD...  ...this role: ~5+ years of experience in SRE, DevOps, or Platform Engineering roles...  ...recovery planning Experience with Zero Trust Networking (ZTNA) or VPN solutions Background in... 
    Cloud
    Network
    Full time
    Work at office
    Immediate start
    Visa sponsorship
    Night shift

    E-Space

    Saratoga, CA
    17 days ago
  •  .... Bitdeer also offers advanced cloud capabilities to customers with...  ...Bitdeer is building an AI-operated GPU cloud — a global fleet of self-...  ..., and operates the fleet. The SRE Platform team builds the...  ...that every other squad — storage, network, GPU, K8S, and L1 operators — depends... 
    Cloud
    Network
    Remote job
    Full time
    Contract work
    Internship
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    13 days ago
  • $207k - $300k

     ...Experience with Large Language Model.Site Reliability Engineering (SRE) combines software and systems engineering to build and run...  ...warranties by taking things apart so we can rebuild them. We keep our networks up and running, ensuring our users have the best and fastest... 
    Cloud
    Network

    Google

    Sunnyvale, CA
    2 days ago
  • $145k - $235.5k

     ...Our Mission At Palo Alto Networks®, we’re united by a shared mission...  ...Summary As an AI-Native Cloud FinOps Engineer , you will build...  .... Programmatically manage GPU cluster provisioning, optimize...  ..., Cloud Engineering, SRE, or Cloud FinOps teams. ~2+... 
    Cloud
    Senior
    Network
    Full time
    Work at office
    Local area
    Visa sponsorship
    Work visa
    Shift work

    Palo Alto Networks

    Santa Clara, CA
    4 days ago
  • $207k - $300k

     ...experience.8 years of technical experience in network protocols (TCP/IP, BGP, DNS, routing,...  ...engineering VPs, Directors, and enterprise Cloud customers.Extensive background in...  ...intelligence pipelines.Site Reliability Engineering (SRE) combines software and systems engineering... 
    Cloud
    Network
    Shift work

    Google

    Sunnyvale, CA
    8 hours ago
  • Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure...  .... One person, one GPU.If you'd like to build the world...  ...practices across infrastructure, networking, platform engineering, and...  ...management frameworks (ITIL, SRE, or equivalent)Excellent communication... 
    Cloud
    Senior
    Network
    Work at office
    Local area
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  • $152k - $241.5k

     ...Performance Computing and Visualization. The GPU, our invention, serves as the visual...  ...are the GPU Communications Libraries and Networking team at NVIDIA. We deliver libraries like...  ...preferably PythonFamiliar with containers, cloud provisioning and scheduling tools (... 
    Cloud
    Senior
    Network
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...are seeking a Senior DevOps / Cloud Simulation Infrastructure Engineer...  ...for deploying a robust, multi-GPU pipeline that supports...  ...addressing function-to-function networking, gRPC bottlenecks, and in-cluster...  ...in production-grade DevOps, SRE, or Infrastructure Engineering... 
    Cloud
    Senior
    Network
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...System Software Engineer, Software Defined Networking to design, build, and operate highly...  ...scalable SDN solutions for NVIDIA's AI Clouds hosting GPU-accelerated workloads — including...  ...and performance tuningCollaborate with SRE, DevOps, and network engineering teams... 
    Cloud
    Senior
    Network
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...next era of computing. An era in which our GPU acts as the brains of computers, robots,...  ...builds for NVIDIA GPUs, CPUs, and networking hardware. Engage early with HW/FW/SW/platform...  ...software stacks (CUDA).Experience with modern cloud and container-based enterprise computing... 
    Cloud
    Senior
    Network
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure...  .... One person, one GPU.If you'd like to build the world...  ...standards for resource management, networking, and RBAC across the platform....  ...Platform, Infrastructure, or SRE roles, including running... 
    Cloud
    Senior
    Network
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  • $200k - $322k

     ...computing. An era in which our GPU acts as the brains of...  ...skilled Senior Staff SRE to join our dynamic team...  ...both on-prem and cloud. Join us in this exciting...  ...hardware optimizations (SR-IOV/ DPU)Experience with...  ...InfrastructureUnderstanding of Network Protocols and... 
    Cloud
    Senior
    Network
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...customers. We are seeking an expert Solutions Architect to assist...  ...critical business needs and support cloud service integration for NVIDIA...  ...including AI accelerators and networking as it relates to the...  ...diagnostics.Hands-on experience with GPU systems in general including... 
    Cloud
    Senior
    Network
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $208k - $327.75k

     ...computing. An era in which our GPU acts as the brains of...  ...building the software-defined network for the AI factory. Our SDN team...  ...ships across NVIDIA and NVIDIA Cloud Partner (NCP) DSX deployments....  ...throughout its entire scope: North/South tenant networking, East/West GPUDirect... 
    Cloud
    Senior
    Network
    Full time

    Nvidia

    Santa Clara, CA
    8 hours ago
  •  ...inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of...  ...are building a high-performance SRE function to support one of the world’s...  ...ownership, and minimal dependency on expert SRE operators.You will collaborate with... 
    Cloud
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    4 days ago
  • $140k - $215k

     ...CrowdStrike, Site Reliability Engineering (SRE) is at the forefront of ensuring the reliability and scalability of our cloud-native security platform. In this role, you'...  ...employees regardless of level or roleEmployee Networks, geographic neighborhood groups, and... 
    Cloud
    Network
    Full time
    Work experience placement
    Work at office
    Local area

    CrowdStrike

    Sunnyvale, CA
    1 day ago
  • $133.5k - $272k

     ...NetskopeNetskope (NASDAQ: NTSK) is a leader in modern security and networking for the cloud and AI era. We secure and accelerate cloud, data, and AI in...  ...-Functional Collaboration: Partner with Product Managers, SRE, and adjacent engineering teams to align on requirements and... 
    Cloud
    Senior
    Network

    Netskope

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

     ...building reliable systems software for cloud-scale GPU infrastructure, we encourage you to contribute...  .../OS state, CPU, memory, disk, and networking.Develop inventory, enrollment, node...  ...Collaborate with backend, infrastructure, SRE, security, and datacenter operations... 
    Cloud
    Senior
    Network
    Full time
    Local area
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Sr GPU Cloud South-North Network SRE Expert (SRE SME) [Remote]. Be the first to apply!