Sr GPU Cloud South-North Network SRE Expert (SRE SME)
Bitdeer Technologies Group
About Bitdeer Technologies Group
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.
Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.
Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
To learn more, visit (Job Description
You own the north-south fabric that connects our AI cloud to the world — and drive the automation that turns EVPN/BGP ops from tickets into policy.
NeoCloud is building an AI-operated GPU cloud spanning 4 US DCs, APAC sites, and Iceland. In this role you own the IP/Ethernet infrastructure that connects that fleet — DCI, WAN, internet edge, tenant-facing networks — and you build the automation that lets the AIOps substrate express network intent as code, not as change tickets.
What you'll own
- EVPN-VxLAN data center fabrics for multi-tenant network isolation and workload mobility.
- Multi-site DCI overlay/underlay connecting 4 US DCs with consistent L2/L3 services.
- BGP (eBGP/iBGP), OSPF, ECMP load balancing, and VxLAN across spine-leaf architectures.
- IP transit, peering relationships, and internet edge infrastructure.
- WAN/IP backbone connecting US DCs to APAC and Iceland sites.
- Network equipment: Arista, Cisco, and Palo Alto switches, routers, and firewalls.
- Network automation with Ansible, Terraform, Nautobot/NetBox for IPAM/DCIM.
- Network health, capacity, and SLA monitoring across all sites.
- Network architecture docs, standard configs, and change procedures.
Feed the AIOps substrate
- Every route policy, ACL, VRF, and peering session lives as code in the platform's config-management substrate — no config drift, no snowflakes.
- Every BGP flap, every DCI degradation feeds the network-stability predictor.
- Every routine change becomes a workflow the remediation actuator can execute under a maintenance window.
Job Requirement:
- 5+ years in data center network engineering, with hands-on EVPN-VxLAN deployment experience
- Strong BGP expertise (eBGP/iBGP, route policies, communities, traffic engineering)
- Experience designing and operating multi-site DCI with EVPN multi-homing
- Proficiency with Arista EOS and/or Cisco NX-OS in spine-leaf data center environments
- Experience with firewall platforms (Palo Alto, Fortinet) for perimeter and inter-tenant security
- Hands-on experience with IP transit provisioning and internet peering
- Familiarity with network automation tools (Ansible, Python/Netmiko, Nautobot/NetBox)
- Strong understanding of QoS, traffic engineering, and capacity planning for DC networks
- AIOps aptitude — you've either driven a network-automation program from ticket-driven ops to intent-driven policy, or you have a clear thesis on how you would.
- Runbook-as-code mindset — configs, changes, and rollbacks should be executable by an agent.
--------------------------------------------------------------------
Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.
- .... Bitdeer also offers advanced cloud capabilities to customers with... ...NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all... ..., NVLink domain awareness, network rail affinity. Custom Resource... ...(ArgoCD/Flux) ~ Strong SRE background: SLI/SLO frameworks,...CloudSeniorNetworkFull timeLocal area
- ...operations. Bitdeer also offers advanced cloud capabilities to customers with high demand... .... Bitdeer is building an AI-operated GPU cloud. Storage is where AI workloads either... ...data paths. Deploy and manage storage networking (NFS over RDMA, NVMe-oF, high-speed...CloudSeniorNetworkRemote jobFull timeLocal area
$101k - $161k
Company DescriptionArista Networks is an industry leader in data-driven, client-to-cloud networking for large data center, campus and routing environments. What sets... ...Arista’s CloudVision-as-a-Service (CVaaS) global SRE team. SREs at Arista combine strong software...CloudSeniorNetwork$167.7k - $245.2k
...Senior Site Reliability Engineer (SRE), you will build, operate, and... ...backbone supporting both cloud and air-gapped customer deployments... ...issues spanning Kubernetes, networking, storage, and application... ...worldwide network of doers and experts, and you’ll see that the opportunities...CloudSeniorNetworkFull timeTemporary workLocal areaFlexible hours2 days per week$98.9k - $228.7k
...and vendors — including media servers, CDN providers, cloud-native services, and edge networking.Ensuring consistent standards for CI/CD pipelines, deployment... ...rollout strategies across teams.Acting as the primary SRE partner for multiple engineering teams building real-...CloudSeniorNetworkFull timeWork at officeRemote workFlexible hours$256k - $414k
...NOW is the global leader in cloud gaming, dedicated to making high... ...of high-performance networking for GPU-based cloud infrastructure. This... ...teams, hardware vendors, and SRE groups to influence technology... ...large-scale configurations using SR-IOV, Xen virtualization, or...CloudSeniorNetworkFull timeLocal area- .... Bitdeer also offers advanced cloud capabilities to customers with... ...Bitdeer is building an AI-operated GPU cloud — a global fleet of self-... ..., and operates the fleet. The SRE Platform team builds the... ...that every other squad — storage, network, GPU, K8S, and L1 operators — depends...CloudNetworkFull timeContract workInternshipLocal area
$186.9k - $267.7k
...approach empowers our customers to expertly deploy and manage AI-powered... ...Site Reliability Engineer (SRE), you will provide technical leadership... ...for operating large-scale cloud and on-prem deployments.Your... ..., databases, and networking.Drive automation to eliminate...CloudNetworkFull timeTemporary workLocal areaFlexible hours2 days per week$177.6k - $257.4k
...Site Reliability Engineering (SRE) Technical Leader on the Intersight... ..., and security of our cloud platforms. The broader team is... ...have hands-on SRE or systems/network administration experience, with... ...complex technical concepts for non-experts while fostering collaboration...CloudNetworkFull timeTemporary workLocal areaImmediate startFlexible hours- ...Job Description Job Description Java SRE Engineer Onsite San Francisco Bay Area... ...Platform Engineer to support large-scale cloud migrations and production systems on AWS... ...understanding of distributed systems and networking Preferred Qualifications: Experience...CloudNetwork
- ...inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of... ...RoleWe are seeking an experienced IT SRE Team Lead to build and run the reliability... ..., collaboration, SaaS, and internal networking. The right candidate will bring a...CloudNetwork
$148.75k - $361k
...infrastructure. The ideal candidate has a strong background in DevOps/SRE practices, cloud infrastructure management, and MLOps tooling — with a passion... ...for ML workloads across AWS and GCP, including GPU/TPU-based training and inference environmentsArchitect and improve...CloudSeniorWork at officeLocal areaRemote workMonday to ThursdayFlexible hours$207k - $300k
...stakeholders.Site Reliability Engineering (SRE) combines software and systems... ...tolerant systems. SRE ensures that Google Cloud's services—both our internally critical and... ...apart so we can rebuild them. We keep our networks up and running, ensuring our users have the...CloudNetwork- ...career.THE TEAM:AMD's Data Center GPU organization is transforming the... ...driving our business success with North America cloud service providers. The role... ...customers, media, analysts, technical experts and senior executivesPossess a network of industry relationships with potential...CloudSeniorNetworkWork experience placement
- .... Bitdeer also offers advanced cloud capabilities to customers with... ...NeoCloud is building an AI-operated GPU cloud. That doesn't mean fewer... ...Monitor GPU cluster health, network status, storage systems, and environmental... ...Collect diagnostic data for L2/SME escalation: logs, DCGM output,...CloudNetworkFull timeLocal areaShift workNight shift
- Webex Central Operations Engineering seeks a Senior Cloud Platform Engineer / SRE to design, operate, automate, and support AWS-based cloud platforms running Kubernetes/EKS in regulated FedRAMP High / IL-5 environments. You will collaborate with Compliance and SecOps to...CloudSenior
$155.8k - $224.2k
...rapidly becoming the new norm. Sr. Staff Cloud EngineerWe are looking for a... ...response.Drive adoption of SRE practices including service level... ..., cloud architecture, cloud networking, security, resiliency, and... ...Experience with AI/ML infrastructure, GPU workloads, data platforms, or...CloudSeniorNetworkFull timeWorldwide$145k - $235.5k
...Our Mission At Palo Alto Networks®, we're united by a shared mission... ...Summary As an AI-Native Cloud FinOps Engineer , you will build... .... Programmatically manage GPU cluster provisioning, optimize... ..., Cloud Engineering, SRE, or Cloud FinOps teams. ~2+...CloudSeniorNetworkFull timeWork at officeLocal areaVisa sponsorshipWork visaShift work$207k - $300k
...experience.8 years of technical experience in network protocols (TCP/IP, BGP, DNS, routing,... ...engineering VPs, Directors, and enterprise Cloud customers.Extensive background in... ...intelligence pipelines.Site Reliability Engineering (SRE) combines software and systems engineering...CloudNetworkShift work$307k - $427k
...practices to enable partners to operate the cloud autonomously while maintaining service... ...and Google rollout velocity.Ensure that SRE principles (SLOs, alerting, runbooks, diagnostics... ...design or with Unix/Linux systems, IP networking, performance and application issues....CloudNetwork- ...Devops Infrastructure Engineer/GPU Infrastructure Engineer.... ...infrastructure engineering, DevOps/SRE, platform engineering, or similar... ..., including compute, networking, storage, and accelerator scheduling... ...Platform Engineering, DevOps, SRE, Cloud Engineering, or AI/ML...CloudNetworkContract work
$152k - $241.5k
...Performance Computing and Visualization. The GPU, our invention, serves as the visual... ...are the GPU Communications Libraries and Networking team at NVIDIA. We deliver libraries like... ...preferably PythonFamiliar with containers, cloud provisioning and scheduling tools (...CloudSeniorNetworkFull time- Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure... .... One person, one GPU.If you'd like to build the world... ...practices across infrastructure, networking, platform engineering, and... ...management frameworks (ITIL, SRE, or equivalent)Excellent communication...CloudSeniorNetworkWork at officeLocal areaFlexible hours
- ...inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of... ...are building a high-performance SRE function to support one of the world’s... ...ownership, and minimal dependency on expert SRE operators.You will collaborate with...CloudShift work
$184k - $287.5k
...are seeking a Senior DevOps / Cloud Simulation Infrastructure Engineer... ...for deploying a robust, multi-GPU pipeline that supports... ...addressing function-to-function networking, gRPC bottlenecks, and in-cluster... ...in production-grade DevOps, SRE, or Infrastructure Engineering...CloudSeniorNetworkFull timeLocal area$184k - $287.5k
...next era of computing. An era in which our GPU acts as the brains of computers, robots,... ...builds for NVIDIA GPUs, CPUs, and networking hardware. Engage early with HW/FW/SW/platform... ...software stacks (CUDA).Experience with modern cloud and container-based enterprise computing...CloudSeniorNetworkFull timeRemote work- Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure... .... One person, one GPU.If you'd like to build the world... ...standards for resource management, networking, and RBAC across the platform.... ...Platform, Infrastructure, or SRE roles, including running...CloudSeniorNetworkWork at officeLocal areaWork from homeFlexible hours
$184k - $287.5k
...System Software Engineer, Software Defined Networking to design, build, and operate highly... ...scalable SDN solutions for NVIDIA's AI Clouds hosting GPU-accelerated workloads — including... ...and performance tuningCollaborate with SRE, DevOps, and network engineering teams...CloudSeniorNetworkFull time$184k - $287.5k
...customers. We are seeking an expert Solutions Architect to assist... ...critical business needs and support cloud service integration for NVIDIA... ...including AI accelerators and networking as it relates to the... ...diagnostics.Hands-on experience with GPU systems in general including...CloudSeniorNetworkFull time$208k - $327.75k
...computing. An era in which our GPU acts as the brains of... ...building the software-defined network for the AI factory. Our SDN team... ...ships across NVIDIA and NVIDIA Cloud Partner (NCP) DSX deployments.... ...throughout its entire scope: North/South tenant networking, East/West GPUDirect...CloudSeniorNetworkFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Sr GPU Cloud South-North Network SRE Expert (SRE SME). Be the first to apply!
- fulfillment expert San Jose, CA
- guest service support expert San Jose, CA
- technology expert San Jose, CA
- senior operations technician San Jose, CA
- senior cloud service delivery manager San Jose, CA
- senior it service manager San Jose, CA
- senior project engineer San Jose, CA
- senior chief engineer San Jose, CA
- sr operations manager San Jose, CA
- senior physical design engineer San Jose, CA



