Sr GPU Cloud K8S Expert (SRE SME)
Bitdeer Technologies Group
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.
Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.
Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
To learn more, visit (Position Overview
You run the control plane where AIOps meets tenants — where topology-aware scheduling, self-healing, and agent-driven remediation actually execute.
NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all of that lands on real customer workloads: the topology-aware scheduler places jobs on the right NVLink domain, the operator drains and reschedules around predicted faults, and the tenant boundary is enforced against a Bare-Metal-as-a-Service backend. In this role you design, deploy, and operate that control plane — and you make sure the AIOps substrate can reach in and remediate without a human on the pager.
What you'll own
- Production Kubernetes clusters optimized for GPU workloads at scale (100–10,000 GPUs).
- Nvidia GPU operator, device plugin, MIG configuration, and GPU time-slicing policies.
- Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity.
- Custom Resource Definitions (CRDs) for GPU workload lifecycle management.
- AI framework integrations: Slurm on K8S, Ray on K8S, Kubeflow.
- Multi-tenant isolation: namespaces, network policies, resource quotas, RBAC, pod security standards.
- Bare-Metal as a Service (BMaaS): automated provisioning, tenant onboarding, lifecycle, reclamation.
- Terraform providers and modules for infrastructure-as-code across GPU clusters.
- SLIs/SLOs for cluster availability, job completion rates, and provisioning latency.
- Incident management: runbook automation, escalation, post-incident reviews.
- Monitoring stack: Prometheus, Grafana, Alertmanager, PagerDuty.
- GPU node failure handling: automated detection, drain/cordon/taint, workload rescheduling.
Feed the AIOps substrate
- The remediation-actuator and workflow engine land here — you make the control plane safe for automated action.
- Your CRDs are the schema the platform's predictors and remediators write against.
- Every human intervention you do this quarter becomes an autonomous workflow next quarter.
What success looks like in year 1
- Automated drain/reschedule around predicted GPU faults, at scale, without customer impact.
- BMaaS live for external tenants with self-service onboarding.
- Cluster availability and job-completion SLOs published and met.
Job Requirement:
- 5+ years in Kubernetes operations, with at least 2 years managing GPU workloads on K8S
- Deep understanding of Nvidia GPU operator, device plugin, and GPU scheduling in K8S
- Experience with topology-aware scheduling and GPU-specific resource management
- Hands-on experience building multi-tenant K8S platforms with strong isolation guarantees
- Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom)
- Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux)
- Strong SRE background: SLI/SLO frameworks, incident management, capacity planning
- Experience with Prometheus, Grafana, and alerting at scale
- Strong programming skills in Go or Python for operator/CRD development
- AIOps aptitude — you think of the K8S control plane as an execution surface for automated remediation, not just a scheduler. You've either wired an autoscaler/remediator loop into K8S or you can design one.
- Runbook-as-code mindset — every SRE playbook you write should be executable by the platform.
--------------------------------------------------------------------
Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.
- ...equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial... ...stalls a $50M training run. Bitdeer is building an AI-operated GPU cloud. Storage is where AI workloads either fly or fall over: a...CloudSeniorFull timeLocal area
- ...equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial... ...from tickets into policy. NeoCloud is building an AI-operated GPU cloud spanning 4 US DCs, APAC sites, and Iceland. In this role you...CloudSeniorFull timeLocal area
- ...together, we’ll advance your career.THE ROLE:The Sr. Director, Pre-Si System Validation leads... ...and pre-silicon validation for Data Center GPU platforms. This role ensures early, system-accurate insights across AI, HPC, and cloud-scale environments, accelerating software...CloudSeniorShift work
$256k - $414k
...GeForce NOW is the global leader in cloud gaming, dedicated to making... ...-performance networking for GPU-based cloud infrastructure. This... ...platform teams, hardware vendors, and SRE groups to influence technology... ...-scale configurations using SR-IOV, Xen virtualization, or...CloudSeniorFull timeLocal area$148.75k - $361k
...infrastructure. The ideal candidate has a strong background in DevOps/SRE practices, cloud infrastructure management, and MLOps tooling — with a passion... ...for ML workloads across AWS and GCP, including GPU/TPU-based training and inference environmentsArchitect and improve...CloudSeniorWork at officeLocal areaRemote workMonday to ThursdayFlexible hours$101k - $161k
...DescriptionArista Networks is an industry leader in data-driven, client-to-cloud networking for large data center, campus and routing environments... ...our growing Arista’s CloudVision-as-a-Service (CVaaS) global SRE team. SREs at Arista combine strong software engineering...CloudSenior- Webex Central Operations Engineering seeks a Senior Cloud Platform Engineer / SRE to design, operate, automate, and support AWS-based cloud platforms running Kubernetes/EKS in regulated FedRAMP High / IL-5 environments. You will collaborate with Compliance and SecOps to...CloudSenior
$152k - $241.5k
...Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and... ...a scripting language, preferably PythonFamiliar with containers, cloud provisioning and scheduling tools (Kubernetes, SLURM, Ansible,...CloudSeniorFull time$208k - $327.75k
...worldwide. As data volumes explode and computational demands rise, GPU-accelerated storage solutions like GPUDirect Storage and cuFile... ...management.Previous roles defining product strategies or requirements in cloud-based or HPC environments.NVIDIA is renowned globally as a top-...CloudSeniorFull timeWorldwide- ...Together, we advance your career. THE ROLE:Join AMD’s Datacenter GPU Product Application Engineering team to lead a high-performing group... ...in customer-facing engineering roles within data center, GPU, cloud, or HPC environments. You are equally comfortable leading teams,...CloudSeniorFor contractors
$184k - $287.5k
...potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars... ...computing software stacks (CUDA).Experience with modern cloud and container-based enterprise computing architectures, with Slurm...CloudSeniorFull timeRemote work$184k - $287.5k
...our largest customers. We are seeking an expert Solutions Architect to assist customers in... ...address critical business needs and support cloud service integration for NVIDIA technology... ...error diagnostics.Hands-on experience with GPU systems in general including but not...CloudSeniorFull time- ...Senior Systems Software Engineer – GPU Performance at Scale We are looking for a dedicated engineer for the Senior Systems Software... ...accelerated computing software stacks (CUDA). Experience with modern cloud and container‑based enterprise computing architectures, with...CloudSenior
- ...advance your career.THE TEAM:AMD's Data Center GPU organization is transforming the industry... ...our business success with North America cloud service providers. The role requires deep... ...- customers, media, analysts, technical experts and senior executivesPossess a network of...CloudSeniorWork experience placement
$200k - $230k
...fast in many key markets such as AI, HPC, Cloud Computing, Storage, etc. To meet the market... ...integrated and tested.Job Summary:The Sr. Manager, Solution Management collaborates... ...infrastructuresHands-on experience with Nvidia and AMD GPU based deployments are desirableSolid...CloudSenior- ...of high-performance backend networks supporting large-scale AMD GPU clusters. The engineer will own the network path from the GPU server... ...and operating production data center networks for AI, GPU, HPC, cloud, or other large-scale distributed computing environments....CloudSenior
$155.8k - $224.2k
...scarcity are rapidly becoming the new norm. Sr. Staff Cloud EngineerWe are looking for a Sr. Staff... ...and incident response.Drive adoption of SRE practices including service level objectives... ....Experience with AI/ML infrastructure, GPU workloads, data platforms, or large-scale...CloudSeniorFull timeWorldwide$98.9k - $228.7k
...platforms, and vendors — including media servers, CDN providers, cloud-native services, and edge networking.Ensuring consistent standards... ...rollout strategies across teams.Acting as the primary SRE partner for multiple engineering teams building real-time features...CloudSeniorFull timeWork at officeRemote workFlexible hours- ...Description Job Description Developer & Infrastructure Expert Role Type: Contractor Location: Remote Job... ...evaluate AI-powered workflows across software development, cloud infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands...CloudRemote jobFor contractors
$168k - $258.75k
...potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars... ...to make sophisticated trade-offsAbility to work closely with cloud partnersWays to stand out from the crowd: Strong background in...CloudFull time- .... Bitdeer also offers advanced cloud capabilities to customers with... ...Bitdeer is building an AI-operated GPU cloud — a global fleet of self-... ..., and operates the fleet. The SRE Platform team builds the... ...squad — storage, network, GPU, K8S, and L1 operators — depends on....CloudFull timeContract workInternshipLocal area
$140k - $165k
...we drive the evolution of advancing mobile technology, empowering cloud computing, and pioneering future technologies. Our cutting-edge... ...: Design and deploy on-prem AI infrastructure — including GPU clusters, model serving (e.g., vLLM, TGI, Triton), vector DBs (e....CloudSeniorFull time$167.7k - $245.2k
...control.As a Senior Site Reliability Engineer (SRE), you will build, operate, and... ...the operational backbone supporting both cloud and air-gapped customer deployments, develop... ...to that our worldwide network of doers and experts, and you’ll see that the opportunities to...CloudSeniorFull timeTemporary workLocal areaFlexible hours2 days per week$7,200 per month
...re UNSTOPPABLE for our employees! Job Overview Experience Experts are our most capable experts with a passion for delivering best-... ...training to keep current and develop knowledge and skills to act as SME in all things T-Mobile. Consistently leverages digital tools...SeniorHourly payFull timeTemporary workPart timeWork experience placementLocal areaFlexible hours$224k - $356.5k
We are now looking for a GPU System Performance Architect:The NVIDIA Architecture group is looking for extraordinary computer architects... ...architectures to extend the state of the art in GPU-accelerated cloud computing.You'll analyze trade-offs in system performance, cost...CloudFull timeWork experience placement$155k - $185k
...advanced server, storage, and networking solutions for Data Center, Cloud Computing, Enterprise IT, Hadoop/ Big Data, Hyperscale, HPC and... ...: Be at the forefront of developing and integrating advanced GPU systems and rack-scale solutions that power cloud providers, hyperscalers...CloudSeniorWorldwide$168k - $270.25k
...Performance Computing, and Visualization. The GPU, our invention, serves as the visual... ...storage solutions while harnessing the power of cloud computing. You will be responsible for... ...such as Docker, Mesosphere DCOS, Kubernetes (k8s).NVIDIA is widely considered to be one of...CloudSeniorFull time$137k - $156k
...advanced server, storage, and networking solutions for Data Center, Cloud Computing, Enterprise IT, Hadoop/ Big Data, Hyperscale, HPC and... ...:Supermicro is seeking an experienced Senior Systems Engineer / GPU Platforms to support the bring-up, qualification, enablement, and...SeniorWorldwide$145k - $235.5k
...opportunity. Job Summary As an AI-Native Cloud FinOps Engineer , you will build... .../ML workloads. Programmatically manage GPU cluster provisioning, optimize model batching... ...Software Engineering, Cloud Engineering, SRE, or Cloud FinOps teams. ~2+ years of hands...CloudSeniorFull timeWork at officeLocal areaVisa sponsorshipWork visaShift work$184k - $287.5k
...looking for a hardworking Sr. Systems Software... ...innovative ways to make GPU accelerated applications... ...the core group working on Cloud Native technologies enabling... ...accelerators in the k8s environment.Work with engineering... ...work experience.Expert level knowledge in a systems...CloudSeniorFull timeWork experience placementRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Sr GPU Cloud K8S Expert (SRE SME). Be the first to apply!
- fulfillment expert San Jose, CA
- guest service support expert San Jose, CA
- technology expert San Jose, CA
- senior operations technician San Jose, CA
- senior cloud service delivery manager San Jose, CA
- senior it service manager San Jose, CA
- senior project engineer San Jose, CA
- senior chief engineer San Jose, CA
- sr operations manager San Jose, CA
- senior physical design engineer San Jose, CA




