Sr GPU Cloud K8S Expert (SRE SME) [Remote]
$180k - $260kBitdeer
- Remote job
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.
Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.
Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
To learn more, visit [ Position OverviewYou run the control plane where AIOps meets tenants — where topology-aware scheduling, self-healing, and agent-driven remediation actually execute.
NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all of that lands on real customer workloads: the topology-aware scheduler places jobs on the right NVLink domain, the operator drains and reschedules around predicted faults, and the tenant boundary is enforced against a Bare-Metal-as-a-Service backend. In this role you design, deploy, and operate that control plane — and you make sure the AIOps substrate can reach in and remediate without a human on the pager.
What you'll own- Production Kubernetes clusters optimized for GPU workloads at scale (100–10,000 GPUs).
- Nvidia GPU operator, device plugin, MIG configuration, and GPU time-slicing policies.
- Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity.
- Custom Resource Definitions (CRDs) for GPU workload lifecycle management.
- AI framework integrations: Slurm on K8S, Ray on K8S, Kubeflow.
- Multi-tenant isolation: namespaces, network policies, resource quotas, RBAC, pod security standards.
- Bare-Metal as a Service (BMaaS): automated provisioning, tenant onboarding, lifecycle, reclamation.
- Terraform providers and modules for infrastructure-as-code across GPU clusters.
- SLIs/SLOs for cluster availability, job completion rates, and provisioning latency.
- Incident management: runbook automation, escalation, post-incident reviews.
- Monitoring stack: Prometheus, Grafana, Alertmanager, PagerDuty.
- GPU node failure handling: automated detection, drain/cordon/taint, workload rescheduling.
- The remediation-actuator and workflow engine land here — you make the control plane safe for automated action.
- Your CRDs are the schema the platform's predictors and remediators write against.
- Every human intervention you do this quarter becomes an autonomous workflow next quarter.
- Automated drain/reschedule around predicted GPU faults, at scale, without customer impact.
- BMaaS live for external tenants with self-service onboarding.
- Cluster availability and job-completion SLOs published and met.
- 5+ years in Kubernetes operations, with at least 2 years managing GPU workloads on K8S
- Deep understanding of Nvidia GPU operator, device plugin, and GPU scheduling in K8S
- Experience with topology-aware scheduling and GPU-specific resource management
- Hands-on experience building multi-tenant K8S platforms with strong isolation guarantees
- Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom)
- Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux)
- Strong SRE background: SLI/SLO frameworks, incident management, capacity planning
- Experience with Prometheus, Grafana, and alerting at scale
- Strong programming skills in Go or Python for operator/CRD development
- AIOps aptitude — you think of the K8S control plane as an execution surface for automated remediation, not just a scheduler. You've either wired an autoscaler/remediator loop into K8S or you can design one.
- Runbook-as-code mindset — every SRE playbook you write should be executable by the platform.
- .... Bitdeer also offers advanced cloud capabilities to customers with... ...NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all... ...integrations: Slurm on K8S, Ray on K8S, Kubeflow. Multi... ...workflows (ArgoCD/Flux) ~ Strong SRE background: SLI/SLO frameworks,...CloudSeniorRemote jobFull timeLocal area
$180k - $320k
...equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial... ...stalls a $50M training run. Bitdeer is building an AI-operated GPU cloud. Storage is where AI workloads either fly or fall over: a...CloudSeniorRemote jobFull timeLocal area$180k - $320k
...equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial... ...from tickets into policy. NeoCloud is building an AI-operated GPU cloud spanning 4 US DCs, APAC sites, and Iceland. In this role you...CloudSeniorRemote jobFull timeLocal area- ...equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial... ...and link-failure predictors. Bitdeer is building an AI-operated GPU cloud where East-West bandwidth is the difference between a...CloudRemote jobFull timeLocal area
- .... Bitdeer also offers advanced cloud capabilities to customers with... ...bounded contexts of the NeoCloud SRE platform — the multi-region... ...observes, protects, and operates a GPU rental fleet across self-built... ...SRE-tool collection plugins for K8s, Slurm, Ray, Volcano, Kueue,...CloudSeniorRemote jobFull timeContract workLocal area
- ...operations. Bitdeer also offers advanced cloud capabilities to customers with high demand... ...is seeking a visionary and hands-on Cloud SRE Architect to lead the design, development,... ...oversee the end-to-end architecture across CPU, GPU, RDS, storage, networking, serverless, and...CloudSeniorRemote jobFull timeContract workLocal areaShift work
- Webex Central Operations Engineering seeks a Senior Cloud Platform Engineer / SRE to design, operate, automate, and support AWS-based cloud platforms running Kubernetes/EKS in regulated FedRAMP High / IL-5 environments. You will collaborate with Compliance and SecOps to...CloudSenior
- ...Senior Systems Software Engineer – GPU Performance at Scale We are looking for a dedicated engineer for the Senior Systems Software... ...accelerated computing software stacks (CUDA). Experience with modern cloud and container‑based enterprise computing architectures, with...CloudSenior
$184k - $287.5k
...potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self‑driving cars... ...computing software stacks (CUDA). Experience with modern cloud and container‑based enterprise computing architectures, with Slurm...CloudSenior$140k - $165k
...we drive the evolution of advancing mobile technology, empowering cloud computing, and pioneering future technologies. Our cutting-edge... ...: Design and deploy on-prem AI infrastructure — including GPU clusters, model serving (e.g., vLLM, TGI, Triton), vector DBs (e....CloudSeniorFull time$152k - $241.5k
...Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and... ...scripting language, preferably Python Familiar with containers, cloud provisioning and scheduling tools (Kubernetes, SLURM, Ansible,...CloudSeniorFull timeRemote work- ...Description Job Description Developer & Infrastructure Expert Role Type: Contractor Location: Remote Job... ...evaluate AI-powered workflows across software development, cloud infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands...CloudRemote jobFor contractors
- ...SRE Role - SRE Location, RTP/NC and San Jose, CA Duration - Fulltime Job Description... ...Candidate should have good knowledge in K8s Mandatory and good knowledge with K8s... ...building CICD pipelines (preferred) Cloud platform knowledge (specifically AWS) is...CloudFull time
- ...and networking solutions for Data Center, Cloud Computing, Enterprise IT, Hadoop/ Big Data... ...us. Job Summary Supermicro is seeking a Sr. Product Manager who can lead the development... ...lead the development and integration of GPU server/workstation system products Develop...CloudSeniorRemote workWorldwide
$180k - $320k
...operations. Bitdeer also offers advanced cloud capabilities to customers with high demand... ...Position Overview We are seeking a Senior GPU Systems & Fabric Engineer to serve as the... ...host networking stacks, including RDMA, SR-IOV, RoCEv2, and InfiniBand, ensuring line...CloudSeniorRemote jobFull timeLocal area$182k - $242k
...CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and... ...role: CoreWeave is the top-rated AI-cloud for high-performance GPU infrastructure across AI/ML, visual effects, rendering, and real-...CloudSeniorPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$183k - $247.6k
...the foundation of the world’s most advanced cloud for AI training and inference — where... ...engineers, supply chain specialists, security experts, operations managers, and other vital... ...failure analysis, server components (e.g. CPU, GPU, SSDs, memory), BIOS, BMC, and networking...CloudSeniorLocal areaFlexible hours$193.5k
...bring-up and validation Platform Integration: - Interface with CPU/GPU vendors (Intel, AMD and Nvidia) for new platform bring-up -... ...Services (AWS) is the world's most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped...CloudSeniorInternshipLocal areaWorldwideFlexible hours$215.18k - $358.63k
...congestion management for large-scale GPU cluster environments, is... ...Familiarity with private and public cloud capabilities, including Linux,... ...roles strongly preferred. ~ Expert proficiency with IEEE 802.3... ...and containerization (Docker, K8s). Working knowledge of public...CloudSeniorWork experience placementShift work- ...for an AI Devops Infrastructure Engineer/GPU Infrastructure Engineer. Job Title:... ...foundation in infrastructure engineering, DevOps/SRE, platform engineering, or similar... ...Engineering, Platform Engineering, DevOps, SRE, Cloud Engineering, or AI/ML Infrastructure....CloudContract work
$143.5k - $212.85k
...that enhance the productivity, reliability, and velocity of Venmo's Cloud Infrastructure and DevOps engineering teams. Job Description:... ...Engineering, our work integrates elements of Site Reliability Engineering (SRE) and DevOps projects as well. On the DevOps front, we oversee...CloudSeniorFull timeWork at officeLocal areaImmediate startFlexible hours- ...ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises... ...complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage...CloudSenior
- ...advance your career.THE TEAM:AMD's Data Center GPU organization is transforming the industry... ...our business success with North America cloud service providers. The role requires deep... ...- customers, media, analysts, technical experts and senior executivesPossess a network of...CloudSeniorWork experience placement
$84k - $112k
...About Supermicro: Supermicro is a Top Tier provider of advanced server, storage, and networking solutions for Data Center, Cloud Computing, Enterprise IT, Hadoop/ Big Data, Hyperscale, HPC and IoT/Embedded customers worldwide. We are the #5 fastest growing company...CloudSeniorWorldwide$110k - $178k
...advanced server, storage, and networking solutions for Data Center, Cloud Computing, Enterprise IT, Hadoop/ Big Data, Hyperscale, HPC and... ...us.Job Summary:Supermicro Computer, Inc. is currently seeking a Sr. Sales Manager responsible for successfully expanding Supermicro'...CloudSeniorWorldwide- ...Jobsbridge ! Jobsbridge, Inc . is a fast growing Silicon Valley based I.T staffing and professional services company specializing in Web, Cloud & Mobility staffing solutions. Be it core Java, full-stack Java, Web/UI designers, Big Data or Cloud or Mobility developers/...CloudSeniorFull time
- ...Bitdeer also offers advanced cloud capabilities to... ...infrastructure using Kubernetes (K8s) and Docker. Manage and... ...resources (e.g., GPU clusters) to support high... ...Reliability Engineering (SRE), or Cloud Infrastructure... ...roles. Networking & OS: Expert-level knowledge of Linux...CloudSeniorRemote jobFull timeLocal area
- ...Head of GPU Cloud About the Company Developing software foundation for next-generation accelerated computing platform for large-scale AI infrastructure. Industry Information Technology and Services Type Privately Held About the Role The Company...Cloud
- .... Bitdeer also offers advanced cloud capabilities to customers with... ...Bitdeer is building an AI-operated GPU cloud — a global fleet of self-... ..., and operates the fleet. The SRE Platform team builds the... ...squad — storage, network, GPU, K8S, and L1 operators — depends on....CloudRemote jobFull timeContract workInternshipLocal area
- Saviynt is seeking a Sr. HR Business Partner Manager to support AMS Sales, Marketing, Hyperscaler/Alliance, and Finance. You will align... ..., ensuring GTM is scalable and high-performing within enterprise cloud environments. In this hybrid role, you will design org structures...CloudSeniorWork at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Sr GPU Cloud K8S Expert (SRE SME) [Remote]. Be the first to apply!
- guest service support expert San Jose, CA
- fulfillment expert San Jose, CA
- technology expert San Jose, CA
- senior computer engineer San Jose, CA
- senior manager customer operations San Jose, CA
- senior software engineer ruby on rails San Jose, CA
- sr finance manager San Jose, CA
- sr marketing manager San Jose, CA
- senior customer service San Jose, CA
- senior business manager San Jose, CA



