Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Sr GPU Cloud K8S Expert (SRE SME) [Remote]

$180k - $260k
Full-time

Bitdeer

San Jose, CA
  • Remote job

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

To learn more, visit [ Position Overview

You run the control plane where AIOps meets tenants — where topology-aware scheduling, self-healing, and agent-driven remediation actually execute.

NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all of that lands on real customer workloads: the topology-aware scheduler places jobs on the right NVLink domain, the operator drains and reschedules around predicted faults, and the tenant boundary is enforced against a Bare-Metal-as-a-Service backend. In this role you design, deploy, and operate that control plane — and you make sure the AIOps substrate can reach in and remediate without a human on the pager.

What you'll own
  • Production Kubernetes clusters optimized for GPU workloads at scale (100–10,000 GPUs).
  • Nvidia GPU operator, device plugin, MIG configuration, and GPU time-slicing policies.
  • Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity.
  • Custom Resource Definitions (CRDs) for GPU workload lifecycle management.
  • AI framework integrations: Slurm on K8S, Ray on K8S, Kubeflow.
  • Multi-tenant isolation: namespaces, network policies, resource quotas, RBAC, pod security standards.
  • Bare-Metal as a Service (BMaaS): automated provisioning, tenant onboarding, lifecycle, reclamation.
  • Terraform providers and modules for infrastructure-as-code across GPU clusters.
  • SLIs/SLOs for cluster availability, job completion rates, and provisioning latency.
  • Incident management: runbook automation, escalation, post-incident reviews.
  • Monitoring stack: Prometheus, Grafana, Alertmanager, PagerDuty.
  • GPU node failure handling: automated detection, drain/cordon/taint, workload rescheduling.
Feed the AIOps substrate
  • The remediation-actuator and workflow engine land here — you make the control plane safe for automated action.
  • Your CRDs are the schema the platform's predictors and remediators write against.
  • Every human intervention you do this quarter becomes an autonomous workflow next quarter.
What success looks like in year 1
  • Automated drain/reschedule around predicted GPU faults, at scale, without customer impact.
  • BMaaS live for external tenants with self-service onboarding.
  • Cluster availability and job-completion SLOs published and met.
Job Requirement:
  • 5+ years in Kubernetes operations, with at least 2 years managing GPU workloads on K8S
  • Deep understanding of Nvidia GPU operator, device plugin, and GPU scheduling in K8S
  • Experience with topology-aware scheduling and GPU-specific resource management
  • Hands-on experience building multi-tenant K8S platforms with strong isolation guarantees
  • Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom)
  • Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux)
  • Strong SRE background: SLI/SLO frameworks, incident management, capacity planning
  • Experience with Prometheus, Grafana, and alerting at scale
  • Strong programming skills in Go or Python for operator/CRD development
  • AIOps aptitude — you think of the K8S control plane as an execution surface for automated remediation, not just a scheduler. You've either wired an autoscaler/remediator loop into K8S or you can design one.
  • Runbook-as-code mindset — every SRE playbook you write should be executable by the platform.
-------------------------------------------------------------------- Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.
Vacancy posted 18 days ago
Similar jobs that could be interesting for youBased on the Sr GPU Cloud K8S Expert (SRE SME) [Remote] in San Jose, CA vacancy
  •  .... Bitdeer also offers advanced cloud capabilities to customers with...  ...NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all...  ...integrations: Slurm on K8S, Ray on K8S, Kubeflow. Multi...  ...workflows (ArgoCD/Flux) ~ Strong SRE background: SLI/SLO frameworks,... 
    Cloud
    Senior
    Remote job
    Full time
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    29 days ago
  • $180k - $320k

     ...equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial...  ...stalls a $50M training run. Bitdeer is building an AI-operated GPU cloud. Storage is where AI workloads either fly or fall over: a... 
    Cloud
    Senior
    Remote job
    Full time
    Local area

    Bitdeer

    San Jose, CA
    18 days ago
  • $180k - $320k

     ...equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial...  ...from tickets into policy. NeoCloud is building an AI-operated GPU cloud spanning 4 US DCs, APAC sites, and Iceland. In this role you... 
    Cloud
    Senior
    Remote job
    Full time
    Local area

    Bitdeer

    San Jose, CA
    18 days ago
  •  ...equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial...  ...and link-failure predictors. Bitdeer is building an AI-operated GPU cloud where East-West bandwidth is the difference between a... 
    Cloud
    Remote job
    Full time
    Local area

    Bitdeer

    San Jose, CA
    a month ago
  •  .... Bitdeer also offers advanced cloud capabilities to customers with...  ...bounded contexts of the NeoCloud SRE platform — the multi-region...  ...observes, protects, and operates a GPU rental fleet across self-built...  ...SRE-tool collection plugins for K8s, Slurm, Ray, Volcano, Kueue,... 
    Cloud
    Senior
    Remote job
    Full time
    Contract work
    Local area

    Bitdeer

    San Jose, CA
    a month ago
  •  ...operations. Bitdeer also offers advanced cloud capabilities to customers with high demand...  ...is seeking a visionary and hands-on Cloud SRE Architect to lead the design, development,...  ...oversee the end-to-end architecture across CPU, GPU, RDS, storage, networking, serverless, and... 
    Cloud
    Senior
    Remote job
    Full time
    Contract work
    Local area
    Shift work

    Bitdeer

    San Jose, CA
    a month ago
  • Webex Central Operations Engineering seeks a Senior Cloud Platform Engineer / SRE to design, operate, automate, and support AWS-based cloud platforms running Kubernetes/EKS in regulated FedRAMP High / IL-5 environments. You will collaborate with Compliance and SecOps to... 
    Cloud
    Senior

    Expedite Talent Solutions

    Milpitas, CA
    5 days ago
  •  ...Senior Systems Software Engineer – GPU Performance at Scale We are looking for a dedicated engineer for the Senior Systems Software...  ...accelerated computing software stacks (CUDA). Experience with modern cloud and container‑based enterprise computing architectures, with... 
    Cloud
    Senior

    NVIDIA

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self‑driving cars...  ...computing software stacks (CUDA). Experience with modern cloud and container‑based enterprise computing architectures, with Slurm... 
    Cloud
    Senior

    NVIDIA Gruppe

    Santa Clara, CA
    3 days ago
  • $140k - $165k

     ...we drive the evolution of advancing mobile technology, empowering cloud computing, and pioneering future technologies. Our cutting-edge...  ...: Design and deploy on-prem AI infrastructure — including GPU clusters, model serving (e.g., vLLM, TGI, Triton), vector DBs (e.... 
    Cloud
    Senior
    Full time

    SK Hynix Memory Solutions America Inc.

    San Jose, CA
    more than 2 months ago
  • $152k - $241.5k

     ...Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and...  ...scripting language, preferably Python Familiar with containers, cloud provisioning and scheduling tools (Kubernetes, SLURM, Ansible,... 
    Cloud
    Senior
    Full time
    Remote work

    NVIDIA

    Santa Clara, CA
    3 days ago
  •  ...Description Job Description Developer & Infrastructure Expert Role Type: Contractor Location: Remote Job...  ...evaluate AI-powered workflows across software development, cloud infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands... 
    Cloud
    Remote job
    For contractors

    YO AI Labs

    San Jose, CA
    3 days ago
  •  ...SRE Role - SRE Location, RTP/NC and San Jose, CA Duration - Fulltime Job Description...  ...Candidate should have good knowledge in K8s Mandatory and good knowledge with K8s...  ...building CICD pipelines (preferred) Cloud platform knowledge (specifically AWS) is... 
    Cloud
    Full time

    RGH - Global Ltd

    San Jose, CA
    1 day ago
  •  ...and networking solutions for Data Center, Cloud Computing, Enterprise IT, Hadoop/ Big Data...  ...us. Job Summary Supermicro is seeking a Sr. Product Manager who can lead the development...  ...lead the development and integration of GPU server/workstation system products Develop... 
    Cloud
    Senior
    Remote work
    Worldwide

    Supermicro

    San Jose, CA
    3 days ago
  • $180k - $320k

     ...operations. Bitdeer also offers advanced cloud capabilities to customers with high demand...  ...Position Overview We are seeking a Senior GPU Systems & Fabric Engineer to serve as the...  ...host networking stacks, including RDMA, SR-IOV, RoCEv2, and InfiniBand, ensuring line... 
    Cloud
    Senior
    Remote job
    Full time
    Local area

    Bitdeer

    San Jose, CA
    29 days ago
  • $182k - $242k

     ...CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and...  ...role: CoreWeave is the top-rated AI-cloud for high-performance GPU infrastructure across AI/ML, visual effects, rendering, and real-... 
    Cloud
    Senior
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    Core Weave

    Sunnyvale, CA
    more than 2 months ago
  • $183k - $247.6k

     ...the foundation of the world’s most advanced cloud for AI training and inference — where...  ...engineers, supply chain specialists, security experts, operations managers, and other vital...  ...failure analysis, server components (e.g. CPU, GPU, SSDs, memory), BIOS, BMC, and networking... 
    Cloud
    Senior
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  • $193.5k

     ...bring-up and validation Platform Integration: - Interface with CPU/GPU vendors (Intel, AMD and Nvidia) for new platform bring-up -...  ...Services (AWS) is the world's most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped... 
    Cloud
    Senior
    Internship
    Local area
    Worldwide
    Flexible hours

    Amazon Development Center U.s. Inc.

    Cupertino, CA
    3 days ago
  • $215.18k - $358.63k

     ...congestion management for large-scale GPU cluster environments, is...  ...Familiarity with private and public cloud capabilities, including Linux,...  ...roles strongly preferred. ~ Expert proficiency with IEEE 802.3...  ...and containerization (Docker, K8s). Working knowledge of public... 
    Cloud
    Senior
    Work experience placement
    Shift work

    Keysight Technologies

    Santa Clara, CA
    4 days ago
  •  ...for an  AI Devops Infrastructure Engineer/GPU Infrastructure Engineer. Job Title:...  ...foundation in infrastructure engineering, DevOps/SRE, platform engineering, or similar...  ...Engineering, Platform Engineering, DevOps, SRE, Cloud Engineering, or AI/ML Infrastructure.... 
    Cloud
    Contract work

    Maxonic

    San Jose, CA
    7 days ago
  • $143.5k - $212.85k

     ...that enhance the productivity, reliability, and velocity of Venmo's Cloud Infrastructure and DevOps engineering teams. Job Description:...  ...Engineering, our work integrates elements of Site Reliability Engineering (SRE) and DevOps projects as well. On the DevOps front, we oversee... 
    Cloud
    Senior
    Full time
    Work at office
    Local area
    Immediate start
    Flexible hours

    Paypal

    San Jose, CA
    24 days ago
  •  ...ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises...  ...complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage... 
    Cloud
    Senior

    Mirantis

    San Jose, CA
    8 days ago
  •  ...advance your career.THE TEAM:AMD's Data Center GPU organization is transforming the industry...  ...our business success with North America cloud service providers. The role requires deep...  ...- customers, media, analysts, technical experts and senior executivesPossess a network of... 
    Cloud
    Senior
    Work experience placement

    AMD

    Santa Clara, CA
    5 hours ago
  • $84k - $112k

     ...About Supermicro: Supermicro is a Top Tier provider of advanced server, storage, and networking solutions for Data Center, Cloud Computing, Enterprise IT, Hadoop/ Big Data, Hyperscale, HPC and IoT/Embedded customers worldwide. We are the #5 fastest growing company... 
    Cloud
    Senior
    Worldwide

    Super Micro Computer

    San Jose, CA
    4 days ago
  • $110k - $178k

     ...advanced server, storage, and networking solutions for Data Center, Cloud Computing, Enterprise IT, Hadoop/ Big Data, Hyperscale, HPC and...  ...us.Job Summary:Supermicro Computer, Inc. is currently seeking a Sr. Sales Manager responsible for successfully expanding Supermicro'... 
    Cloud
    Senior
    Worldwide

    Super Micro Computer

    San Jose, CA
    17 hours ago
  •  ...Jobsbridge ! Jobsbridge, Inc . is a fast growing Silicon Valley based I.T staffing and professional services company specializing in Web, Cloud & Mobility staffing solutions. Be it core Java, full-stack Java, Web/UI designers, Big Data or Cloud or Mobility developers/... 
    Cloud
    Senior
    Full time

    Jobsbridge

    Santa Clara, CA
    more than 2 months ago
  •  ...Bitdeer also offers advanced cloud capabilities to...  ...infrastructure using Kubernetes (K8s) and Docker. Manage and...  ...resources (e.g., GPU clusters) to support high...  ...Reliability Engineering (SRE), or Cloud Infrastructure...  ...roles. Networking & OS: Expert-level knowledge of Linux... 
    Cloud
    Senior
    Remote job
    Full time
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    29 days ago
  •  ...Head of GPU Cloud About the Company Developing software foundation for next-generation accelerated computing platform for large-scale AI infrastructure. Industry Information Technology and Services Type Privately Held About the Role The Company... 
    Cloud

    Confidential

    San Jose, CA
    3 days ago
  •  .... Bitdeer also offers advanced cloud capabilities to customers with...  ...Bitdeer is building an AI-operated GPU cloud — a global fleet of self-...  ..., and operates the fleet. The SRE Platform team builds the...  ...squad — storage, network, GPU, K8S, and L1 operators — depends on.... 
    Cloud
    Remote job
    Full time
    Contract work
    Internship
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    29 days ago
  • Saviynt is seeking a Sr. HR Business Partner Manager to support AMS Sales, Marketing, Hyperscaler/Alliance, and Finance. You will align...  ..., ensuring GTM is scalable and high-performing within enterprise cloud environments. In this hybrid role, you will design org structures... 
    Cloud
    Senior
    Work at office

    Saviynt

    Milpitas, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Sr GPU Cloud K8S Expert (SRE SME) [Remote]. Be the first to apply!