Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Sr GPU Cloud K8S Expert (SRE SME)

Full-time

Bitdeer Technologies Group

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

To learn more, visit (

Position Overview

You run the control plane where AIOps meets tenants — where topology-aware scheduling, self-healing, and agent-driven remediation actually execute.

NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all of that lands on real customer workloads: the topology-aware scheduler places jobs on the right NVLink domain, the operator drains and reschedules around predicted faults, and the tenant boundary is enforced against a Bare-Metal-as-a-Service backend. In this role you design, deploy, and operate that control plane — and you make sure the AIOps substrate can reach in and remediate without a human on the pager.

What you'll own

  • Production Kubernetes clusters optimized for GPU workloads at scale (100–10,000 GPUs).
  • Nvidia GPU operator, device plugin, MIG configuration, and GPU time-slicing policies.
  • Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity.
  • Custom Resource Definitions (CRDs) for GPU workload lifecycle management.
  • AI framework integrations: Slurm on K8S, Ray on K8S, Kubeflow.
  • Multi-tenant isolation: namespaces, network policies, resource quotas, RBAC, pod security standards.
  • Bare-Metal as a Service (BMaaS): automated provisioning, tenant onboarding, lifecycle, reclamation.
  • Terraform providers and modules for infrastructure-as-code across GPU clusters.
  • SLIs/SLOs for cluster availability, job completion rates, and provisioning latency.
  • Incident management: runbook automation, escalation, post-incident reviews.
  • Monitoring stack: Prometheus, Grafana, Alertmanager, PagerDuty.
  • GPU node failure handling: automated detection, drain/cordon/taint, workload rescheduling.

Feed the AIOps substrate

  • The remediation-actuator and workflow engine land here — you make the control plane safe for automated action.
  • Your CRDs are the schema the platform's predictors and remediators write against.
  • Every human intervention you do this quarter becomes an autonomous workflow next quarter.

What success looks like in year 1

  • Automated drain/reschedule around predicted GPU faults, at scale, without customer impact.
  • BMaaS live for external tenants with self-service onboarding.
  • Cluster availability and job-completion SLOs published and met.

Job Requirement:

  • 5+ years in Kubernetes operations, with at least 2 years managing GPU workloads on K8S
  • Deep understanding of Nvidia GPU operator, device plugin, and GPU scheduling in K8S
  • Experience with topology-aware scheduling and GPU-specific resource management
  • Hands-on experience building multi-tenant K8S platforms with strong isolation guarantees
  • Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom)
  • Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux)
  • Strong SRE background: SLI/SLO frameworks, incident management, capacity planning
  • Experience with Prometheus, Grafana, and alerting at scale
  • Strong programming skills in Go or Python for operator/CRD development
  • AIOps aptitude — you think of the K8S control plane as an execution surface for automated remediation, not just a scheduler. You've either wired an autoscaler/remediator loop into K8S or you can design one.
  • Runbook-as-code mindset — every SRE playbook you write should be executable by the platform.

--------------------------------------------------------------------

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Vacancy posted 15 days ago
Similar jobs that could be interesting for youBased on the Sr GPU Cloud K8S Expert (SRE SME) in San Jose, CA vacancy
  •  ...equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial...  ...stalls a $50M training run. Bitdeer is building an AI-operated GPU cloud. Storage is where AI workloads either fly or fall over: a... 
    Cloud
    Senior
    Full time
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    15 days ago
  •  ...equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial...  ...from tickets into policy. NeoCloud is building an AI-operated GPU cloud spanning 4 US DCs, APAC sites, and Iceland. In this role you... 
    Cloud
    Senior
    Full time
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    15 days ago
  •  ...together, we’ll advance your career.THE ROLE:The Sr. Director, Pre-Si System Validation leads...  ...and pre-silicon validation for Data Center GPU platforms. This role ensures early, system-accurate insights across AI, HPC, and cloud-scale environments, accelerating software... 
    Cloud
    Senior
    Shift work

    AMD

    San Jose, CA
    4 hours ago
  • $256k - $414k

     ...GeForce NOW is the global leader in cloud gaming, dedicated to making...  ...-performance networking for GPU-based cloud infrastructure. This...  ...platform teams, hardware vendors, and SRE groups to influence technology...  ...-scale configurations using SR-IOV, Xen virtualization, or... 
    Cloud
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    2 days ago
  • $148.75k - $361k

     ...infrastructure. The ideal candidate has a strong background in DevOps/SRE practices, cloud infrastructure management, and MLOps tooling — with a passion...  ...for ML workloads across AWS and GCP, including GPU/TPU-based training and inference environmentsArchitect and improve... 
    Cloud
    Senior
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    4 days ago
  • $101k - $161k

     ...DescriptionArista Networks is an industry leader in data-driven, client-to-cloud networking for large data center, campus and routing environments...  ...our growing Arista’s CloudVision-as-a-Service (CVaaS) global SRE team. SREs at Arista combine strong software engineering... 
    Cloud
    Senior

    Arista Networks

    Santa Clara, CA
    2 days ago
  • Webex Central Operations Engineering seeks a Senior Cloud Platform Engineer / SRE to design, operate, automate, and support AWS-based cloud platforms running Kubernetes/EKS in regulated FedRAMP High / IL-5 environments. You will collaborate with Compliance and SecOps to... 
    Cloud
    Senior

    Expedite Talent Solutions

    Milpitas, CA
    2 days ago
  • $152k - $241.5k

     ...Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and...  ...a scripting language, preferably PythonFamiliar with containers, cloud provisioning and scheduling tools (Kubernetes, SLURM, Ansible,... 
    Cloud
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $208k - $327.75k

     ...worldwide. As data volumes explode and computational demands rise, GPU-accelerated storage solutions like GPUDirect Storage and cuFile...  ...management.Previous roles defining product strategies or requirements in cloud-based or HPC environments.NVIDIA is renowned globally as a top-... 
    Cloud
    Senior
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...Together, we advance your career. THE ROLE:Join AMD’s Datacenter GPU Product Application Engineering team to lead a high-performing group...  ...in customer-facing engineering roles within data center, GPU, cloud, or HPC environments. You are equally comfortable leading teams,... 
    Cloud
    Senior
    For contractors

    AMD

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars...  ...computing software stacks (CUDA).Experience with modern cloud and container-based enterprise computing architectures, with Slurm... 
    Cloud
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...our largest customers. We are seeking an expert Solutions Architect to assist customers in...  ...address critical business needs and support cloud service integration for NVIDIA technology...  ...error diagnostics.Hands-on experience with GPU systems in general including but not... 
    Cloud
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...Senior Systems Software Engineer – GPU Performance at Scale We are looking for a dedicated engineer for the Senior Systems Software...  ...accelerated computing software stacks (CUDA). Experience with modern cloud and container‑based enterprise computing architectures, with... 
    Cloud
    Senior

    NVIDIA

    Santa Clara, CA
    5 days ago
  •  ...advance your career.THE TEAM:AMD's Data Center GPU organization is transforming the industry...  ...our business success with North America cloud service providers. The role requires deep...  ...- customers, media, analysts, technical experts and senior executivesPossess a network of... 
    Cloud
    Senior
    Work experience placement

    AMD

    Santa Clara, CA
    1 day ago
  • $200k - $230k

     ...fast in many key markets such as AI, HPC, Cloud Computing, Storage, etc. To meet the market...  ...integrated and tested.Job Summary:The Sr. Manager, Solution Management collaborates...  ...infrastructuresHands-on experience with Nvidia and AMD GPU based deployments are desirableSolid... 
    Cloud
    Senior

    Super Micro Computer

    San Jose, CA
    3 days ago
  •  ...of high-performance backend networks supporting large-scale AMD GPU clusters. The engineer will own the network path from the GPU server...  ...and operating production data center networks for AI, GPU, HPC, cloud, or other large-scale distributed computing environments.... 
    Cloud
    Senior

    AMD

    San Jose, CA
    5 days ago
  • $155.8k - $224.2k

     ...scarcity are rapidly becoming the new norm.  Sr. Staff Cloud EngineerWe are looking for a Sr. Staff...  ...and incident response.Drive adoption of SRE practices including service level objectives...  ....Experience with AI/ML infrastructure, GPU workloads, data platforms, or large-scale... 
    Cloud
    Senior
    Full time
    Worldwide

    Bloom Energy

    San Jose, CA
    1 day ago
  • $98.9k - $228.7k

     ...platforms, and vendors — including media servers, CDN providers, cloud-native services, and edge networking.Ensuring consistent standards...  ...rollout strategies across teams.Acting as the primary SRE partner for multiple engineering teams building real-time features... 
    Cloud
    Senior
    Full time
    Work at office
    Remote work
    Flexible hours

    Zoom

    San Jose, CA
    4 hours ago
  •  ...Description Job Description Developer & Infrastructure Expert Role Type: Contractor Location: Remote Job...  ...evaluate AI-powered workflows across software development, cloud infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands... 
    Cloud
    Remote job
    For contractors

    YO AI Labs

    San Jose, CA
    3 days ago
  • $168k - $258.75k

     ...potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars...  ...to make sophisticated trade-offsAbility to work closely with cloud partnersWays to stand out from the crowd: Strong background in... 
    Cloud
    Full time

    Nvidia

    Santa Clara, CA
    4 hours ago
  •  .... Bitdeer also offers advanced cloud capabilities to customers with...  ...Bitdeer is building an AI-operated GPU cloud — a global fleet of self-...  ..., and operates the fleet. The SRE Platform team builds the...  ...squad — storage, network, GPU, K8S, and L1 operators — depends on.... 
    Cloud
    Full time
    Contract work
    Internship
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    15 days ago
  • $140k - $165k

     ...we drive the evolution of advancing mobile technology, empowering cloud computing, and pioneering future technologies. Our cutting-edge...  ...: Design and deploy on-prem AI infrastructure — including GPU clusters, model serving (e.g., vLLM, TGI, Triton), vector DBs (e.... 
    Cloud
    Senior
    Full time

    SK Hynix Memory Solutions America Inc.

    San Jose, CA
    more than 2 months ago
  • $167.7k - $245.2k

     ...control.As a Senior Site Reliability Engineer (SRE), you will build, operate, and...  ...the operational backbone supporting both cloud and air-gapped customer deployments, develop...  ...to that our worldwide network of doers and experts, and you’ll see that the opportunities to... 
    Cloud
    Senior
    Full time
    Temporary work
    Local area
    Flexible hours
    2 days per week

    CISCO Systems

    San Jose, CA
    1 day ago
  • $7,200 per month

     ...re UNSTOPPABLE for our employees! Job Overview Experience Experts are our most capable experts with a passion for delivering best-...  ...training to keep current and develop knowledge and skills to act as SME in all things T-Mobile. Consistently leverages digital tools... 
    Senior
    Hourly pay
    Full time
    Temporary work
    Part time
    Work experience placement
    Local area
    Flexible hours

    T-Mobile

    San Jose, CA
    5 days ago
  • $224k - $356.5k

    We are now looking for a GPU System Performance Architect:The NVIDIA Architecture group is looking for extraordinary computer architects...  ...architectures to extend the state of the art in GPU-accelerated cloud computing.You'll analyze trade-offs in system performance, cost... 
    Cloud
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    1 day ago
  • $155k - $185k

     ...advanced server, storage, and networking solutions for Data Center, Cloud Computing, Enterprise IT, Hadoop/ Big Data, Hyperscale, HPC and...  ...: Be at the forefront of developing and integrating advanced GPU systems and rack-scale solutions that power cloud providers, hyperscalers... 
    Cloud
    Senior
    Worldwide

    Super Micro Computer

    San Jose, CA
    4 days ago
  • $168k - $270.25k

     ...Performance Computing, and Visualization. The GPU, our invention, serves as the visual...  ...storage solutions while harnessing the power of cloud computing. You will be responsible for...  ...such as Docker, Mesosphere DCOS, Kubernetes (k8s).NVIDIA is widely considered to be one of... 
    Cloud
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $137k - $156k

     ...advanced server, storage, and networking solutions for Data Center, Cloud Computing, Enterprise IT, Hadoop/ Big Data, Hyperscale, HPC and...  ...:Supermicro is seeking an experienced Senior Systems Engineer / GPU Platforms to support the bring-up, qualification, enablement, and... 
    Senior
    Worldwide

    Super Micro Computer

    San Jose, CA
    4 days ago
  • $145k - $235.5k

     ...opportunity. Job Summary As an AI-Native Cloud FinOps Engineer , you will build...  .../ML workloads. Programmatically manage GPU cluster provisioning, optimize model batching...  ...Software Engineering, Cloud Engineering, SRE, or Cloud FinOps teams. ~2+ years of hands... 
    Cloud
    Senior
    Full time
    Work at office
    Local area
    Visa sponsorship
    Work visa
    Shift work

    Palo Alto Networks

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...looking for a hardworking Sr. Systems Software...  ...innovative ways to make GPU accelerated applications...  ...the core group working on Cloud Native technologies enabling...  ...accelerators in the k8s environment.Work with engineering...  ...work experience.Expert level knowledge in a systems... 
    Cloud
    Senior
    Full time
    Work experience placement
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Sr GPU Cloud K8S Expert (SRE SME). Be the first to apply!