Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

K8 Site Reliability SME [Remote]

Full-time

Bitdeer Technologies Group

Remote
  • Remote job

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

To learn more, visit  (

Position Overview

You run the control plane where AIOps meets tenants — where topology-aware scheduling, self-healing, and agent-driven remediation actually execute.

NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all of that lands on real customer workloads: the topology-aware scheduler places jobs on the right NVLink domain, the operator drains and reschedules around predicted faults, and the tenant boundary is enforced against a Bare-Metal-as-a-Service backend. In this role you design, deploy, and operate that control plane — and you make sure the AIOps substrate can reach in and remediate without a human on the pager.

What you'll own

  • Production Kubernetes clusters optimized for GPU workloads at scale (100–10,000 GPUs).
  • Nvidia GPU operator, device plugin, MIG configuration, and GPU time-slicing policies.
  • Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity.
  • Custom Resource Definitions (CRDs) for GPU workload lifecycle management.
  • AI framework integrations: Slurm on K8S, Ray on K8S, Kubeflow.
  • Multi-tenant isolation: namespaces, network policies, resource quotas, RBAC, pod security standards.
  • Bare-Metal as a Service (BMaaS): automated provisioning, tenant onboarding, lifecycle, reclamation.
  • Terraform providers and modules for infrastructure-as-code across GPU clusters.
  • SLIs/SLOs for cluster availability, job completion rates, and provisioning latency.
  • Incident management: runbook automation, escalation, post-incident reviews.
  • Monitoring stack: Prometheus, Grafana, Alertmanager, PagerDuty.
  • GPU node failure handling: automated detection, drain/cordon/taint, workload rescheduling.

Feed the AIOps substrate

  • The remediation-actuator and workflow engine land here — you make the control plane safe for automated action.
  • Your CRDs are the schema the platform's predictors and remediators write against.
  • Every human intervention you do this quarter becomes an autonomous workflow next quarter.

What success looks like in year 1

  • Automated drain/reschedule around predicted GPU faults, at scale, without customer impact.
  • BMaaS live for external tenants with self-service onboarding.
  • Cluster availability and job-completion SLOs published and met.

Job Requirement:

  • 5+ years in Kubernetes operations, with at least 2 years managing GPU workloads on K8S
  • Deep understanding of Nvidia GPU operator, device plugin, and GPU scheduling in K8S
  • Experience with topology-aware scheduling and GPU-specific resource management
  • Hands-on experience building multi-tenant K8S platforms with strong isolation guarantees
  • Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom)
  • Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux)
  • Strong SRE background: SLI/SLO frameworks, incident management, capacity planning
  • Experience with Prometheus, Grafana, and alerting at scale
  • Strong programming skills in Go or Python for operator/CRD development
  • AIOps aptitude — you think of the K8S control plane as an execution surface for automated remediation, not just a scheduler. You've either wired an autoscaler/remediator loop into K8S or you can design one.
  • Runbook-as-code mindset — every SRE playbook you write should be executable by the platform.

--------------------------------------------------------------------

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Vacancy posted 25 days ago
Similar jobs that could be interesting for youBased on the K8 Site Reliability SME [Remote] in Remote vacancy
  •  ...CEdge has an opportunity for a SME Web Developer , located in Norfolk , VA . If you are ready to work alongside World Renowned Technology...  ...5.2 Lead architecture, design, and development of enterprise web sites and web applications supporting NAVSUP global logistics... 
    Website
    Full time
    For contractors
    Local area
    Relocation

    Cedge Inc.

    Remote
    16 hours ago
  •  ...company. And people who care—about each other, about UiPath, and about our larger purpose.Could that be you?Your missionAt UiPath’s Site Reliability team, we build the platforms and systems that the entire company depends on to deliver on our compliance and SLA promises to... 
    Website
    Work at office
    Immediate start
    Remote work

    UiPath

    Bellevue, WA
    3 days ago
  •  ...SourceFly seeks a motivated, career and customer oriented  SME Full Stack Developer - Python \ to join our team in Ashburn, VA. This...  ...Design, develop, and implement scalable, high -performance and reliable applications and microservices using Java (including frameworks... 
    Suggested
    Full time
    Work at office
    Remote work
    2 days per week
    1 day per week

    Sourcefly

    Ashburn, VA
    16 hours ago
  •  ...shape a brighter way forward. Electrical SME - Data Center Operations     What this...  ...center facilities.   You'll   ensure the reliable operation of high-voltage electrical...  ...without sponsorship. Location: On-site –Amarillo, TX If this job description resonates... 
    Website
    Daily paid
    Full time
    Flexible hours

    *US AMR-Jones Lang LaSalle Americas, Inc.

    Amarillo, TX
    2 days ago
  •  ...brighter way forward. Mechanical   SME- Critical Facilities     As a Mechanical...  ..., energy efficiency, and operational reliability across multiple facility types ....  ...without sponsorship. Location: On-site –Amarillo, TX If this job description... 
    Website
    Daily paid
    Full time
    Apprenticeship
    Work at office
    Local area

    *US AMR-Jones Lang LaSalle Americas, Inc.

    Amarillo, TX
    2 days ago
  •  ...Senior Site Reliability Engineer We're looking for an experienced Site Reliability Engineer to help scale and maintain our hosting infrastructure in AWS. You'll work closely with our engineering teams to build secure, scalable systems that keep Framer's websites running... 
    Website
    Remote work

    Framer

    United States
    4 days ago
  •  ...shape a brighter way forward. Controls SME - Data Center Operations     What this...  ...to ensure   optimal   performance and reliability of mission-critical facilities. This role...  ...without sponsorship. Location: On-site –Amarillo, TX If this job description... 
    Website
    Daily paid
    Full time
    Immediate start
    Flexible hours

    *US AMR-Jones Lang LaSalle Americas, Inc.

    Amarillo, TX
    2 days ago
  •  ...apply today!IntroductionChevron Phillips is currently seeking a Reliability Engineer - Fixed Equipment to join the Golden Triangle Polymers...  ...corporate Mechanical Integrity program at the site.Collects and analyzes reliability metrics to identify bad actors... 
    Website
    Relocation

    Chevron Phillips Chemical

    Orange, TX
    4 days ago
  •  ...achievement. We make software robots, so people donʼt have to be robots. Would you like to be part of this journey? At UiPath's Site Reliability team, we build the platforms and systems the entire company depends on to deliver on our compliance and SLA promises to... 
    Website
    Work at office
    Immediate start
    Remote work

    Socket

    Denver, CO
    3 days ago
  •  ...seeking an accomplished Embedded Software SME with expertise in Extended Reality (XR)...  ...that ensure scalability, performance, and reliability in mission‑critical environments.Integrate...  ...a multi-disciplinary team across several sites.Travel/Physical Requirements:Location:... 
    Website
    Remote work

    Elbit Systems of America

    Merrimack, NH
    2 days ago
  •  ...opening for SharePoint Online Migration Specialist / Techno-Functional SME rec 807555 This position is 6 months, with the option of...  ...migrating data to Sharepoint specifically from a Canvas LMS site Highly desired Prior experience migrating data to Sharepoint... 
    Website
    Local area
    Remote work

    FHR

    Richmond, VA
    29 days ago
  • DescriptionJob Description SummaryThe Digital Site Reliability Engineer (SRE) - GCP Cloud Adoption Engineer is responsible for facilitating the migration, adoption, and optimization of Google Cloud Platform (GCP) services within the organization.Job DescriptionSummary:... 
    Website
    Full time
    H1b
    Work at office
    Remote work
    Work from home
    Flexible hours

    The Huntington National Bank

    Columbus, OH
    1 day ago
  •  ...Site Reliability Engineer ValidaTek is building teams of Site Reliability Engineers (SRE's) to support internal and external engineering and operations of a large scale and world-wide Enterprise IT environment that covers application hosting and support, enterprise... 
    Website

    ClearanceJobs

    Washington DC
    1 day ago
  • $98.28k - $154.44k

     ...is seeking a highly experienced Maintenance & Reliability Engineer to serve as the Company's global subject matter expert (SME) and program owner for vibration analysis and...  ...and deployment of asset strategies at plant sites, identifies gaps in a site's maintenance strategies... 
    Website
    Contract work
    For contractors

    DuPont

    Wilmington, DE
    4 days ago
  • $112k - $137k

     ...rewarded.The selected colleague will work at an MUFG office or client sites four days per week and work remotely one day. A member of our...  ...: MUFG is seeking a highly motivated Certified Sr. Cloud Site Reliability Engineer to build a robust, scalable, and reliable web... 
    Website
    Full time
    Work at office
    Local area
    Remote work

    MUFG

    Tampa, FL
    2 days ago
  • $127k - $249k

     ...scale solutions that have the ability to impact our customer’s most crucial workloads. Role OverviewWe are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This role requires engineers to have a customer-first mindset to ensure... 
    Website
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Boston, MA
    16 hours ago
  •  ...About the Role: Wrike’s Backend Reliability (BRE) team is the backbone of our backend infrastructure and the guardian of our uptime. Our...  ...Standout Qualities: Background in infrastructure engineering or Site Reliability Engineering (SRE) , including infrastructure-as-... 
    Website
    Full time

    Wrike

    Remote
    27 days ago
  • $102.9k - $133.75k

     ...frameworks and web technologies. • Act as the SME for assigned technology platforms—driving...  ...for scalability, performance, and reliability. • Analyze, resolve, and document complex...  ...departments. May require travel to attend on-site meetings/events for collaboration,... 
    Website
    Full time
    Live in
    Immediate start
    Home office
    Monday to Friday
    Flexible hours

    Affinity Plus

    Remote
    16 hours ago
  • $125k - $140k

     ...character, perspective, and passion for achieving great things in the world are equally as important to us. The role The Site Reliability Engineer is a fundamental piece of the Site Reliability Engineering team. Site Reliability Engineering is accountable for the... 
    Website
    Full time
    Local area
    Remote work
    Flexible hours

    Hitachi Digital Services

    Remote
    3 days ago
  • Reliability Lead - Manufacturing (KC Onsite-Remote)Job DescriptionJoin the team behind iconic brands like Huggies, Kleenex, Cottonelle, Scott...  ..., (4) deliver accurate and timely reporting that enables site and sector leadership to act, (5) actively support the vision of... 
    Website
    Full time
    Temporary work
    Work at office
    Remote work
    Flexible hours

    Kimberly-Clark

    Atlanta, GA
    2 days ago
  •  ...infrastructure, applications, and programs supported in the environment.As a SME Computer User Support Specialist, your primary responsibility is...  .... Customer service is key to this position. You will provide on site technical assistance to computer users by answering questions to... 
    Website
    Full time
    Temporary work
    For contractors
    Work experience placement
    Remote work
    Shift work

    Empower ai

    Washington DC
    4 days ago
  •  ...surveillance, and support services. Summary The User Interface SME will be the UI team lead. Responsibilities Designs and builds web...  ..., and tools. Designs and develops user interface features, site animation, and special-effects elements. Contributes to the design... 
    Website
    Contract work
    For contractors
    Flexible hours

    Goldbelt, Inc.

    Arlington, VA
    4 days ago
  • $132.4k - $251.6k

     ...~ 1151 E Hermans Rd ~ BLDG 801 (External Site)Country: United States of AmericaTime Type...  ...responsible for ensuring our products are Safe, Reliable, Maintainable and delivered on time. Life...  ....Serve as a Subject Matter Expert (SME) in reliability disciplines including FMECA... 
    Website
    Temporary work
    Work experience placement
    Interim role
    Remote work
    Flexible hours

    Raytheon

    Tucson, AZ
    16 hours ago
  • $134.25k - $214.8k

     ...that makes our platform self-service for the teams that depend on it.You will be part of the Observability team within Axon's Site Reliability organization — a focused team responsible for Axon's metrics, logging, tracing, and alerting infrastructure across dozens of environments... 
    Website
    Work experience placement
    Work at office
    Remote work

    Axon

    Seattle, WA
    3 days ago
  • $139k - $257.55k

    The ChallengeThe Adobe Creative Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous AI workflows, and cloud-native infrastructure. Adobe Stock gives designers and businesses... 
    Website
    Full time
    Temporary work
    Local area
    Remote work
    Worldwide

    Adobe Systems

    New York, NY
    4 days ago
  •  ...Staff Site Reliability Engineer Remote We're Bolt.new by StackBlitz! We're the team that brought you WebContainers, the first-of-its-kind technology that made it possible to run Node.js right inside your browser. That breakthrough kicked off our journey in 2019,... 
    Website
    Immediate start
    Remote work

    Bolt

    United States
    5 days ago
  •  ...challenges of productionizing AI for software engineering at scale. The Role We're looking for a skilled and passionate Site Reliability Engineer to join our team. As a SRE, you'll be responsible for the reliability, observability, performance, and security of... 
    Website
    Remote work
    1 day per week

    Runloop AI, Inc

    San Francisco, CA
    1 day ago
  •  ...Mid-Level Site Reliability Engineer We're looking for a mid-level Site Reliability Engineer to help build and operate the critical cloud infrastructure behind our platform in Microsoft Azure. You'll define the observability standards (SLOs, SLIs, dashboards, alerting... 
    Website
    Remote work

    System Automation

    United States
    2 days ago
  •  ...Network Reliability EngineerHybridAt Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world...  ...Skills, Knowledge, and Experience3 years of relevant Network/Site Reliability Engineering experienceBA/BS in Computer Science or... 
    Website
    Local area

    Cloudflare Inc

    Atlanta, GA
    3 days ago
  •  ...Opportunity SourceFly seeks a motivated, career and customer oriented SME Full Stack Developer - Python to join our team in Ashburn, VA....  ...Design, develop, and implement scalable, high-performance and reliable applications and microservices using Java (including frameworks... 
    Work at office
    Remote work
    2 days per week
    1 day per week

    SourceFly LLC

    Ashburn, VA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to K8 Site Reliability SME [Remote]. Be the first to apply!