K8 Site Reliability SME [Remote]
Bitdeer Technologies Group
- Remote job
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.
Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.
Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
To learn more, visit (Position Overview
You run the control plane where AIOps meets tenants — where topology-aware scheduling, self-healing, and agent-driven remediation actually execute.
NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all of that lands on real customer workloads: the topology-aware scheduler places jobs on the right NVLink domain, the operator drains and reschedules around predicted faults, and the tenant boundary is enforced against a Bare-Metal-as-a-Service backend. In this role you design, deploy, and operate that control plane — and you make sure the AIOps substrate can reach in and remediate without a human on the pager.
What you'll own
- Production Kubernetes clusters optimized for GPU workloads at scale (100–10,000 GPUs).
- Nvidia GPU operator, device plugin, MIG configuration, and GPU time-slicing policies.
- Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity.
- Custom Resource Definitions (CRDs) for GPU workload lifecycle management.
- AI framework integrations: Slurm on K8S, Ray on K8S, Kubeflow.
- Multi-tenant isolation: namespaces, network policies, resource quotas, RBAC, pod security standards.
- Bare-Metal as a Service (BMaaS): automated provisioning, tenant onboarding, lifecycle, reclamation.
- Terraform providers and modules for infrastructure-as-code across GPU clusters.
- SLIs/SLOs for cluster availability, job completion rates, and provisioning latency.
- Incident management: runbook automation, escalation, post-incident reviews.
- Monitoring stack: Prometheus, Grafana, Alertmanager, PagerDuty.
- GPU node failure handling: automated detection, drain/cordon/taint, workload rescheduling.
Feed the AIOps substrate
- The remediation-actuator and workflow engine land here — you make the control plane safe for automated action.
- Your CRDs are the schema the platform's predictors and remediators write against.
- Every human intervention you do this quarter becomes an autonomous workflow next quarter.
What success looks like in year 1
- Automated drain/reschedule around predicted GPU faults, at scale, without customer impact.
- BMaaS live for external tenants with self-service onboarding.
- Cluster availability and job-completion SLOs published and met.
Job Requirement:
- 5+ years in Kubernetes operations, with at least 2 years managing GPU workloads on K8S
- Deep understanding of Nvidia GPU operator, device plugin, and GPU scheduling in K8S
- Experience with topology-aware scheduling and GPU-specific resource management
- Hands-on experience building multi-tenant K8S platforms with strong isolation guarantees
- Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom)
- Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux)
- Strong SRE background: SLI/SLO frameworks, incident management, capacity planning
- Experience with Prometheus, Grafana, and alerting at scale
- Strong programming skills in Go or Python for operator/CRD development
- AIOps aptitude — you think of the K8S control plane as an execution surface for automated remediation, not just a scheduler. You've either wired an autoscaler/remediator loop into K8S or you can design one.
- Runbook-as-code mindset — every SRE playbook you write should be executable by the platform.
--------------------------------------------------------------------
Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.
- Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles...WebsiteRemote jobFull timeLocal area
$168k - $270.25k
NVIDIA is looking for a Senior Site Reliability Engineer (SRE) to join its GeForce Now (GFN) team. SRE at NVIDIA ensures that our internal and... ...Kubernetes with complex and highly available VMI setup on K8's. Lead significant production improvements including change...WebsiteFull time$122.5k - $175k
...age, we invite you to bring your talents to Zscaler and help shape the future of cybersecurity.RoleWe are looking for a Staff Site Reliability Engineer to join our team. This is a hybrid role going into the San Jose, CA office 3 days a week, reporting to the Chief Architect...WebsiteFull timeWork at officeLocal area3 days per week- ...teams to make sure new features and changes are deployed quickly and safely. • Constantly improve our system performance and reliability through better tools, process and monitoring system. • Staffing an on-call rotation with HQ at Beijing to ensure our...WebsiteWorldwide
- ...sustainable and more connected world. Job OverviewA Quality & Reliability Engineer Supervisor uses their engineering skills to assist in... ...System in support of business requirements, working closely with site Quality leaders. • As a site matures (such as moving from component...WebsiteRemote work
$230k - $250k
...team: curious people who'd rather build what doesn't exist than accept how things have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on" SRE role. As our first or early SRE hire you will be building the...WebsiteNight shift$184k - $287.5k
At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability. This demanding position merges software and systems engineering efforts to guarantee flawless service operation...WebsiteFull time- A leading technology firm is in search of a Senior Wireless Network Site Reliability Engineer to manage and enhance their wireless network infrastructure. The ideal candidate has over 8 years of experience in wireless network operations and a strong background in wireless...Website
- ...unit serving Muon's developers. The role is hybrid, requiring on-site presence in the San Jose, CA office three days per week. You will hire, coach, and set technical roadmaps while ensuring reliability, scalability, and performance of production services and...WebsiteWork at office3 days per week
$262k - $364k
...automation, and evolve systems by pushing for changes that improve reliability and velocity.Practice sustainable incident response and... ...qualifications:Master's degree in Computer Science or Engineering.Site Reliability Engineering (SRE) combines software and systems engineering...Website$168k - $264.5k
...impact on the world.Are you ready to be part of something outstanding? NVIDIA's Digital Marketing Organization seeks a senior Site Reliability Engineer (SRE) to join our Santa Clara, CA team. As an SRE at NVIDIA, you will have a meaningful role in keeping our Digital Marketing...WebsiteFull time- Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work with...WebsiteFlexible hours
- ...A global technology company is seeking a Site Reliability Engineer for its San Jose office. The successful candidate will work in a hybrid environment to ensure the reliability and scalability of critical data platforms, supporting a leading mobile app. Candidates should...WebsiteWork at office
$110.5k - $152k
...ResponsibilitiesDevelops, applies, revises, maintains and/ or tests quality/ reliability standards to ensure alignment with customer expectations.... ...by law.In addition, Applied endeavors to make our careers site accessible to all users. If you would like to contact us...WebsiteFull time$138k - $183.5k
...more about our benefits. Key ResponsibilitiesEvaluates, from a reliability standpoint, the materials, properties and techniques used in production... ...by law.In addition, Applied endeavors to make our careers site accessible to all users. If you would like to contact us...WebsiteFull time- ...with software, platform, and networking teams to improve service reliability and deployment workflowsDeploy and maintain network monitoring,... ...in the on-call rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering, or a similar roleHave...WebsiteWork at officeLocal areaWork from homeFlexible hours
- ...A leading social media platform is seeking a Site Reliability Engineer to develop and run an AI/ML recommendation system. Responsibilities include designing scalable systems, monitoring performance, and maintaining security practices. Applicants should have expertise...Website
$168k - $270.25k
...to deploy and run an AI data center. We take great pride in providing excellent, comprehensive support to our customers! Sr Site Reliability Engineer in this role will significantly impact and contribute to the overall success of both external customers running their...WebsiteFull timeWorldwide- ...provisioning, upgrades, patching, and deletion.Define and implement SLOs and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or similar role, with a deep knowledge of running Linux clusters and...WebsiteWork at officeLocal areaWork from homeFlexible hours
- LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is...WebsiteFull timeWork at office2 days per week
- Eurofins USA Consumer Product Testing is seeking an on-site Operations Leader to manage daily ESL Environmental/Reliability laboratory activities in Santa Clara, ensuring on-time delivery and quality across customer projects. The role partners with engineering, sales,...Website
$122.44k - $232.19k
...the chip development flow. Mission: Define and own the pod-level reliability specifications that ensure the availability, resilience, and... ...hiring process.Work Model for this RoleThis role will require an on-site presence. * Job posting details (such as work model, location...WebsiteFull timeLocal areaImmediate startShift work- ...A leading video platform company in San Jose is looking for an Entry Level Site Reliability Engineer to ensure the reliability of their expansive video system. The role involves managing production systems, responding to incidents, and enhancing system capabilities. Ideal...Website
- ...Platform powers compute provisioning and infrastructure orchestration across our physical data centers. We are looking for a Senior Site Reliability Engineer to improve the reliability, scalability, and operational maturity of these systems as Lambda’s fleet and customer...WebsiteWork at officeLocal areaWork from homeFlexible hours
$168k - $270.25k
...phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance Computing...WebsiteFull time$271.62k - $383.46k
...ImpactAs the leader of Intel's Foundry pre-Silicon Quality and Reliability team, you will lead a high-performing team of Design and Quality... ...Skills and ExperienceExperience in managing multi-level, cross-site, and global engineering teams.Exceptional communication skills,...WebsiteFull timeWork experience placementLocal areaImmediate startShift work$207k - $300k
...design and lead the implementation of solutions to enhance the reliability of systems that support F1.Scale systems sustainably through mechanisms... ...Computer Science or a related technical field.Experience in a Site Reliability Engineering role.Experience in designing, analyzing...Website$116k - $159.5k
...join our winning team that is focused on Quality Engineering and Reliability Engineering. We are working with engineers and scientists... ...equipment failures during stress tests in the lab, or at customer sites. Our engineering judgment is in demand every day as we are striving...WebsiteFull timeWorldwide$148.99k - $232.8k
...us. Information about Agilent is available at .The Santa Clara Site Quality Engineering Services (QES) organization provides quality... ...to determine whether products meet defined quality and reliability expectations.Develop and document test plans that define objectives...WebsiteFull timeWork at officeLocal areaWorldwide$133.5k - $183.5k
...strategies, processes and resources. Leads Business Unit quality and reliability improvement projects, driving corrective and preventive actions... ...by law.In addition, Applied endeavors to make our careers site accessible to all users. If you would like to contact us...WebsiteFull timeWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to K8 Site Reliability SME [Remote]. Be the first to apply!
- website content developer San Jose, CA
- site leader San Jose, CA
- on-site clinical research associate (traveling/remote) San Jose, CA
- on site coordinator San Jose, CA
- official site San Jose, CA
- historic site San Jose, CA
- IT site lead San Jose, CA
- junior website developer San Jose, CA
- site safety San Jose, CA
- website coordinator San Jose, CA


