Senior Slurm Cluster & HPC Engineer
Full-time
Bitdeer
About Bitdeer:
Bitdeer is a world-leading technology company for Bitcoin mining and AI cloud.
Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers. Apart from designing industry-leading ASIC chips and manufacturing mining rigs, the Group handles complex processes involved in computing across the value chain. This includes equipment procurement, transport logistics, datacenter design and construction, equipment management, and network and facility operations. Bitdeer also offers advanced cloud capabilities to customers with a high demand for artificial intelligence. Headquartered in Singapore, Bitdeer operates globally with a diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia. What you will be responsible for:- Slurm cluster architecture and lifecycle — Design, deploy, and operate production Slurm clusters on bare metal and VMs: slurmctld/slurmdbd high availability, slurmrestd, configless slurmd, SACK/MUNGE and JWT authentication, and rolling version upgrades on live clusters without losing running jobs.
- Topology-aware scheduling for GPU fabrics — Model the physical fabric in topology.conf — topology/tree for rail-optimized InfiniBand/RoCE designs and topology/block for NVLink domains such as GB200/GB300 NVL72 — and prove placement quality with NCCL bandwidth and multi-node training validation rather than assumption.
- Multi-tenant scheduling policy — Own the account/association tree, partitions, QOS, fairshare, preemption, reservations, and per-tenant TRES limits. Enforce fail-closed defaults: an unresolved tenant identity or an empty entitlement set must deny, never degrade into unrestricted access.
- Slinky on Kubernetes — Lead implementation of the Slinky slurm-operator, including its NodeSet, LoginSet, Accounting, RestAPI, and Token custom resources, cert-manager and Helm-based delivery, shared parallel-storage mounts, and login pods running sackd/sshd. Evaluate and pilot slurm-bridge for co-scheduling Kubernetes Pods, PodGroups, Jobs, JobSets, and LeaderWorkerSets through the Slurm scheduler, and document its constraints — notably exclusive whole-node allocation — before any customer exposure.
- Elastic capacity between Slurm and Kubernetes — Use Slurm cloud and power-save mechanisms (ResumeProgram/SuspendProgram, SuspendTime, ResumeTimeout) together with fleet automation to shift GPU nodes between batch training queues and Kubernetes inference capacity as demand moves.
- Container and job runtime — Operate Pyxis/Enroot and OCI/containerd job paths with correct gres.conf, cgroup v2 device constraints, and CUDA VISIBLE DEVICES behavior; support MPI/PMIx, module/Spack environments, and customer-supplied images.
- Cluster health and reliability engineering — Build the passive and active health-check system expected of a top-tier GPU cloud: prolog/epilog checks, LBNL NHC or equivalent, DCGM diagnostics, and detection of XID/SXID errors, ECC faults, PCIe errors, GPUs falling off the bus, IB/RoCE link flaps, and NCCL stalls — with automatic drain and job requeue. Own burn-in and acceptance testing for every new rack before it carries paid work.
- Automation and infrastructure as code — Deliver clusters through Terraform/Ansible, golden images, and bare-metal provisioning (PXE, Redfish, IPMI) so that a cluster build is reproducible, reviewable, and auditable rather than hand-tuned.
- Observability, accounting, and billing integration — Instrument queue wait time, allocation efficiency, GPU utilization, and job failure taxonomy through a Slurm exporter into Prometheus/Grafana; configure AccountingStorageTRES and TRESBillingWeights, and reconcile sacct/sreport GPU-hours against the platform's metering and invoicing pipeline.
- Technical leadership and customer engagement — Write runbooks and tenant-facing documentation, onboard and support enterprise customers, act as escalation point for cluster incidents, and mentor platform engineers on Slurm and HPC scheduling practice.
- 8+ years in HPC, systems, or cloud infrastructure engineering, including 4+ years operating production Slurm clusters at 100+ GPU-node scale with real users and service-level commitments.
- Deep hands-on Slurm expertise: slurm.conf, gres.conf, topology.conf, cgroup.conf, partitions/QOS/fairshare/preemption/reservations, slurmdbd accounting, slurmrestd, MUNGE/SACK and JWT authentication, and version upgrades performed on live clusters.
- Strong GPU and fabric fundamentals: NVIDIA drivers and Fabric Manager, DCGM, MIG, InfiniBand/RoCEv2 (subnet manager/UFM, rail-optimized topology), GPUDirect RDMA, and practical NCCL tuning and failure diagnosis.
- Production Kubernetes experience and working knowledge of the operator/CRD pattern, plus hands-on exposure to at least one Slurm-on-Kubernetes stack — Slinky slurm-operator or slurm-bridge, CoreWeave SUNK, or Nebius Soperator — with an informed view of the tradeoffs between them.
- Experience delivering both bare-metal and virtualized compute: bare-metal provisioning and firmware/BIOS lifecycle management, hypervisor or VM-based clusters (KVM/QEMU or a public-cloud equivalent), and Terraform/Ansible-driven automation.
- Working knowledge of parallel and shared storage for AI workloads — Lustre, GPFS/Spectrum Scale, WEKA, VAST, or NFS — and of how storage behavior shapes job performance and failure modes.
- Proficient in Python and Bash for cluster automation; Go experience is a plus for integrating with Bitdeer AI's platform control plane and with Slurm/Slinky REST client code.
- Multi-tenant security discipline: derives tenant scope from a verified identity rather than client-supplied fields, designs authorization to fail closed, and treats isolation across accounts, namespaces, storage, and networks as a hard requirement.
- Clear written and verbal communication in English, with the maturity to work directly with enterprise customers and to translate scheduling and reliability tradeoffs for product, sales, and executive stakeholders.
- A culture that values authenticity and diversity of thoughts and backgrounds;
- An inclusive and respectable environment with open workspaces and exciting start-up spirit;
- Fast-growing company with the chance to network with industrial pioneers and enthusiasts;
- Ability to contribute directly and make an impact on the future of the digital asset industry;
- Involvement in new projects, developing processes/systems;
- Personal accountability, autonomy, fast growth, and learning opportunities;
- Attractive welfare benefits and developmental opportunities such as training and mentoring.
#LI-ST1
Vacancy posted 6 days ago
Similar jobs that could be interesting for youBased on the Senior Slurm Cluster & HPC Engineer in Singapore vacancy
- ...portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan,... ...-tenant inference workloads and maximize cluster utilization. Profile and tune kernel-... ...Collaborate with the Scheduling and Storage engineering teams to ensure topology-aware placement...SeniorFull timeLocal area
- ...portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan,... ...plane architectures using tools like Cluster API to enable fleet-wide automation. Develop... ...with AI Scheduling and Fabric engineering teams to ensure the control plane integrates...SeniorFull timeLocal area
- ...portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan... ...centers (AIDCs), covering detection engineering, incident response, host and network hardening... ...-speed networks, and large-scale GPU clusters — hands-on from writing detection rules...SeniorFull timeLocal areaWorldwide
- ...diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia. About the team: We are seeking a Senior Inference Runtime Engineer to own the performance-critical serving layer of the MaaS...SeniorFull timeLocal area
- ...diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada,... ...Chinese and English. Preferred Qualifications: Large-Scale Cluster Experience: Hands-on experience in the construction or operations...SeniorFull timeTemporary workLocal areaNight shift
- ...diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada,... ...silently degrade quality, latency, or cost. Work with runtime engineers to identify bottlenecks and verify improvements; work with SRE to...SeniorFull timeLocal area
- ...portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan,... ...experience. Production Kubernetes clusters optimized for GPU workloads at scale (100... ...The remediation-actuator and workflow engine land here — you make the control plane safe...SeniorFull timeLocal area
- ...diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada,... ...Role Overview: We are seeking a talented Design Verification Engineer to join our IC development team and help ensure the functional...SeniorFull timeLocal areaNight shift
- ...Motional is seeking a highly skilled Senior Cybersecurity Engineer to join our Defense Operations team. This operations-focused role puts you on the front lines of our security program, acting as a senior resource for security monitoring, system ownership, and the continuous...SeniorFull timeWork experience placement
- ...to challenge consensus. As a Software Engineer, Research Technology, you will... ...research tooling, distributed compute across HPC (Slurm today, evolving). You will help deliver... ...reliability of data and compute pipelines on HPC clusters. Create ad-hoc computation frameworks...Full timeInternshipFlexible hours
- ...testing using both manual and automation tools to provide quality and timely output before release to market. As Automation Test Engineer you are primarily responsible to convert the BDD tests written gherkin language to automated scripts across applicable digital...SeniorFull timeWork experience placement
- ...portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan... ...BS/MS/PhD in EE, Materials, Mechanical Engineering, or a related field. ~5+ years in 3... ...decisions set the stack. Early-team seniority with room to shape the 3D methodology and...SeniorFull timeLocal area
- ...portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan,... ...gang scheduling. Develop and manage cluster-wide admission control and sophisticated... ...scale HPC environments. Mentor junior engineers and conduct design reviews to maintain architectural...SeniorFull timeLocal area
- ...diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada,... ...optimization to reduce energy consumption. Collaborate with DFT engineers for scan chain insertion, MBIST for large SRAMs, and JTAG integration...SeniorFull timeLocal areaNight shift
- ...portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan,... ...stack to provide a unified view of the cluster. Build eBPF-based diagnostic tools to... ...s degree in Computer Science, Electrical Engineering, or a related field. ~6+ years of software...SeniorFull timeLocal area
- ...diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada,... ...IC designs. In this role, you will lead a team of verification engineers to ensure robust functional coverage and high-quality silicon through...Full timeLocal areaNight shift
- ...technology professionals. We are inviting experienced AI Harness Engineer to register their interest in current and upcoming... ...'s data, knowledge, APIs and operational systems. This is a senior individual-contributor role combining deep hands-on engineering...Senior
- ...and deploys Bitcoin mining and HPC datacenters in the United... ...own Production Kubernetes clusters optimized for GPU workloads at... ...AI framework integrations: Slurm on K8S, Ray on K8S, Kubeflow.... ...remediation-actuator and workflow engine land here — you make the control...SeniorFull timeLocal area
- ...a related field. Minimum 3-5 years of experience as a Data Engineer or Database Administrator. Proficient in database technologies... ...backup and recovery strategies. Familiarity with cluster-based setups, delayed replication, and other advanced database...SeniorFull time
- ...internet backends with the transaction, signing, and on-chain interaction problems unique to blockchain. We're looking for a Senior Java Engineer with roughly 5–10 years of backend experience: someone with strong engineering fundamentals who can independently own core...SeniorFull timeContract workWorldwide
- ...growing demands of our business. Our Systems Engineers are responsible for designing, building,... ...experience a plus ~ HT Condor, HTC clusters, Team City, Keycloak, Atlassian suite... ...Direct exposure to the decision makers and senior leaders on the business side A...Full timeWorldwideShift work
- ...diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada,... ...designs (such as but not limited to structured cabling, data center engineering, security systems, network communications, audio-visual systems...Full timeLocal area
- ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Platform Engineer Overview Network Operations Center (NOC) Senior Platform Engineer is a key position responsible for ensuring the stable...SeniorFull timeWork at officeWorldwideShift workNight shift
- ...diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada,... ...and data-center regions. How you will stand out: ~3+ years (Senior: 6+ years) in large-scale bare-metal / server-fleet operations,...Full timeLocal area
- ...diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada,... ...you will stand out: Bachelor’s degree or above in Electrical Engineering and Automation, Power System Automation, Power Distribution Engineering...Full timeFor contractorsLocal area
- ...technology professionals. We are inviting experienced AI Knowledge Engineer to register their interest in current and upcoming... ...support activities for knowledge-system components, working with senior engineers, architects, and delivery partners where required....Senior
- ...diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada,... ...fair resource sharing for customer workloads. Mentor junior engineers and drive architectural design reviews to maintain high standards...SeniorFull timeLocal area
- ...diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada,... ...self-healing workflows. Partner with runtime and performance engineers to debug incidents from public API edge to model worker. How...SeniorFull timeLocal area
- ...of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior AI Engineer Senior Software Engineer (Generative AI) - Foundry R&D, Singapore We are looking for a Senior Software Engineer (Generative...SeniorFull timeWorldwideShift work
- ...ATE Application / Field Application Engineer- Singapore Location: Singapore Experience: 5-12+ years Mobility: Open to qualified... ..., qualification, training and production ramp. This role suits senior test engineers who can combine deep technical expertise with...Relocation
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Slurm Cluster & HPC Engineer. Be the first to apply!
Related searches
- senior operations technician Singapore
- senior cloud service delivery manager Singapore
- senior director coding Singapore
- senior java full-stack developer Singapore
- senior staff engineer Singapore
- senior storage engineer Singapore
- senior staff systems engineer Singapore
- senior manager tax Singapore
- senior associate architect Singapore
- senior associate vice president Singapore

