Staff Slurm Cluster & HPC Engineer
Bitdeer Technologies Group
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.
Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.
Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
To learn more, visit (Position Overview
- We are seeking a Staff Slurm Cluster & HPC Scheduling Engineer to own Slurm as a first-class, productized scheduling layer across that fleet. This person is the single technical owner of Slurm cluster architecture, multi-tenant scheduling policy, and cluster reliability on both bare-metal and VM-based GPU nodes, and will lead our adoption of the Slinky operator stack (slurm-operator, and slurm-bridge where it fits) so that Slurm and Kubernetes workloads can share the same GPU pool. The role is deeply hands-on, customer-facing during onboarding and escalations, and sets the engineering standard the rest of the platform team builds on.
Key Responsibilities
- Slurm cluster architecture and lifecycle — Design, deploy, and operate production Slurm clusters on bare metal and VMs: slurmctld/slurmdbd high availability, slurmrestd, configless slurmd, SACK/MUNGE and JWT authentication, and rolling version upgrades on live clusters without losing running jobs.
- Topology-aware scheduling for GPU fabrics — Model the physical fabric in topology.conf — topology/tree for rail-optimized InfiniBand/RoCE designs and topology/block for NVLink domains such as GB200/GB300 NVL72 — and prove placement quality with NCCL bandwidth and multi-node training validation rather than assumption.
- Multi-tenant scheduling policy — Own the account/association tree, partitions, QOS, fairshare, preemption, reservations, and per-tenant TRES limits. Enforce fail-closed defaults: an unresolved tenant identity or an empty entitlement set must deny, never degrade into unrestricted access.
- Slinky on Kubernetes — Lead implementation of the Slinky slurm-operator, including its NodeSet, LoginSet, Accounting, RestAPI, and Token custom resources, cert-manager and Helm-based delivery, shared parallel-storage mounts, and login pods running sackd/sshd. Evaluate and pilot slurm-bridge for co-scheduling Kubernetes Pods, PodGroups, Jobs, JobSets, and LeaderWorkerSets through the Slurm scheduler, and document its constraints — notably exclusive whole-node allocation — before any customer exposure.
- Elastic capacity between Slurm and Kubernetes — Use Slurm cloud and power-save mechanisms (ResumeProgram/SuspendProgram, SuspendTime, ResumeTimeout) together with fleet automation to shift GPU nodes between batch training queues and Kubernetes inference capacity as demand moves.
- Container and job runtime — Operate Pyxis/Enroot and OCI/containerd job paths with correct gres.conf, cgroup v2 device constraints, and CUDA_VISIBLE_DEVICES behavior; support MPI/PMIx, module/Spack environments, and customer-supplied images.
- Cluster health and reliability engineering — Build the passive and active health-check system expected of a top-tier GPU cloud: prolog/epilog checks, LBNL NHC or equivalent, DCGM diagnostics, and detection of XID/SXID errors, ECC faults, PCIe errors, GPUs falling off the bus, IB/RoCE link flaps, and NCCL stalls — with automatic drain and job requeue. Own burn-in and acceptance testing for every new rack before it carries paid work.
- Automation and infrastructure as code — Deliver clusters through Terraform/Ansible, golden images, and bare-metal provisioning (PXE, Redfish, IPMI) so that a cluster build is reproducible, reviewable, and auditable rather than hand-tuned.
- Observability, accounting, and billing integration — Instrument queue wait time, allocation efficiency, GPU utilization, and job failure taxonomy through a Slurm exporter into Prometheus/Grafana; configure AccountingStorageTRES and TRESBillingWeights, and reconcile sacct/sreport GPU-hours against the platform's metering and invoicing pipeline.
- Technical leadership and customer engagement — Write runbooks and tenant-facing documentation, onboard and support enterprise customers, act as escalation point for cluster incidents, and mentor platform engineers on Slurm and HPC scheduling practice.
Qualifications
- 8+ years in HPC, systems, or cloud infrastructure engineering, including 4+ years operating production Slurm clusters at 100+ GPU-node scale with real users and service-level commitments.
- Deep hands-on Slurm expertise: slurm.conf, gres.conf, topology.conf, cgroup.conf, partitions/QOS/fairshare/preemption/reservations, slurmdbd accounting, slurmrestd, MUNGE/SACK and JWT authentication, and version upgrades performed on live clusters.
- Strong GPU and fabric fundamentals: NVIDIA drivers and Fabric Manager, DCGM, MIG, InfiniBand/RoCEv2 (subnet manager/UFM, rail-optimized topology), GPUDirect RDMA, and practical NCCL tuning and failure diagnosis.
- Production Kubernetes experience and working knowledge of the operator/CRD pattern, plus hands-on exposure to at least one Slurm-on-Kubernetes stack — Slinky slurm-operator or slurm-bridge, CoreWeave SUNK, or Nebius Soperator — with an informed view of the tradeoffs between them.
- Experience delivering both bare-metal and virtualized compute: bare-metal provisioning and firmware/BIOS lifecycle management, hypervisor or VM-based clusters (KVM/QEMU or a public-cloud equivalent), and Terraform/Ansible-driven automation.
- Working knowledge of parallel and shared storage for AI workloads — Lustre, GPFS/Spectrum Scale, WEKA, VAST, or NFS — and of how storage behavior shapes job performance and failure modes.
- Proficient in Python and Bash for cluster automation; Go experience is a plus for integrating with Bitdeer AI's platform control plane and with Slurm/Slinky REST client code.
- Multi-tenant security discipline: derives tenant scope from a verified identity rather than client-supplied fields, designs authorization to fail closed, and treats isolation across accounts, namespaces, storage, and networks as a hard requirement.
- Clear written and verbal communication in English, with the maturity to work directly with enterprise customers and to translate scheduling and reliability tradeoffs for product, sales, and executive stakeholders
--------------------------------------------------------------------
Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.
$145.92k - $209.24k
...and international. Job ID: 1750The Role: We're looking for an HPC Cluster Engineer to join our Infrastructure Team. Our mission is to build and... ...management.Familiarity with HPC workload schedulers such as Slurm, PBS, or Grid Engine.Comfort with Git or other version...SuggestedPermanent employmentContract workWork at officeRemote work- ...as a Senior High Performance Computing (HPC) Engineer for Classified Computing to lead the design... ...experience in HPC architecture, cluster management, and parallel computing, with... ...including cluster management tools (e.g., SLURM, PBS, Moab).Linux system administration...SuggestedWork at officeLocal areaRemote workRelocation packageFlexible hours
$224k - $356.5k
...on the world.We are seeking a Senior HPC & Quantum Systems Engineer to help architect, deploy, and operate... ...combining large-scale NVIDIA GPU clusters with physical quantum processors (neutral... ...such as CUDA-Q, cuQuantum, NVQlink, Slurm, and related toolchains.HPC Systems &...SuggestedFull timeWork at officeRemote work$110.3k - $155.7k
...Performance Computing Engineer to plan, implement, and... ...involves deep expertise in HPC architectures, parallel... ...medium-scale HPC clusters and associated storage... ...optimization using tools like SLURM. Administer parallel... ...Provide mentorship to junior staff and knowledge sharing...SuggestedPermanent employmentRemote workVisa sponsorshipRelocation package- ...defense and research programs, the full-time Senior HPC Systems Engineer will build and operate production Slurm clusters, manage hybrid federation of customer-owned... ...provisioning, networking, and fault coordination with site staff or vendors Required qualifications 10+ years of...SuggestedFull timeRemote work
- ...their own on-premises clusters, Government and commercial... .... We are a small engineering company, so engineers here... ...Works is hiring a Senior HPC Systems Engineer to... ...covers GPU node bring-up, Slurm configuration, fabric and... ...coordination with site staff or vendors. GPU and...Temporary workRemote work
- ...moves the world forward.THE ROLEGlobal Cluster Engineering (GCE) at AMD has a unique opportunity for... ...experts, marketing, and technical support staff.Understand and promote usability and... ...Familiarity with High Performance Computing (HPC) and cluster networks.Experience with...Remote workShift work
$165k - $242k
...'ll do:CoreWeave is seeking a highly skilled and motivated HPC Performance Engineer to join our HAVOCK Team, reporting into the Manager of Systems... ...bare-metal systems from POST through joining a Kubernetes cluster. The team’s primary responsibilities include maintaining a...Permanent employmentFull timeTemporary workCasual workWork at officeRemote workFlexible hours$15k
...lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale... ...will provide a world-class HPC platform for researchers to... ...to our technical staff. You will leverage IaC, Automation... .../batch compute frameworks (Slurm, Kueue, AWS/GCP Batch) and/...Work at officeLocal areaRemote work$260k - $290k
About the role As a Customer Support Engineer at Together AI, you will serve as the named technical... ...providers, ensuring SLA compliance and cluster availability Act as project manager for... ...architecture for large‑scale AI or HPC infrastructure Deep expertise in GPU infrastructure...Full timeRemote work$250k
...company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments powering... ...frameworks for GPU compute clusters Collaborate with ML, data, and... ...Support and optimize Slurm-based GPU cluster environments...Full timeRemote work- ...experienced AI Infrastructure Engineer to architect, deploy, optimize, and operate large-scale GPU clusters supporting state-of-the-art AI... ...~CUDA ~UCX ~MPI ~Slurm ~Pyxis/Enroot ~Optimize... ...production GPU clusters for AI or HPC workloads. ~Demonstrated...Full time
- HPC Systems Engineer | NYC | Stealth Quant Fund One of the most low-key, high-firepower quant trading firms in the game is hiring. No brand recognition... ...trading, and enterprise environments. What you’ll touch: SLURM job scheduling + workload tuning at scale InfiniBand + high...Remote workFlexible hours
$255k - $340k
...from home day is currently Tuesday.Hardware Engineering at Lambda is responsible for building and... ...system integration validation for new HPC AI/ML, general purpose compute, storage,... ...Hardware Engineer, Data Center Engineering, and Cluster Network Design to ensure new platforms...Work at officeLocal areaWork from homeFlexible hours$156.86k - $191.72k
...seeking a System Infrastructure / Platform Engineer to help build and manage HPC systems and Linux-based infrastructure. NERSC... ...cutting-edge technologies such as CPU/GPU clusters, parallel storage, high-speed networking, Slurm, and Kubernetes, balancing innovation with...Permanent employmentFull timeRemote workFlexible hours$150k - $160k
...biomedical science, software engineering, and program management, we focus... ...High-Performance Computing (HPC) Systems Engineer to join... ...configure, and maintain scalable HPC clusters for optimal performance.... ...and job schedulers (e.g., Slurm) for improved interactivity....Local areaRemote workFlexible hours$184k - $287.5k
NVIDIA Math Libraries team is looking for a senior engineer to join our development efforts in the area of kernel generation for AI and HPC, specifically targeting matrix operations, JITing and fusions. Around the world, leading commercial and academic organizations are...Full timeRemote work$152k - $241.5k
...inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.We are looking for a HPC Performance Engineer in our NVHPC compilers & tools group. Our performance engineers analyze High Performance Computing (HPC) applications with...Full timeRemote work$86.8k - $165.2k
...world leader in the design, manufacture and service of aircraft engines and auxiliary power systems and has been revolutionizing modern... ...Join us and help shape the future of aerospace and defense.The GTF HPC Design team is seeking an experienced engineer to support IBRs/...Contract workTemporary workWork experience placementWork at officeRemote workFlexible hours- ...a world-renowned science and engineering institute that marshals some... ...platforms like Foreman or MaaS.Clustering: Establish, maintain, and... ...like Proxmox, Kubernetes, and Slurm to provide high-level deployment... ...high-performance computing (HPC) systems.Working knowledge of...Remote work
- ...Job Description SYSTEMS ENGINEER PRINCIPAL Advance how our customers... ...High Performance Computing (HPC) systems that generate the... ...developers, and NWS operational staff to troubleshoot issues,... ...scheduler expertise (PBS Pro/Slurm), scripting languages, performance...Work from homeFlexible hours
$25k
...Job Description Job Description Software Integration Engineer - HPC - 10+ yrs of Experience - TS/SCI w/Poly Clearance Halogen Engineering... ...Bash/Python to develop scripts Solid understanding of the Slurm resource management and job scheduling tool Experience with...Full timeContract workRemote workShift work- AMD, Inc. is seeking a PMTS Systems Design Engineer to research, design, develop, and test... ...operations, and to integrate software for GPU clusters supporting AI inferencing and training.... ...in GPU performance, RDMA networking, and HPC system design. #J-18808-Ljbffr Socket....Remote job
- ...professionals to accelerate scientific discovery and engineering advances across a broad range of subject... ...the broader High-Performance Computing (HPC) infrastructure, the division also hosts... ...bugs in conjunction with other technical staff.Work with vendors to resolve issues and...Work at officeLocal areaRelocation packageFlexible hours
- ...Description Job Description Position Summary The Systems Engineer is responsible for designing, implementing, maintaining, and... ...Virtualization Design, deploy, and administer Microsoft Hyper-V clusters. Create and manage virtual machines, virtual networking,...Work at officeRemote workWeekend workAfternoon shift
- ...leader, take a look at the exciting employment opportunities that are currently available and apply online.Job SummaryThe Mechanical Engineer - Gas and Steam Turbine provides technical engineering support for gas turbine, steam turbine, generator, and associated balance-...Full timeLocal areaRelocation
$122.8k - $184.2k
...is seeking a Guidance Navigation Control Engineer: Modeling and Simulation (M&S) Product... ...Experience running 6DOF/3DOF simulations on HPC cluster for modeling vehicle trajectory and... ...Experience using job schedulers such as SLURM or PBS.Experience with the Atlassian toolset...Full timeRemote workRelocation packageShift work$74 - $84 per hour
Job Title: CAE Engineer Position Description: Protingent Staffing has... ...Infrastructure Engineer to operate the HPC compute infrastructure and... ...support for compute clusters, application environments, license... ...knowledge of job schedulers (LSF, SLURM, Grid Engine, or equivalent),...Contract workRemote work- ...Department Summary Advanced Research Computing melds expert staff and technical infrastructure to amplify and accelerate the impact... ...high-performance computing centers. We are seeking a Senior HPC Storage Engineer to lead the design, deployment, and operation of large-scale...Work at officeLocal areaRemote workAfternoon shift
$170k - $260k
...Are you a Senior HPC Systems Engineer who is ready for a new challenge that will launch your career to the next level? Tired of being treated like a company drone? Tired of promised adventures during the hiring phase, then dropped off on a remote contract and never...Full timeContract workRemote workWork from homeRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Slurm Cluster & HPC Engineer. Be the first to apply!



