Staff Slurm Cluster & HPC Engineer [Remote]
Bitdeer Technologies Group
- Remote job
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.
Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.
Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
To learn more, visit (Position Overview
- We are seeking a Staff Slurm Cluster & HPC Scheduling Engineer to own Slurm as a first-class, productized scheduling layer across that fleet. This person is the single technical owner of Slurm cluster architecture, multi-tenant scheduling policy, and cluster reliability on both bare-metal and VM-based GPU nodes, and will lead our adoption of the Slinky operator stack (slurm-operator, and slurm-bridge where it fits) so that Slurm and Kubernetes workloads can share the same GPU pool. The role is deeply hands-on, customer-facing during onboarding and escalations, and sets the engineering standard the rest of the platform team builds on.
Key Responsibilities
- Slurm cluster architecture and lifecycle — Design, deploy, and operate production Slurm clusters on bare metal and VMs: slurmctld/slurmdbd high availability, slurmrestd, configless slurmd, SACK/MUNGE and JWT authentication, and rolling version upgrades on live clusters without losing running jobs.
- Topology-aware scheduling for GPU fabrics — Model the physical fabric in topology.conf — topology/tree for rail-optimized InfiniBand/RoCE designs and topology/block for NVLink domains such as GB200/GB300 NVL72 — and prove placement quality with NCCL bandwidth and multi-node training validation rather than assumption.
- Multi-tenant scheduling policy — Own the account/association tree, partitions, QOS, fairshare, preemption, reservations, and per-tenant TRES limits. Enforce fail-closed defaults: an unresolved tenant identity or an empty entitlement set must deny, never degrade into unrestricted access.
- Slinky on Kubernetes — Lead implementation of the Slinky slurm-operator, including its NodeSet, LoginSet, Accounting, RestAPI, and Token custom resources, cert-manager and Helm-based delivery, shared parallel-storage mounts, and login pods running sackd/sshd. Evaluate and pilot slurm-bridge for co-scheduling Kubernetes Pods, PodGroups, Jobs, JobSets, and LeaderWorkerSets through the Slurm scheduler, and document its constraints — notably exclusive whole-node allocation — before any customer exposure.
- Elastic capacity between Slurm and Kubernetes — Use Slurm cloud and power-save mechanisms (ResumeProgram/SuspendProgram, SuspendTime, ResumeTimeout) together with fleet automation to shift GPU nodes between batch training queues and Kubernetes inference capacity as demand moves.
- Container and job runtime — Operate Pyxis/Enroot and OCI/containerd job paths with correct gres.conf, cgroup v2 device constraints, and CUDA_VISIBLE_DEVICES behavior; support MPI/PMIx, module/Spack environments, and customer-supplied images.
- Cluster health and reliability engineering — Build the passive and active health-check system expected of a top-tier GPU cloud: prolog/epilog checks, LBNL NHC or equivalent, DCGM diagnostics, and detection of XID/SXID errors, ECC faults, PCIe errors, GPUs falling off the bus, IB/RoCE link flaps, and NCCL stalls — with automatic drain and job requeue. Own burn-in and acceptance testing for every new rack before it carries paid work.
- Automation and infrastructure as code — Deliver clusters through Terraform/Ansible, golden images, and bare-metal provisioning (PXE, Redfish, IPMI) so that a cluster build is reproducible, reviewable, and auditable rather than hand-tuned.
- Observability, accounting, and billing integration — Instrument queue wait time, allocation efficiency, GPU utilization, and job failure taxonomy through a Slurm exporter into Prometheus/Grafana; configure AccountingStorageTRES and TRESBillingWeights, and reconcile sacct/sreport GPU-hours against the platform's metering and invoicing pipeline.
- Technical leadership and customer engagement — Write runbooks and tenant-facing documentation, onboard and support enterprise customers, act as escalation point for cluster incidents, and mentor platform engineers on Slurm and HPC scheduling practice.
Qualifications
- 8+ years in HPC, systems, or cloud infrastructure engineering, including 4+ years operating production Slurm clusters at 100+ GPU-node scale with real users and service-level commitments.
- Deep hands-on Slurm expertise: slurm.conf, gres.conf, topology.conf, cgroup.conf, partitions/QOS/fairshare/preemption/reservations, slurmdbd accounting, slurmrestd, MUNGE/SACK and JWT authentication, and version upgrades performed on live clusters.
- Strong GPU and fabric fundamentals: NVIDIA drivers and Fabric Manager, DCGM, MIG, InfiniBand/RoCEv2 (subnet manager/UFM, rail-optimized topology), GPUDirect RDMA, and practical NCCL tuning and failure diagnosis.
- Production Kubernetes experience and working knowledge of the operator/CRD pattern, plus hands-on exposure to at least one Slurm-on-Kubernetes stack — Slinky slurm-operator or slurm-bridge, CoreWeave SUNK, or Nebius Soperator — with an informed view of the tradeoffs between them.
- Experience delivering both bare-metal and virtualized compute: bare-metal provisioning and firmware/BIOS lifecycle management, hypervisor or VM-based clusters (KVM/QEMU or a public-cloud equivalent), and Terraform/Ansible-driven automation.
- Working knowledge of parallel and shared storage for AI workloads — Lustre, GPFS/Spectrum Scale, WEKA, VAST, or NFS — and of how storage behavior shapes job performance and failure modes.
- Proficient in Python and Bash for cluster automation; Go experience is a plus for integrating with Bitdeer AI's platform control plane and with Slurm/Slinky REST client code.
- Multi-tenant security discipline: derives tenant scope from a verified identity rather than client-supplied fields, designs authorization to fail closed, and treats isolation across accounts, namespaces, storage, and networks as a hard requirement.
- Clear written and verbal communication in English, with the maturity to work directly with enterprise customers and to translate scheduling and reliability tradeoffs for product, sales, and executive stakeholders
--------------------------------------------------------------------
Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.
$152k - $241.5k
...on the world.We are seeking a highly skilled and experienced HPC Cluster Engineer to design, deploy, and operate GPU Compute Clusters for EDA... ...Experience with AI/HPC job schedulers and orchestrators, such as Slurm, LSF, PBS or K8s. Applied experience with AI/HPC workflows...SuggestedFull time$184k - $287.5k
...next-gen distributed storage services for HPC workloads, optimizing both performance... ...our researchers to run their flows on our clusters including performance analysis and... ...degree in Computer Science, Electrical Engineering or related field or equivalent experience...SuggestedFull time- Job-ID18347302Reference21-13197 Job Description: Design and implementation of high-performance compute clusters Solid knowledge on the HPC cluster systems, including scalable/robust storage, high-bandwidth inter-connects, CPU / GPU architecture, and a knowledge of cloud...Suggested
- ...THE ROLE:We are seeking an AI Systems Engineer to join our AMD IT compute platforms... ...of High-Performance Computing (HPC) infrastructure, GPU clusters, and AI workload schedulers. THE PERSON... ...infrastructure engineering for AI/HPC domain SLURM and Kubernetes management Managing...Suggested
$136.3k - $231.7k
...into R&D. Our expert teams of physicists, engineers, data scientists and problem-solvers work... ...of a world class team of physicists, HPC system designers, machine learning and application... ...Linux administration, Kubernetes/Docker, Slurm, Ray, TensorFlow, PyTorch, GPU/TPU...SuggestedMinimum wageFull timeWork experience placementFlexible hours- ...deployment of large-scale AI/ML clustered infrastructure. You will be working... ...team of multi-disciplined engineers that operates across industry verticals... ...performance tuningOrchestration models: Slurm configuration patterns, Kubernetes for HPC/AI (GPU operators, device plugins...Flexible hours
$124k - $195.5k
As an HPC Operations Engineer at NVIDIA, you will play a pivotal role in ensuring the flawless operation of our high-performance computing (HPC)... ...tasksSolid understanding of workload schedulers such as LSF, Slurm, or similar systemsStrong grasp of network computing...Full time$95k - $161.5k
...15% of sales back into R&D. Our expert teams of physicists, engineers, data scientists and problem-solvers work together with the world... ...QualificationsKey ResponsibilitiesDesign & configure HPC clusters - Support development of compute cluster architectures optimized...Minimum wageFull timeWork experience placementWorldwideFlexible hours$162.8k - $217.6k
...well as computing infrastructure to enable engineers to solve problems faster and more... ...efficiently utilize High-Performance Computing (HPC) resources, and make informed decisions... ...Experience with HPC management software (Slurm/PBS/Torque, OpenHPC/Bright, Warewulf/XCat...Local area$152k - $241.5k
..., measurable, and aligned with long-term engineering demands.What you'll be doing:Manage, scale... ...optimize job scheduling systems (LSF, Slurm, etc.) in a large-scale, multi-site environment... ...scheduling systems (LSF, Slurm, etc.) in HPC or silicon design environmentsProficiency...Full time$255k - $340k
...from home day is currently Tuesday.Hardware Engineering at Lambda is responsible for building and... ...system integration validation for new HPC AI/ML, general purpose compute, storage,... ...Hardware Engineer, Data Center Engineering, and Cluster Network Design to ensure new platforms...Work at officeLocal areaWork from homeFlexible hours$184k - $287.5k
NVIDIA Math Libraries team is looking for a senior engineer to join our development efforts in the area of kernel generation for AI and HPC, specifically targeting matrix operations, JITing and fusions. Around the world, leading commercial and academic organizations are...Full timeRemote work$184k - $287.5k
NVIDIA has become the platform upon which every new AI-powered application is built. We are seeking a Sr. HPC Performance engineer to join our team of scientists and engineers passionate about building the next generation of scientific machine learning (ML) frameworks....Full time- AMD, Inc. is seeking a PMTS Systems Design Engineer to research, design, develop, and test... ...operations, and to integrate software for GPU clusters supporting AI inferencing and training.... ...in GPU performance, RDMA networking, and HPC system design. #J-18808-Ljbffr Socket....Remote job
$165k - $242k
...HPC Performance Engineer CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology... ...bare-metal systems from POST through joining a Kubernetes cluster. The team's primary responsibilities include maintaining a...Full timeTemporary workCasual workWork at officeRemote workFlexible hours$155k - $185k
..., Cloud Computing, Enterprise IT, Hadoop/ Big Data, Hyperscale, HPC and IoT/Embedded customers worldwide. We are the #5 fastest growing... ...community. We seek talented, passionate, and committed engineers, technologists, and business leaders to join us.Job Summary:As a...Contract workImmediate startWorldwide$74 - $84 per hour
Job Title: CAE Engineer Position Description: Protingent Staffing has... ...Infrastructure Engineer to operate the HPC compute infrastructure and... ...support for compute clusters, application environments, license... ...knowledge of job schedulers (LSF, SLURM, Grid Engine, or equivalent),...Contract workRemote work- NVIDIA has become the platform upon which every new AI-powered application is built. We are seeking a Sr. HPC Performance engineer to join our team of scientists and engineers passionate about building the next generation of scientific machine learning (ML) frameworks....
$2,000 per month
...and staffed by leading engineers, Etched is redefining the... ...performance computing (HPC), building systems that... ...Optimize and manage Slurm-based job scheduling for... ...infrastructure, and compute clusters. Manage and optimize... ...all of our technical staff to contribute to both...Work at officeRelocation package- ...availability of large-scale GPU clusters. • Respond to incidents and... ...Computer Science, Computer Engineering, Software Engineering,... ...SRE, DevOps, cloud operations, HPC, or infrastructure operations... ...Preferred Qualifications • Slurm. • GPU infrastructure. •...Night shift
$189k - $301k
A leading semiconductor company in San Jose is seeking a Senior Staff Engineer specializing in High-Speed I/O Analog-Mixed Circuit design. The role requires extensive experience in analog circuit design and high-speed interfaces, with a focus on optimization and simulation...$183k - $247.6k
...powers breakthrough innovation in AI/ML and HPC workloads. If you're passionate about... ...team of software, hardware, and network engineers, supply chain specialists, security experts... ...cooperatively with other employees, supervisors, and staff; adhere to standards of excellence...Local areaFlexible hours$150.98k - $218.62k
...innovators stay Ahead of What's Possible. Learn more at and on LinkedIn and X.Staff System Architecture and Design EngineerData Center & Energy (DCE)San Jose, CAAbout the RoleAs a System Engineer, you will own the system development of high-voltage power converters — from...Permanent employmentFull timeWork at officeShift workDay shift$148.7k - $201.2k
...product line. The Nitro Team is looking for engineers with systems knowledge and experience in... ...workloads.The Nitro High Memory and HPC team owns the purpose built platform development... ...with other employees, supervisors, and staff; adhere to standards of excellence...InternshipLocal areaFlexible hours$131k - $175k
...excellence has earned us several prestigious awards, such as Best Engineering Team, Best Company for Diversity, Compensation, and Work-Life... ...platforms integrate cleanly into large-scale AI and cloud clusters.What You’ll Do Lead the design, validation, and deployment of...Remote workFlexible hours$140k - $224.25k
...include gaming, automotive, vision, HPC, datacenters and networking in... ...telemetries, scale out cluster, test plan development, track... ...a STEM (Science, Technology, Engineering, Math or Physics) field5+ years... ...in GitHub/Gitlab/Gerrit, PXE, SLURM, Stack/Kubernetes/Docker) - huge...Full time- ...looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of... ...and tensor parallelism across multi-node clusters. Analyze network-level bottlenecks using... ...in GPU performance engineering, HPC, or systems performance analysisHands-on...
$147.75k - $221.63k
...jobWe are seeking a highly motivated High Performance Computing (HPC) Engineer to design, develop, deploy, and maintain cutting-edge... ...Unix.Troubleshooting and resolving hardware problems on Linux cluster hardware and networking equipment and applying logical methods...Full timeWork experience placement$100.55k - $194.27k
Job Details:Job Description: The world is transforming - and so is Intel. Intel is a company of bold and curious inventors and problem solvers who create some of the most astounding technology advancements and experiences in the world. With a legacy of relentless innovation...Full timeInternshipLocal areaImmediate startShift work$320k
...CPUs, and a fully optimized NVIDIA AI and HPC software stack. We're searching for a... ...motivated, technical leader to drive the engineering roadmap and innovation for our rack system... ...crowd:Knowledge of large-scale cloud and cluster level deployment and management systems....Full timeShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Slurm Cluster & HPC Engineer [Remote]. Be the first to apply!
- staff engineer San Jose, CA
- assistant engineer San Jose, CA
- staff design engineer San Jose, CA
- engineering aide San Jose, CA
- senior staff engineer San Jose, CA
- senior staff systems engineer San Jose, CA
- software engineer staff San Jose, CA
- technology administrator San Jose, CA
- staff data engineer San Jose, CA
- assistant mechanical engineer San Jose, CA

