Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Slurm Cluster & HPC Engineer

Full-time

Bitdeer Technologies Group

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

To learn more, visit (

Position Overview

  • We are seeking a Staff Slurm Cluster & HPC Scheduling Engineer to own Slurm as a first-class, productized scheduling layer across that fleet. This person is the single technical owner of Slurm cluster architecture, multi-tenant scheduling policy, and cluster reliability on both bare-metal and VM-based GPU nodes, and will lead our adoption of the Slinky operator stack (slurm-operator, and slurm-bridge where it fits) so that Slurm and Kubernetes workloads can share the same GPU pool. The role is deeply hands-on, customer-facing during onboarding and escalations, and sets the engineering standard the rest of the platform team builds on.

Key Responsibilities

  • Slurm cluster architecture and lifecycle — Design, deploy, and operate production Slurm clusters on bare metal and VMs: slurmctld/slurmdbd high availability, slurmrestd, configless slurmd, SACK/MUNGE and JWT authentication, and rolling version upgrades on live clusters without losing running jobs.
  • Topology-aware scheduling for GPU fabrics — Model the physical fabric in topology.conf — topology/tree for rail-optimized InfiniBand/RoCE designs and topology/block for NVLink domains such as GB200/GB300 NVL72 — and prove placement quality with NCCL bandwidth and multi-node training validation rather than assumption.
  • Multi-tenant scheduling policy — Own the account/association tree, partitions, QOS, fairshare, preemption, reservations, and per-tenant TRES limits. Enforce fail-closed defaults: an unresolved tenant identity or an empty entitlement set must deny, never degrade into unrestricted access.
  • Slinky on Kubernetes — Lead implementation of the Slinky slurm-operator, including its NodeSet, LoginSet, Accounting, RestAPI, and Token custom resources, cert-manager and Helm-based delivery, shared parallel-storage mounts, and login pods running sackd/sshd. Evaluate and pilot slurm-bridge for co-scheduling Kubernetes Pods, PodGroups, Jobs, JobSets, and LeaderWorkerSets through the Slurm scheduler, and document its constraints — notably exclusive whole-node allocation — before any customer exposure.
  • Elastic capacity between Slurm and Kubernetes — Use Slurm cloud and power-save mechanisms (ResumeProgram/SuspendProgram, SuspendTime, ResumeTimeout) together with fleet automation to shift GPU nodes between batch training queues and Kubernetes inference capacity as demand moves.
  • Container and job runtime — Operate Pyxis/Enroot and OCI/containerd job paths with correct gres.conf, cgroup v2 device constraints, and CUDA_VISIBLE_DEVICES behavior; support MPI/PMIx, module/Spack environments, and customer-supplied images.
  • Cluster health and reliability engineering — Build the passive and active health-check system expected of a top-tier GPU cloud: prolog/epilog checks, LBNL NHC or equivalent, DCGM diagnostics, and detection of XID/SXID errors, ECC faults, PCIe errors, GPUs falling off the bus, IB/RoCE link flaps, and NCCL stalls — with automatic drain and job requeue. Own burn-in and acceptance testing for every new rack before it carries paid work.
  • Automation and infrastructure as code — Deliver clusters through Terraform/Ansible, golden images, and bare-metal provisioning (PXE, Redfish, IPMI) so that a cluster build is reproducible, reviewable, and auditable rather than hand-tuned.
  • Observability, accounting, and billing integration — Instrument queue wait time, allocation efficiency, GPU utilization, and job failure taxonomy through a Slurm exporter into Prometheus/Grafana; configure AccountingStorageTRES and TRESBillingWeights, and reconcile sacct/sreport GPU-hours against the platform's metering and invoicing pipeline.
  • Technical leadership and customer engagement — Write runbooks and tenant-facing documentation, onboard and support enterprise customers, act as escalation point for cluster incidents, and mentor platform engineers on Slurm and HPC scheduling practice.

Qualifications

  • 8+ years in HPC, systems, or cloud infrastructure engineering, including 4+ years operating production Slurm clusters at 100+ GPU-node scale with real users and service-level commitments.
  • Deep hands-on Slurm expertise: slurm.conf, gres.conf, topology.conf, cgroup.conf, partitions/QOS/fairshare/preemption/reservations, slurmdbd accounting, slurmrestd, MUNGE/SACK and JWT authentication, and version upgrades performed on live clusters.
  • Strong GPU and fabric fundamentals: NVIDIA drivers and Fabric Manager, DCGM, MIG, InfiniBand/RoCEv2 (subnet manager/UFM, rail-optimized topology), GPUDirect RDMA, and practical NCCL tuning and failure diagnosis.
  • Production Kubernetes experience and working knowledge of the operator/CRD pattern, plus hands-on exposure to at least one Slurm-on-Kubernetes stack — Slinky slurm-operator or slurm-bridge, CoreWeave SUNK, or Nebius Soperator — with an informed view of the tradeoffs between them.
  • Experience delivering both bare-metal and virtualized compute: bare-metal provisioning and firmware/BIOS lifecycle management, hypervisor or VM-based clusters (KVM/QEMU or a public-cloud equivalent), and Terraform/Ansible-driven automation.
  • Working knowledge of parallel and shared storage for AI workloads — Lustre, GPFS/Spectrum Scale, WEKA, VAST, or NFS — and of how storage behavior shapes job performance and failure modes.
  • Proficient in Python and Bash for cluster automation; Go experience is a plus for integrating with Bitdeer AI's platform control plane and with Slurm/Slinky REST client code.
  • Multi-tenant security discipline: derives tenant scope from a verified identity rather than client-supplied fields, designs authorization to fail closed, and treats isolation across accounts, namespaces, storage, and networks as a hard requirement.
  • Clear written and verbal communication in English, with the maturity to work directly with enterprise customers and to translate scheduling and reliability tradeoffs for product, sales, and executive stakeholders

--------------------------------------------------------------------

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Staff Slurm Cluster & HPC Engineer in Remote vacancy
  • $145.92k - $209.24k

     ...and international. Job ID: 1750The Role: We're looking for an HPC Cluster Engineer to join our Infrastructure Team. Our mission is to build and...  ...management.Familiarity with HPC workload schedulers such as Slurm, PBS, or Grid Engine.Comfort with Git or other version... 
    Suggested
    Permanent employment
    Contract work
    Work at office
    Remote work

    IONQ

    College Park, MD
    2 days ago
  •  ...as a Senior High Performance Computing (HPC) Engineer for Classified Computing to lead the design...  ...experience in HPC architecture, cluster management, and parallel computing, with...  ...including cluster management tools (e.g., SLURM, PBS, Moab).Linux system administration... 
    Suggested
    Work at office
    Local area
    Remote work
    Relocation package
    Flexible hours

    Oak Ridge National Laboratory

    Oak Ridge, TN
    4 days ago
  • $224k - $356.5k

     ...on the world.We are seeking a Senior HPC & Quantum Systems Engineer to help architect, deploy, and operate...  ...combining large-scale NVIDIA GPU clusters with physical quantum processors (neutral...  ...such as CUDA-Q, cuQuantum, NVQlink, Slurm, and related toolchains.HPC Systems &... 
    Suggested
    Full time
    Work at office
    Remote work

    Nvidia

    Westford, MA
    1 day ago
  • $110.3k - $155.7k

     ...Performance Computing Engineer to plan, implement, and...  ...involves deep expertise in HPC architectures, parallel...  ...medium-scale HPC clusters and associated storage...  ...optimization using tools like SLURM. Administer parallel...  ...Provide mentorship to junior staff and knowledge sharing... 
    Suggested
    Permanent employment
    Remote work
    Visa sponsorship
    Relocation package

    Federal Reserve Bank of Kansas City

    Kansas City, MO
    5 days ago
  •  ...defense and research programs, the full-time Senior HPC Systems Engineer will build and operate production Slurm clusters, manage hybrid federation of customer-owned...  ...provisioning, networking, and fault coordination with site staff or vendors Required qualifications 10+ years of... 
    Suggested
    Full time
    Remote work

    Virtual Vocations Inc

    United States
    2 days ago
  •  ...their own on-premises clusters, Government and commercial...  .... We are a small engineering company, so engineers here...  ...Works is hiring a Senior HPC Systems Engineer to...  ...covers GPU node bring-up, Slurm configuration, fabric and...  ...coordination with site staff or vendors. GPU and... 
    Temporary work
    Remote work

    Parallel Works

    United States
    3 days ago
  •  ...moves the world forward.THE ROLEGlobal Cluster Engineering (GCE) at AMD has a unique opportunity for...  ...experts, marketing, and technical support staff.Understand and promote usability and...  ...Familiarity with High Performance Computing (HPC) and cluster networks.Experience with... 
    Remote work
    Shift work

    AMD

    Texas
    3 days ago
  • $165k - $242k

     ...'ll do:CoreWeave is seeking a highly skilled and motivated HPC Performance Engineer to join our HAVOCK Team, reporting into the Manager of Systems...  ...bare-metal systems from POST through joining a Kubernetes cluster. The team’s primary responsibilities include maintaining a... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours

    CoreWeave

    New York, NY
    1 day ago
  • $15k

     ...lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale...  ...will provide a world-class HPC platform for researchers to...  ...to our technical staff. You will leverage IaC, Automation...  .../batch compute frameworks (Slurm, Kueue, AWS/GCP Batch) and/... 
    Work at office
    Local area
    Remote work

    The Voleon Group

    Berkeley, CA
    4 days ago
  • $260k - $290k

    About the role As a Customer Support Engineer at Together AI, you will serve as the named technical...  ...providers, ensuring SLA compliance and cluster availability Act as project manager for...  ...architecture for large‑scale AI or HPC infrastructure Deep expertise in GPU infrastructure... 
    Full time
    Remote work

    Together AI

    San Francisco, CA
    5 days ago
  • $250k

     ...company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments powering...  ...frameworks for GPU compute clusters Collaborate with ML, data, and...  ...Support and optimize Slurm-based GPU cluster environments... 
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  •  ...experienced AI Infrastructure Engineer to architect, deploy, optimize, and operate large-scale GPU clusters supporting state-of-the-art AI...  ...~CUDA ~UCX ~MPI ~Slurm ~Pyxis/Enroot ~Optimize...  ...production GPU clusters for AI or HPC workloads. ~Demonstrated... 
    Full time

    STN Inc

    Remote
    10 days ago
  • HPC Systems Engineer | NYC | Stealth Quant Fund One of the most low-key, high-firepower quant trading firms in the game is hiring. No brand recognition...  ...trading, and enterprise environments. What you’ll touch: SLURM job scheduling + workload tuning at scale InfiniBand + high... 
    Remote work
    Flexible hours

    Hunter Bond

    New York, NY
    2 days ago
  • $255k - $340k

     ...from home day is currently Tuesday.Hardware Engineering at Lambda is responsible for building and...  ...system integration validation for new HPC AI/ML, general purpose compute, storage,...  ...Hardware Engineer, Data Center Engineering, and Cluster Network Design to ensure new platforms... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  • $156.86k - $191.72k

     ...seeking a System Infrastructure / Platform Engineer to help build and manage HPC systems and Linux-based infrastructure. NERSC...  ...cutting-edge technologies such as CPU/GPU clusters, parallel storage, high-speed networking, Slurm, and Kubernetes, balancing innovation with... 
    Permanent employment
    Full time
    Remote work
    Flexible hours

    Berkeley Lab

    Berkeley, CA
    3 days ago
  • $150k - $160k

     ...biomedical science, software engineering, and program management, we focus...  ...High-Performance Computing (HPC) Systems Engineer to join...  ...configure, and maintain scalable HPC clusters for optimal performance....  ...and job schedulers (e.g., Slurm) for improved interactivity.... 
    Local area
    Remote work
    Flexible hours

    Axle

    Rockville, MD
    more than 2 months ago
  • $184k - $287.5k

    NVIDIA Math Libraries team is looking for a senior engineer to join our development efforts in the area of kernel generation for AI and HPC, specifically targeting matrix operations, JITing and fusions. Around the world, leading commercial and academic organizations are... 
    Full time
    Remote work

    Nvidia

    Pennsylvania
    1 day ago
  • $152k - $241.5k

     ...inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.We are looking for a HPC Performance Engineer in our NVHPC compilers & tools group. Our performance engineers analyze High Performance Computing (HPC) applications with... 
    Full time
    Remote work

    Nvidia

    Texas
    10 hours ago
  • $86.8k - $165.2k

     ...world leader in the design, manufacture and service of aircraft engines and auxiliary power systems and has been revolutionizing modern...  ...Join us and help shape the future of aerospace and defense.The GTF HPC Design team is seeking an experienced engineer to support IBRs/... 
    Contract work
    Temporary work
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Raytheon

    Middletown, RI
    4 days ago
  •  ...a world-renowned science and engineering institute that marshals some...  ...platforms like Foreman or MaaS.Clustering: Establish, maintain, and...  ...like Proxmox, Kubernetes, and Slurm to provide high-level deployment...  ...high-performance computing (HPC) systems.Working knowledge of... 
    Remote work

    California Institute of Technology

    Pasadena, CA
    1 day ago
  •  ...Job Description SYSTEMS ENGINEER PRINCIPAL Advance how our customers...  ...High Performance Computing (HPC) systems that generate the...  ...developers, and NWS operational staff to troubleshoot issues,...  ...scheduler expertise (PBS Pro/Slurm), scripting languages, performance... 
    Work from home
    Flexible hours

    General Dynamics Information Technology

    Remote
    24 days ago
  • $25k

     ...Job Description Job Description Software Integration Engineer - HPC - 10+ yrs of Experience - TS/SCI w/Poly Clearance Halogen Engineering...  ...Bash/Python to develop scripts   Solid understanding of the Slurm resource management and job scheduling tool Experience with... 
    Full time
    Contract work
    Remote work
    Shift work

    Halogen Engineering Group, Inc

    Annapolis Junction, MD
    3 days ago
  • AMD, Inc. is seeking a PMTS Systems Design Engineer to research, design, develop, and test...  ...operations, and to integrate software for GPU clusters supporting AI inferencing and training....  ...in GPU performance, RDMA networking, and HPC system design. #J-18808-Ljbffr Socket.... 
    Remote job

    Socket.dev

    Santa Clara, CA
    2 days ago
  •  ...professionals to accelerate scientific discovery and engineering advances across a broad range of subject...  ...the broader High-Performance Computing (HPC) infrastructure, the division also hosts...  ...bugs in conjunction with other technical staff.Work with vendors to resolve issues and... 
    Work at office
    Local area
    Relocation package
    Flexible hours

    Oak Ridge National Laboratory

    Oak Ridge, TN
    1 day ago
  •  ...Description Job Description Position Summary The Systems Engineer is responsible for designing, implementing, maintaining, and...  ...Virtualization Design, deploy, and administer Microsoft Hyper-V clusters. Create and manage virtual machines, virtual networking,... 
    Work at office
    Remote work
    Weekend work
    Afternoon shift

    Spy Ego Media

    Havre, MT
    4 days ago
  •  ...leader, take a look at the exciting employment opportunities that are currently available and apply online.Job SummaryThe Mechanical Engineer - Gas and Steam Turbine provides technical engineering support for gas turbine, steam turbine, generator, and associated balance-... 
    Full time
    Local area
    Relocation

    TXU Energy

    Pennsylvania
    1 day ago
  • $122.8k - $184.2k

     ...is seeking a Guidance Navigation Control Engineer: Modeling and Simulation (M&S) Product...  ...Experience running 6DOF/3DOF simulations on HPC cluster for modeling vehicle trajectory and...  ...Experience using job schedulers such as SLURM or PBS.Experience with the Atlassian toolset... 
    Full time
    Remote work
    Relocation package
    Shift work

    Northrop Grumman

    Chandler, AZ
    3 days ago
  • $74 - $84 per hour

    Job Title: CAE Engineer Position Description: Protingent Staffing has...  ...Infrastructure Engineer to operate the HPC compute infrastructure and...  ...support for compute clusters, application environments, license...  ...knowledge of job schedulers (LSF, SLURM, Grid Engine, or equivalent),... 
    Contract work
    Remote work

    Protingent

    San Jose, CA
    3 days ago
  •  ...Department Summary Advanced Research Computing melds expert staff and technical infrastructure to amplify and accelerate the impact...  ...high-performance computing centers. We are seeking a Senior HPC Storage Engineer to lead the design, deployment, and operation of large-scale... 
    Work at office
    Local area
    Remote work
    Afternoon shift

    The Regents of the University of California on behalf of the...

    Los Angeles, CA
    10 hours ago
  • $170k - $260k

     ...Are you a Senior HPC Systems Engineer who is ready for a new challenge that will launch your career to the next level? Tired of being treated like a company drone? Tired of promised adventures during the hiring phase, then dropped off on a remote contract and never... 
    Full time
    Contract work
    Remote work
    Work from home
    Relocation package

    GliaCell Technologies LLC

    Annapolis Junction, MD
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Slurm Cluster & HPC Engineer. Be the first to apply!