Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Slurm Cluster & HPC Engineer

Full-time

Bitdeer Technologies Group

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

To learn more, visit (

Position Overview

  • We are seeking a Staff Slurm Cluster & HPC Scheduling Engineer to own Slurm as a first-class, productized scheduling layer across that fleet. This person is the single technical owner of Slurm cluster architecture, multi-tenant scheduling policy, and cluster reliability on both bare-metal and VM-based GPU nodes, and will lead our adoption of the Slinky operator stack (slurm-operator, and slurm-bridge where it fits) so that Slurm and Kubernetes workloads can share the same GPU pool. The role is deeply hands-on, customer-facing during onboarding and escalations, and sets the engineering standard the rest of the platform team builds on.

Key Responsibilities

  • Slurm cluster architecture and lifecycle — Design, deploy, and operate production Slurm clusters on bare metal and VMs: slurmctld/slurmdbd high availability, slurmrestd, configless slurmd, SACK/MUNGE and JWT authentication, and rolling version upgrades on live clusters without losing running jobs.
  • Topology-aware scheduling for GPU fabrics — Model the physical fabric in topology.conf — topology/tree for rail-optimized InfiniBand/RoCE designs and topology/block for NVLink domains such as GB200/GB300 NVL72 — and prove placement quality with NCCL bandwidth and multi-node training validation rather than assumption.
  • Multi-tenant scheduling policy — Own the account/association tree, partitions, QOS, fairshare, preemption, reservations, and per-tenant TRES limits. Enforce fail-closed defaults: an unresolved tenant identity or an empty entitlement set must deny, never degrade into unrestricted access.
  • Slinky on Kubernetes — Lead implementation of the Slinky slurm-operator, including its NodeSet, LoginSet, Accounting, RestAPI, and Token custom resources, cert-manager and Helm-based delivery, shared parallel-storage mounts, and login pods running sackd/sshd. Evaluate and pilot slurm-bridge for co-scheduling Kubernetes Pods, PodGroups, Jobs, JobSets, and LeaderWorkerSets through the Slurm scheduler, and document its constraints — notably exclusive whole-node allocation — before any customer exposure.
  • Elastic capacity between Slurm and Kubernetes — Use Slurm cloud and power-save mechanisms (ResumeProgram/SuspendProgram, SuspendTime, ResumeTimeout) together with fleet automation to shift GPU nodes between batch training queues and Kubernetes inference capacity as demand moves.
  • Container and job runtime — Operate Pyxis/Enroot and OCI/containerd job paths with correct gres.conf, cgroup v2 device constraints, and CUDA_VISIBLE_DEVICES behavior; support MPI/PMIx, module/Spack environments, and customer-supplied images.
  • Cluster health and reliability engineering — Build the passive and active health-check system expected of a top-tier GPU cloud: prolog/epilog checks, LBNL NHC or equivalent, DCGM diagnostics, and detection of XID/SXID errors, ECC faults, PCIe errors, GPUs falling off the bus, IB/RoCE link flaps, and NCCL stalls — with automatic drain and job requeue. Own burn-in and acceptance testing for every new rack before it carries paid work.
  • Automation and infrastructure as code — Deliver clusters through Terraform/Ansible, golden images, and bare-metal provisioning (PXE, Redfish, IPMI) so that a cluster build is reproducible, reviewable, and auditable rather than hand-tuned.
  • Observability, accounting, and billing integration — Instrument queue wait time, allocation efficiency, GPU utilization, and job failure taxonomy through a Slurm exporter into Prometheus/Grafana; configure AccountingStorageTRES and TRESBillingWeights, and reconcile sacct/sreport GPU-hours against the platform's metering and invoicing pipeline.
  • Technical leadership and customer engagement — Write runbooks and tenant-facing documentation, onboard and support enterprise customers, act as escalation point for cluster incidents, and mentor platform engineers on Slurm and HPC scheduling practice.

Qualifications

  • 8+ years in HPC, systems, or cloud infrastructure engineering, including 4+ years operating production Slurm clusters at 100+ GPU-node scale with real users and service-level commitments.
  • Deep hands-on Slurm expertise: slurm.conf, gres.conf, topology.conf, cgroup.conf, partitions/QOS/fairshare/preemption/reservations, slurmdbd accounting, slurmrestd, MUNGE/SACK and JWT authentication, and version upgrades performed on live clusters.
  • Strong GPU and fabric fundamentals: NVIDIA drivers and Fabric Manager, DCGM, MIG, InfiniBand/RoCEv2 (subnet manager/UFM, rail-optimized topology), GPUDirect RDMA, and practical NCCL tuning and failure diagnosis.
  • Production Kubernetes experience and working knowledge of the operator/CRD pattern, plus hands-on exposure to at least one Slurm-on-Kubernetes stack — Slinky slurm-operator or slurm-bridge, CoreWeave SUNK, or Nebius Soperator — with an informed view of the tradeoffs between them.
  • Experience delivering both bare-metal and virtualized compute: bare-metal provisioning and firmware/BIOS lifecycle management, hypervisor or VM-based clusters (KVM/QEMU or a public-cloud equivalent), and Terraform/Ansible-driven automation.
  • Working knowledge of parallel and shared storage for AI workloads — Lustre, GPFS/Spectrum Scale, WEKA, VAST, or NFS — and of how storage behavior shapes job performance and failure modes.
  • Proficient in Python and Bash for cluster automation; Go experience is a plus for integrating with Bitdeer AI's platform control plane and with Slurm/Slinky REST client code.
  • Multi-tenant security discipline: derives tenant scope from a verified identity rather than client-supplied fields, designs authorization to fail closed, and treats isolation across accounts, namespaces, storage, and networks as a hard requirement.
  • Clear written and verbal communication in English, with the maturity to work directly with enterprise customers and to translate scheduling and reliability tradeoffs for product, sales, and executive stakeholders

--------------------------------------------------------------------

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Staff Slurm Cluster & HPC Engineer in San Jose, CA vacancy
  • $136.3k - $231.7k

     ...Our expert teams of physicists, engineers, data scientists and problem-...  ...class team of physicists, HPC system designers, machine learning...  ..., Kubernetes/Docker, Slurm, Ray, TensorFlow, PyTorch, GPU...  ...on experience building AI/GPU cluster Minimum Qualifications... 
    Suggested
    Minimum wage
    Work experience placement
    Flexible hours

    KLA

    Milpitas, CA
    3 days ago
  •  ....Collaborate with product and engineering teams to map workload requirements...  ...baselines, rack-level, and cluster design.Act as a technical lead...  ...large-scale 10k-100k+ GPU HPC or cloud compute platforms.Deep...  ...orchestration tools used in HPC. (Slurm, Kubernetes, etc)Experience... 
    Suggested
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    1 day ago
  •  ...THE ROLE:We are seeking an AI Systems Engineer to join our AMD IT compute platforms...  ...of High-Performance Computing (HPC) infrastructure, GPU clusters, and AI workload schedulers. THE PERSON...  ...infrastructure engineering for AI/HPC domain SLURM and Kubernetes management Managing... 
    Suggested

    AMD

    San Jose, CA
    3 days ago
  • $124k - $241.5k

     ...As an HPC Operations Engineer at NVIDIA, you will play a pivotal role in ensuring the flawless operation of our high-performance computing (HPC...  ...Solid understanding of workload schedulers such as LSF, Slurm, or similar systems Strong grasp of network computing supporting... 
    Suggested
    Full time

    NVIDIA

    Santa Clara, CA
    2 days ago
  • $2,500 per month

     ...and staffed by leading engineers, Etched is redefining the...  ...performance computing (HPC), building systems that...  ...Optimize and manage Slurm-based job scheduling for...  ...infrastructure, and compute clusters. Manage and optimize...  ...all of our technical staff to contribute to both... 
    Suggested
    Work at office
    Relocation package

    Etched

    San Jose, CA
    9 days ago
  • $165k - $242k

    HPC Performance Engineer CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology,...  ...bare-metal systems from POST through joining a Kubernetes cluster. The team's primary responsibilities include maintaining a... 
    Full time
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    15 hours ago
  •  ...availability of large-scale GPU clusters. • Respond to incidents and...  ...Computer Science, Computer Engineering, Software Engineering,...  ...SRE, DevOps, cloud operations, HPC, or infrastructure operations...  ...Preferred Qualifications • Slurm. • GPU infrastructure. •... 
    Night shift

    Institute of Foundation Models

    Sunnyvale, CA
    28 days ago
  • Job-ID18347302Reference21-13197 Job Description: Design and implementation of high-performance compute clusters Solid knowledge on the HPC cluster systems, including scalable/robust storage, high-bandwidth inter-connects, CPU / GPU architecture, and a knowledge of cloud... 

    Intelliswift

    Milpitas, CA
    4 days ago
  • $184k - $287.5k

     ...next-gen distributed storage services for HPC workloads, optimizing both performance...  ...our researchers to run their flows on our clusters including performance analysis and...  ...degree in Computer Science, Electrical Engineering or related field or equivalent experience... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...and operate monitoring and alerting for cluster health — fabric, GPU, power/thermal, and...  ...proactivelyRemotely deploy and configure large-scale HPC clusters for AI workloads using...  ...and feed clear requirements back to other engineering teams on simplification, stability, and... 
    Work at office
    Local area
    Remote work
    Flexible hours

    Lambda Labs

    San Jose, CA
    1 day ago
  • $162.8k - $227.6k

     ...multi-dimensional lookup-table models. HPC & Scale: Build and scale High-...  ...infrastructure and job-scheduling workflows (e.g., Slurm) for large-scale 6DOF Monte Carlo...  ...release reusable analysis tools, and drive engineering best practices, including version control... 
    Local area
    Worldwide
    Visa sponsorship
    Flexible hours

    Archer Aviation

    San Jose, CA
    3 days ago
  • $152k - $241.5k

     ..., ensuring they integrate cleanly with HPC schedulers, storage, and network fabrics...  ....Experience supporting large‑scale HPC clusters using Slurm, LSF or Kubernetes clusters, including...  ...Python, Go, Perl, or Ruby.Mentored other engineers and influenced technical direction... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...architectural patterns for large-scale GPU clusters, storage backends, and multi-tenant...  ...telemetry, and operational visibility.Mentor engineers and cross-functional teams on advanced network...  ...data center networks, preferably for HPC, AI/ML, or large-scale cloud... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    1 day ago
  • $159k - $239k

     ...experienced Senior Linux Systems Engineer to join our dynamic...  ...orchestration, and HPC environments. This role...  ..., and mentor junior staff while ensuring high availability...  ...or OpenShift clusters to orchestrate container...  ...scheduling systems (e.g., Slurm, LSF), optimizing... 
    Full time
    Work experience placement
    Work at office
    Local area

    Ampere

    Santa Clara, CA
    1 day ago
  • $95k - $161.5k

     ...15% of sales back into R&D. Our expert teams of physicists, engineers, data scientists and problem-solvers work together with the world...  ...QualificationsKey ResponsibilitiesDesign & configure HPC clusters - Support development of compute cluster architectures optimized... 
    Minimum wage
    Full time
    Work experience placement
    Worldwide
    Flexible hours

    KLA-Tencor

    Milpitas, CA
    4 days ago
  • $152k - $241.5k

     ..., measurable, and aligned with long-term engineering demands.What you'll be doing:Manage, scale...  ...optimize job scheduling systems (LSF, Slurm, etc.) in a large-scale, multi-site environment...  ...scheduling systems (LSF, Slurm, etc.) in HPC or silicon design environmentsProficiency... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $120k - $140k

     ...Enterprise IT, Hadoop/ Big Data, Hyperscale, HPC and IoT/Embedded customers worldwide....  ...talented, passionate, and committed engineers, technologists, and business leaders...  ...with workload/scheduler Managers (Slurm) for rack/cluster * Familiar with MLPerf Training/Inference... 
    Worldwide

    Supermicro

    San Jose, CA
    2 days ago
  • $136.3k - $199.9k

     ...systems is a powerful High-Performance Computing (HPC) infrastructure. We are seeking a highly skilled engineer to architect, develop, deploy, and support...  ...you will:Architect, deploy, and support Kubernetes clusters running on enterprise Linux platforms across HPC... 
    Minimum wage
    Full time
    Work experience placement

    KLA-Tencor

    Milpitas, CA
    3 days ago
  • $184k - $287.5k

     ...world. We are looking for a dedicated engineer for the Senior Systems Software Engineer...  ...develop new, leading solutions. Engage with HPC, OS, CPU, GPU compute, and systems...  ...enterprise computing architectures, with Slurm preferred. Strong programming and scripting... 

    NVIDIA

    Santa Clara, CA
    2 days ago
  • $167k - $193k

     ...invented the world’s first 3D-stacked photonics engine, Passage™, capable of connecting thousands...  ...data centers for the most advanced AI and HPC workloads. Lightmatter raised $400...  ...computing with light! We are seeking a Staff Hardware Test & Bring-Up Engineer to lead... 
    Full time
    Temporary work
    Flexible hours

    Lightmatter

    Mountain View, CA
    3 days ago
  •  ...edge, and cloud. About the Role As a Forward Deployed Engineer (FDE), you are a core member of our engineering team embedded...  ...-scale AI training environments or high-performance compute (HPC) clusters is highly desirable. Modern AI Infrastructure (Plus):... 

    VAST Data

    San Jose, CA
    4 days ago
  •  ...Job Description Title: CAE Engineer 598268 Location: San Jose, CA Description...  ...3+ years of hands-on experience in HPC infrastructure, EDA/engineering IT infrastructure...  ...knowledge of job schedulers (LSF, SLURM, Grid Engine, or equivalent), license... 

    West Coast Consulting LLC

    San Jose, CA
    1 day ago
  • $137k - $156k

     ...Enterprise IT, Hadoop/ Big Data, Hyperscale, HPC and IoT/Embedded customers worldwide. We...  ...seek talented, passionate, and committed engineers, technologists, and business leaders to...  ...may also be assigned): * Deploy Rack/Cluster infrastructure and execute comprehensive... 
    Worldwide

    Supermicro

    San Jose, CA
    3 days ago
  • $153k - $204k

     ...Senior Systems Engineer, Test Frameworks & Validation Platform Livingston, NJ / New...  ...we're extending that framework into HPC verification, Slurm-on-Kubernetes, and further down the stack...  ...) experience. HPC or large-cluster experience — InfiniBand/RoCE, GPU/accelerator... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Live in
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    2 days ago
  • $110k - $160k

     ...overall code quality. Own test reporting and articulately communicate technical challenges, solutions, and mitigation plans to both engineering teams and management.   Education / Experience: ~ Bachelors in Electrical Engineering, Digital Sciences, Computer... 

    SK hynix memory solutions America Inc.

    San Jose, CA
    18 days ago
  • $170k - $210k

     ...Senior Manager, Power Electronics Firmware ChargePoint is seeking a seasoned Power Electronics Controls and Firmware Engineer with over 5 years of experience in embedded power control firmware development. The ideal candidate should have expertise in Control... 
    Work experience placement

    ChargePoint

    Campbell, CA
    4 days ago
  • $136.3k - $231.7k

    A leading global tech firm is seeking a motivated algo engineer in Milpitas, CA. This role involves developing algorithmic solutions for image modeling, encompassing tasks from conception to productization. Required qualifications include a PhD or MS in a relevant field... 

    KLA-Belgium

    Milpitas, CA
    3 days ago
  • $136.3k - $199.9k

     ...advancing business priorities by delivering high-impact work across your area of expertise.Preferred QualificationsWe are looking for Engineers who are passionate and driven to develop and release products that address unique customer needs and problems.Required Technical... 
    Minimum wage
    Full time
    Work experience placement
    Flexible hours

    KLA-Tencor

    Milpitas, CA
    2 days ago
  • $184k - $287.5k

    NVIDIA has become the platform upon which every new AI-powered application is built. We are seeking a Sr. HPC Performance engineer to join our team of scientists and engineers passionate about building the next generation of scientific machine learning (ML) frameworks.... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

    NVIDIA Math Libraries team is looking for a senior engineer to join our development efforts in the area of kernel generation for AI and HPC, specifically targeting matrix operations, JITing and fusions. Around the world, leading commercial and academic organizations are... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Slurm Cluster & HPC Engineer. Be the first to apply!