Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Slurm Cluster & HPC Engineer [Remote]

Full-time

Bitdeer Technologies Group

San Jose, CA
  • Remote job

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

To learn more, visit (

Position Overview

  • We are seeking a Staff Slurm Cluster & HPC Scheduling Engineer to own Slurm as a first-class, productized scheduling layer across that fleet. This person is the single technical owner of Slurm cluster architecture, multi-tenant scheduling policy, and cluster reliability on both bare-metal and VM-based GPU nodes, and will lead our adoption of the Slinky operator stack (slurm-operator, and slurm-bridge where it fits) so that Slurm and Kubernetes workloads can share the same GPU pool. The role is deeply hands-on, customer-facing during onboarding and escalations, and sets the engineering standard the rest of the platform team builds on.

Key Responsibilities

  • Slurm cluster architecture and lifecycle — Design, deploy, and operate production Slurm clusters on bare metal and VMs: slurmctld/slurmdbd high availability, slurmrestd, configless slurmd, SACK/MUNGE and JWT authentication, and rolling version upgrades on live clusters without losing running jobs.
  • Topology-aware scheduling for GPU fabrics — Model the physical fabric in topology.conf — topology/tree for rail-optimized InfiniBand/RoCE designs and topology/block for NVLink domains such as GB200/GB300 NVL72 — and prove placement quality with NCCL bandwidth and multi-node training validation rather than assumption.
  • Multi-tenant scheduling policy — Own the account/association tree, partitions, QOS, fairshare, preemption, reservations, and per-tenant TRES limits. Enforce fail-closed defaults: an unresolved tenant identity or an empty entitlement set must deny, never degrade into unrestricted access.
  • Slinky on Kubernetes — Lead implementation of the Slinky slurm-operator, including its NodeSet, LoginSet, Accounting, RestAPI, and Token custom resources, cert-manager and Helm-based delivery, shared parallel-storage mounts, and login pods running sackd/sshd. Evaluate and pilot slurm-bridge for co-scheduling Kubernetes Pods, PodGroups, Jobs, JobSets, and LeaderWorkerSets through the Slurm scheduler, and document its constraints — notably exclusive whole-node allocation — before any customer exposure.
  • Elastic capacity between Slurm and Kubernetes — Use Slurm cloud and power-save mechanisms (ResumeProgram/SuspendProgram, SuspendTime, ResumeTimeout) together with fleet automation to shift GPU nodes between batch training queues and Kubernetes inference capacity as demand moves.
  • Container and job runtime — Operate Pyxis/Enroot and OCI/containerd job paths with correct gres.conf, cgroup v2 device constraints, and CUDA_VISIBLE_DEVICES behavior; support MPI/PMIx, module/Spack environments, and customer-supplied images.
  • Cluster health and reliability engineering — Build the passive and active health-check system expected of a top-tier GPU cloud: prolog/epilog checks, LBNL NHC or equivalent, DCGM diagnostics, and detection of XID/SXID errors, ECC faults, PCIe errors, GPUs falling off the bus, IB/RoCE link flaps, and NCCL stalls — with automatic drain and job requeue. Own burn-in and acceptance testing for every new rack before it carries paid work.
  • Automation and infrastructure as code — Deliver clusters through Terraform/Ansible, golden images, and bare-metal provisioning (PXE, Redfish, IPMI) so that a cluster build is reproducible, reviewable, and auditable rather than hand-tuned.
  • Observability, accounting, and billing integration — Instrument queue wait time, allocation efficiency, GPU utilization, and job failure taxonomy through a Slurm exporter into Prometheus/Grafana; configure AccountingStorageTRES and TRESBillingWeights, and reconcile sacct/sreport GPU-hours against the platform's metering and invoicing pipeline.
  • Technical leadership and customer engagement — Write runbooks and tenant-facing documentation, onboard and support enterprise customers, act as escalation point for cluster incidents, and mentor platform engineers on Slurm and HPC scheduling practice.

Qualifications

  • 8+ years in HPC, systems, or cloud infrastructure engineering, including 4+ years operating production Slurm clusters at 100+ GPU-node scale with real users and service-level commitments.
  • Deep hands-on Slurm expertise: slurm.conf, gres.conf, topology.conf, cgroup.conf, partitions/QOS/fairshare/preemption/reservations, slurmdbd accounting, slurmrestd, MUNGE/SACK and JWT authentication, and version upgrades performed on live clusters.
  • Strong GPU and fabric fundamentals: NVIDIA drivers and Fabric Manager, DCGM, MIG, InfiniBand/RoCEv2 (subnet manager/UFM, rail-optimized topology), GPUDirect RDMA, and practical NCCL tuning and failure diagnosis.
  • Production Kubernetes experience and working knowledge of the operator/CRD pattern, plus hands-on exposure to at least one Slurm-on-Kubernetes stack — Slinky slurm-operator or slurm-bridge, CoreWeave SUNK, or Nebius Soperator — with an informed view of the tradeoffs between them.
  • Experience delivering both bare-metal and virtualized compute: bare-metal provisioning and firmware/BIOS lifecycle management, hypervisor or VM-based clusters (KVM/QEMU or a public-cloud equivalent), and Terraform/Ansible-driven automation.
  • Working knowledge of parallel and shared storage for AI workloads — Lustre, GPFS/Spectrum Scale, WEKA, VAST, or NFS — and of how storage behavior shapes job performance and failure modes.
  • Proficient in Python and Bash for cluster automation; Go experience is a plus for integrating with Bitdeer AI's platform control plane and with Slurm/Slinky REST client code.
  • Multi-tenant security discipline: derives tenant scope from a verified identity rather than client-supplied fields, designs authorization to fail closed, and treats isolation across accounts, namespaces, storage, and networks as a hard requirement.
  • Clear written and verbal communication in English, with the maturity to work directly with enterprise customers and to translate scheduling and reliability tradeoffs for product, sales, and executive stakeholders

--------------------------------------------------------------------

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Vacancy posted 29 days ago
Similar jobs that could be interesting for youBased on the Staff Slurm Cluster & HPC Engineer [Remote] in San Jose, CA vacancy
  • $152k - $241.5k

     ...on the world.We are seeking a highly skilled and experienced HPC Cluster Engineer to design, deploy, and operate GPU Compute Clusters for EDA...  ...Experience with AI/HPC job schedulers and orchestrators, such as Slurm, LSF, PBS or K8s. Applied experience with AI/HPC workflows... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...next-gen distributed storage services for HPC workloads, optimizing both performance...  ...our researchers to run their flows on our clusters including performance analysis and...  ...degree in Computer Science, Electrical Engineering or related field or equivalent experience... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • Job-ID18347302Reference21-13197 Job Description: Design and implementation of high-performance compute clusters Solid knowledge on the HPC cluster systems, including scalable/robust storage, high-bandwidth inter-connects, CPU / GPU architecture, and a knowledge of cloud... 
    Suggested

    Intelliswift

    Milpitas, CA
    1 day ago
  •  ...THE ROLE:We are seeking an AI Systems Engineer to join our AMD IT compute platforms...  ...of High-Performance Computing (HPC) infrastructure, GPU clusters, and AI workload schedulers. THE PERSON...  ...infrastructure engineering for AI/HPC domain SLURM and Kubernetes management Managing... 
    Suggested

    AMD

    San Jose, CA
    1 day ago
  • $136.3k - $231.7k

     ...into R&D. Our expert teams of physicists, engineers, data scientists and problem-solvers work...  ...of a world class team of physicists, HPC system designers, machine learning and application...  ...Linux administration, Kubernetes/Docker, Slurm, Ray, TensorFlow, PyTorch, GPU/TPU... 
    Suggested
    Minimum wage
    Full time
    Work experience placement
    Flexible hours

    KLA-Tencor

    Milpitas, CA
    1 day ago
  •  ...deployment of large-scale AI/ML clustered infrastructure. You will be working...  ...team of multi-disciplined engineers that operates across industry verticals...  ...performance tuningOrchestration models: Slurm configuration patterns, Kubernetes for HPC/AI (GPU operators, device plugins... 
    Flexible hours

    AMD

    Santa Clara, CA
    1 day ago
  • $124k - $195.5k

    As an HPC Operations Engineer at NVIDIA, you will play a pivotal role in ensuring the flawless operation of our high-performance computing (HPC)...  ...tasksSolid understanding of workload schedulers such as LSF, Slurm, or similar systemsStrong grasp of network computing... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $95k - $161.5k

     ...15% of sales back into R&D. Our expert teams of physicists, engineers, data scientists and problem-solvers work together with the world...  ...QualificationsKey ResponsibilitiesDesign & configure HPC clusters - Support development of compute cluster architectures optimized... 
    Minimum wage
    Full time
    Work experience placement
    Worldwide
    Flexible hours

    KLA-Tencor

    Milpitas, CA
    1 day ago
  • $162.8k - $217.6k

     ...well as computing infrastructure to enable engineers to solve problems faster and more...  ...efficiently utilize High-Performance Computing (HPC) resources, and make informed decisions...  ...Experience with HPC management software (Slurm/PBS/Torque, OpenHPC/Bright, Warewulf/XCat... 
    Local area

    Archer Aviation

    San Jose, CA
    4 days ago
  • $152k - $241.5k

     ..., measurable, and aligned with long-term engineering demands.What you'll be doing:Manage, scale...  ...optimize job scheduling systems (LSF, Slurm, etc.) in a large-scale, multi-site environment...  ...scheduling systems (LSF, Slurm, etc.) in HPC or silicon design environmentsProficiency... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $255k - $340k

     ...from home day is currently Tuesday.Hardware Engineering at Lambda is responsible for building and...  ...system integration validation for new HPC AI/ML, general purpose compute, storage,...  ...Hardware Engineer, Data Center Engineering, and Cluster Network Design to ensure new platforms... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  • $184k - $287.5k

    NVIDIA Math Libraries team is looking for a senior engineer to join our development efforts in the area of kernel generation for AI and HPC, specifically targeting matrix operations, JITing and fusions. Around the world, leading commercial and academic organizations are... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

    NVIDIA has become the platform upon which every new AI-powered application is built. We are seeking a Sr. HPC Performance engineer to join our team of scientists and engineers passionate about building the next generation of scientific machine learning (ML) frameworks.... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • AMD, Inc. is seeking a PMTS Systems Design Engineer to research, design, develop, and test...  ...operations, and to integrate software for GPU clusters supporting AI inferencing and training....  ...in GPU performance, RDMA networking, and HPC system design. #J-18808-Ljbffr Socket.... 
    Remote job

    Socket.dev

    Santa Clara, CA
    2 days ago
  • $165k - $242k

     ...HPC Performance Engineer CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology...  ...bare-metal systems from POST through joining a Kubernetes cluster. The team's primary responsibilities include maintaining a... 
    Full time
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    5 days ago
  • $155k - $185k

     ..., Cloud Computing, Enterprise IT, Hadoop/ Big Data, Hyperscale, HPC and IoT/Embedded customers worldwide. We are the #5 fastest growing...  ...community. We seek talented, passionate, and committed engineers, technologists, and business leaders to join us.Job Summary:As a... 
    Contract work
    Immediate start
    Worldwide

    Supermicro

    San Jose, CA
    1 day ago
  • $74 - $84 per hour

    Job Title: CAE Engineer Position Description: Protingent Staffing has...  ...Infrastructure Engineer to operate the HPC compute infrastructure and...  ...support for compute clusters, application environments, license...  ...knowledge of job schedulers (LSF, SLURM, Grid Engine, or equivalent),... 
    Contract work
    Remote work

    Protingent

    San Jose, CA
    3 days ago
  • NVIDIA has become the platform upon which every new AI-powered application is built. We are seeking a Sr. HPC Performance engineer to join our team of scientists and engineers passionate about building the next generation of scientific machine learning (ML) frameworks.... 

    NVIDIA

    Santa Clara, CA
    5 days ago
  • $2,000 per month

     ...and staffed by leading engineers, Etched is redefining the...  ...performance computing (HPC), building systems that...  ...Optimize and manage Slurm-based job scheduling for...  ...infrastructure, and compute clusters. Manage and optimize...  ...all of our technical staff to contribute to both... 
    Work at office
    Relocation package

    ETCHED LLC

    San Jose, CA
    5 days ago
  •  ...availability of large-scale GPU clusters. • Respond to incidents and...  ...Computer Science, Computer Engineering, Software Engineering,...  ...SRE, DevOps, cloud operations, HPC, or infrastructure operations...  ...Preferred Qualifications • Slurm. • GPU infrastructure. •... 
    Night shift

    Institute of Foundation Models

    Sunnyvale, CA
    a month ago
  • $189k - $301k

    A leading semiconductor company in San Jose is seeking a Senior Staff Engineer specializing in High-Speed I/O Analog-Mixed Circuit design. The role requires extensive experience in analog circuit design and high-speed interfaces, with a focus on optimization and simulation... 

    Jobleads-US

    San Jose, CA
    3 days ago
  • $183k - $247.6k

     ...powers breakthrough innovation in AI/ML and HPC workloads. If you're passionate about...  ...team of software, hardware, and network engineers, supply chain specialists, security experts...  ...cooperatively with other employees, supervisors, and staff; adhere to standards of excellence... 
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $150.98k - $218.62k

     ...innovators stay Ahead of What's Possible. Learn more at and on LinkedIn and X.Staff System Architecture and Design EngineerData Center & Energy (DCE)San Jose, CAAbout the RoleAs a System Engineer, you will own the system development of high-voltage power converters — from... 
    Permanent employment
    Full time
    Work at office
    Shift work
    Day shift

    Analog Devices

    San Jose, CA
    2 days ago
  • $148.7k - $201.2k

     ...product line. The Nitro Team is looking for engineers with systems knowledge and experience in...  ...workloads.The Nitro High Memory and HPC team owns the purpose built platform development...  ...with other employees, supervisors, and staff; adhere to standards of excellence... 
    Internship
    Local area
    Flexible hours

    Amazon

    Santa Clara, CA
    3 days ago
  • $131k - $175k

     ...excellence has earned us several prestigious awards, such as Best Engineering Team, Best Company for Diversity, Compensation, and Work-Life...  ...platforms integrate cleanly into large-scale AI and cloud clusters.What You’ll Do Lead the design, validation, and deployment of... 
    Remote work
    Flexible hours

    Arista Networks

    Santa Clara, CA
    3 days ago
  • $140k - $224.25k

     ...include gaming, automotive, vision, HPC, datacenters and networking in...  ...telemetries, scale out cluster, test plan development, track...  ...a STEM (Science, Technology, Engineering, Math or Physics) field5+ years...  ...in GitHub/Gitlab/Gerrit, PXE, SLURM, Stack/Kubernetes/Docker) - huge... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of...  ...and tensor parallelism across multi-node clusters. Analyze network-level bottlenecks using...  ...in GPU performance engineering, HPC, or systems performance analysisHands-on... 

    AMD

    Santa Clara, CA
    1 day ago
  • $147.75k - $221.63k

     ...jobWe are seeking a highly motivated High Performance Computing (HPC) Engineer to design, develop, deploy, and maintain cutting-edge...  ...Unix.Troubleshooting and resolving hardware problems on Linux cluster hardware and networking equipment and applying logical methods... 
    Full time
    Work experience placement

    ASML Holding

    San Jose, CA
    4 days ago
  • $100.55k - $194.27k

    Job Details:Job Description: The world is transforming - and so is Intel. Intel is a company of bold and curious inventors and problem solvers who create some of the most astounding technology advancements and experiences in the world. With a legacy of relentless innovation...
    Full time
    Internship
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    8 hours ago
  • $320k

     ...CPUs, and a fully optimized NVIDIA AI and HPC software stack. We're searching for a...  ...motivated, technical leader to drive the engineering roadmap and innovation for our rack system...  ...crowd:Knowledge of large-scale cloud and cluster level deployment and management systems.... 
    Full time
    Shift work

    Nvidia

    Santa Clara, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Slurm Cluster & HPC Engineer [Remote]. Be the first to apply!