Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior HPC Systems Engineer

Parallel Works

About Parallel Works

Parallel Works builds and operates ACTIVATE, a control plane for high performance computing and AI. Our customers run large scientific and AI workloads across their own on-premises clusters, Government and commercial cloud, and commercial GPU providers, and ACTIVATE gives them one way in to all of it. The high security boundary is authorized at Impact Level 5, with FIPS validated cryptography and STIG hardening throughout.

The work reaches most fields that depend on computing at scale: weather and climate forecasting, defense and intelligence programs, aerospace and structural analysis, molecular and materials science, energy, and AI research. A quarter here can include standing up a GPU cluster for one of those communities, federating a laboratory's existing on-premises system with burst capacity it did not have before, and getting a domain code written decades ago to run on current hardware.

Customer success sets our priorities. We are a small engineering company, so engineers here work directly with the people using the systems and carry a problem from the first report through to the fix. This is what we call mission engineering: understanding what a customer is trying to accomplish and why the computing matters to it.

About the role

Parallel Works is hiring a Senior HPC Systems Engineer to build and run the clusters behind our defense and research programs. The work covers GPU node bring-up, Slurm configuration, fabric and storage troubleshooting, security hardening, and Tier 3 escalation.

The computing environments are hybrid. Some clusters are customer owned hardware on site, some run in accredited Government cloud regions, and some are dedicated GPU clusters at commercial providers. On several programs the on-premises systems carry the primary load and cloud takes the overflow. The position is senior: it handles the escalations the rest of the team cannot resolve, and it trains the junior engineers.

What you will do
  • Cluster operations: build and operate production Slurm clusters. slurmctld and slurmdbd, partitions and QOS, accounts and fair share, GPU GRES, prolog and epilog, cgroup enforcement.
  • Hybrid federation: connect customer owned clusters to the control plane, reconciling their site scheduler, storage, and identity source so accounts and allocations behave the same in every venue.
  • On-premises hardware: bare metal provisioning, out of band management, firmware, rack networking, and fault coordination with site staff or vendors.
  • GPU and fabric: validate GPU nodes before users arrive. Driver and CUDA stack, DCGM health checks, XID triage, fabric manager and NVLink checks, InfiniBand verification, NCCL tuning.
  • Storage and automation: tune parallel and high throughput storage, and write the Ansible, Terraform, and image build pipelines that make a cluster reproducible.
  • Security and escalation: STIG hardening, scan remediation, FIPS validated cryptography, security package artifacts, Tier 3 escalations, and a share of the on call rotation.

Requirements

  • 10 or more years operating production Linux systems across more than one distribution family. RHEL, Rocky, or Alma on the Government side and Debian or Ubuntu on the commercial GPU side, since those clusters usually ship Ubuntu. Kernel and network tuning, systemd, cgroups, NUMA.
  • Production Slurm administration. You have configured, debugged, and upgraded a scheduler other people depended on.
  • At least one parallel or high throughput filesystem in production, plus InfiniBand or RoCE fabric operations.
  • NVIDIA GPU node operations at multi-node scale, including driver stack management and fault triage.
  • Experience on customer owned or on-premises clusters as well as public cloud, including bare metal provisioning and out of band management.
  • DevOps experience with infrastructure as code frameworks such as Ansible or Terraform.
  • Proficiency with common scripting languages such as Bash and Python.
  • United States citizenship and eligibility for a Secret clearance, since the work reaches export controlled Government environments. An active clearance is helpful but not required. We sponsor candidates who are eligible but not currently cleared.

You do not need every item on this list. If you have most of it and work well with other people, apply.

Preferred Qualifications
  • Time at a Government supercomputing center, national laboratory, or university research computing center.
  • Work with STIG, SCAP, Tenable, eMASS, or RMF, or time inside FedRAMP or Impact Level boundaries.
  • Running GPU workloads on Kubernetes or OpenShift with Helm and operators, and using HPC containers such as Apptainer, Enroot, or Pyxis.
  • Running PBS Pro, Slurm, and monitoring with Prometheus and Grafana, including utilization and chargeback reporting.

Benefits

Medical, vision, and dental coverage, a 401(k) with company match, short term disability, and generous paid vacation and sick time.

Equal employment opportunity

Parallel Works is an equal opportunity employer. We consider all qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, protected veteran status, or any other characteristic protected by law.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior HPC Systems Engineer in United States vacancy
  • $122.85k

    Posting Description SENIOR HPC SYSTEMS ENGINEER, The Massachusetts Green High Performance Computing Center (MGHPCC), to support the computing infrastructure behind the Massachusetts AI Hub’s AI Computing Resource (AICR). This hands-on role will be responsible for deploying... 
    Senior
    Full time
    Visa sponsorship

    Massachusetts Institute of Technology

    Cambridge, MA
    13 hours ago
  •  ...Cybotic System seeks an HPC Administrator/Engineer to manage a large-scale HPC cluster in Savannah, GA. The role requires strong Linux expertise (RPM-based distributions) and 7+ years in HPC or scientific computing environments, with the ability to lead upgrades and drive... 
    Senior

    Cybotic System

    Savannah, GA
    5 days ago
  •  ...Senior HPC Operations Engineer | Chicago Join one of the world's most advanced high-performance computing environments supporting cutting-edge quantitative...  ...research. We're looking for an experienced HPC Systems Engineer to help operate and scale a large production HPC... 
    Senior

    Autonomai Recruitment

    Chicago, IL
    5 days ago
  •  ...To support defense and research programs, the full-time Senior HPC Systems Engineer will build and operate production Slurm clusters, manage hybrid federation of customer-owned clusters, and oversee GPU node operations in a remote environment. Key responsibilities Build... 
    Senior
    Full time
    Remote work

    Virtual Vocations Inc

    United States
    1 day ago
  • $224k - $356.5k

     ...inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.We are seeking a Senior HPC & Quantum Systems Engineer to help architect, deploy, and operate a first-of-its-kind Accelerated Quantum Computing Center in the Boston area!... 
    Senior
    Full time
    Work at office
    Remote work

    Nvidia

    Westford, MA
    9 hours ago
  • $170k - $260k

     ...Are you a Senior HPC Systems Engineer who is ready for a new challenge that will launch your career to the next level? Tired of being treated like a company drone? Tired of promised adventures during the hiring phase, then dropped off on a remote contract and never... 
    Senior
    Full time
    Contract work
    Remote work
    Work from home
    Relocation package

    GliaCell Technologies LLC

    Annapolis Junction, MD
    2 days ago
  • Parallel Works seeks a Senior HPC Systems Engineer to build and run production Slurm clusters and oversee GPU nodes across on-prem and cloud environments. You will handle bare metal provisioning, network and firmware tasks, and ensure secure, scalable operation for defense... 
    Senior

    Parallel Works

    Chicago, IL
    4 days ago
  •  ...of Kansas City seeks an experienced High Performance Computing Engineer to plan, implement, and maintain advanced cyberinfrastructure for...  ...economic research initiatives. Responsibilities include deploying HPC clusters, managing storage, and optimizing workflows. The role... 
    Senior

    Federal Reserve Bank of Kansas City

    Kansas City, MO
    4 days ago
  • A growing infrastructure company is seeking a Senior Systems Engineer to support the Department of Energy. This role involves guiding national labs...  .... Strong technical presentation skills and understanding of HPC and AI/ML are crucial. Join us for this pivotal opportunity... 
    Senior

    VAST Data

    New York, NY
    3 days ago
  • The Massachusetts Institute of Technology is seeking a Senior HPC Systems Engineer for the Massachusetts Green High Performance Computing Center (MGHPCC). You will support the computing infrastructure needed for the Massachusetts AI Hub’s AI Computing Resource (AICR). Your... 
    Senior

    Massachusetts Institute of Technology

    Cambridge, MA
    4 days ago
  • $180k - $220k

     ...owned IT business delivering innovative solutions since 1982. The Systems Engineer 4 role requires active TS/SCI with Poly and involves supporting a high‑profile government program focused on cloud, HPC, and enterprise architecture in Maryland. The position demands excellent... 
    Senior

    DCCA

    Annapolis, MD
    4 days ago
  •  ...implementation, and management of High-Performance Computing (HPC) systems within a classified environment. We are looking for...  ...critical computing environments.Serve as a technical mentor for HPC engineers, guiding best practices in automation, performance tuning, and... 
    Senior
    Work at office

    Oak Ridge National Laboratory

    Oak Ridge, TN
    3 days ago
  • SpaceX in Hawthorne, California, seeks a Sr. HPC Systems Engineer to administer HPC clusters, storage, and high-speed networks across a world-class engineering org. The role involves providing application support to SpaceX engineers, installing and integrating Linux compute... 
    Senior

    InvestedintheMission

    Hawthorne, CA
    1 day ago
  • SpaceX is seeking an Senior HPC Systems Engineer to administer HPC clusters, storage systems, and networking, providing cross‑disciplinary support and building out Linux compute environments. The role emphasizes experience with cluster management, scripting for automation... 
    Senior

    SpaceX

    Hawthorne, CA
    4 days ago
  • GliaCell Technologies is looking for a Senior HPC Systems Engineer for a full‑time position supporting U.S. Government projects. You will develop requirements management processes and deliver system engineering documentation, ensuring consistency with strategic plans and... 
    Senior
    Remote job
    Full time

    GliaCell Technologies

    Annapolis, MD
    13 hours ago
  • $170k - $260k

    GliaCell Technologies is hiring a Senior HPC Systems Engineer for a full-time role to provide technical expertise in sustaining mission-critical software and systems for a U.S. Government subcontract. The ideal candidate will develop requirements management processes, deliver... 
    Senior
    Remote job
    Full time

    GliaCell Technologies

    Annapolis, MD
    4 days ago
  • SpaceX is hiring a Sr. HPC Systems Engineer in Hawthorne, CA to administer HPC clusters, storage systems, and fast networks. You will provide application support to engineers, install Linux compute clusters, and document procedures for diverse teams, conveying complex... 
    Senior

    SPACE EXPLORATION TECHNOLOGIES CORP

    Hawthorne, CA
    4 days ago
  • $200k - $220k

     ...delivering innovative IT solutions since 1982. The role requires TS/SCI w/Poly and involves engineering across the full system life cycle, emphasizing reliability and resiliency for HPC and IT programs. Salary range Maryland: $200,000-$220,000. Benefits include healthcare,... 
    Senior

    International Executive Service Corps

    Annapolis, MD
    4 days ago
  • $255k - $340k

     ...designated work from home day is currently Tuesday.Hardware Engineering at Lambda is responsible for building and scaling the...  ...operate at, this is that team.What You’ll DoOwn system integration validation for new HPC AI/ML, general purpose compute, storage, and network hardware... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    3 days ago
  • $224k - $356.5k

     ...impact on the world. What you'll be doing: Hybrid Quantum-HPC Platform Engineering Build, deploy, and operate a hybrid computing platform combining...  ..., and future modalities). Integrate quantum control systems and access nodes with HPC infrastructure using APIs, scientific... 
    Senior
    Work at office

    NVIDIA

    Hartford, CT
    2 days ago
  • ExxonMobil is seeking a highly skilled HPC Systems Engineer to manage a performant and reliable high-performance computing environment at its Spring, TX campus. The role involves overseeing system operations, evaluating new hardware, and consulting with users for effective... 
    Senior

    ExxonMobil

    Spring, Montgomery County, TX
    3 days ago
  •  ...Radix Trading is looking for an experienced HPC Systems Engineer to enhance its research infrastructure. The role involves troubleshooting technical issues, optimizing workloads, and managing performance across HPC systems. The ideal candidate has over 5 years of HPC... 
    Senior

    Radix Trading Experienced Job Board

    Chicago, IL
    2 days ago
  •  ...company based in Houston is seeking a Sr. Linux IT Specialist to support our global HPC and Digital Platform team. The ideal candidate will have extensive experience in Linux systems and a Bachelor's degree in Computer Science or related field. Responsibilities include... 
    Senior

    Viridiengroup

    Houston, TX
    1 day ago
  • $95.93k - $110k

    The University Of Chicago is looking for a skilled professional to install, configure, and maintain large computer clusters and servers. This role requires expertise in Linux build automation and experience in scripting. The position offers a competitive salary ranging...
    Senior

    The University of Chicago

    Chicago, IL
    2 days ago
  •  ...data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and...  ...industry, academia, and national labs.The AMD HPC & Sovereign AI applications team seeks a...  ...developers, customers, and product engineers to deliver programs on-time. You delight... 
    Senior

    AMD

    Austin, TX
    13 hours ago
  • NVIDIA is seeking a Senior Software Engineer in Westford, Massachusetts to improve their HPC infrastructure. The role includes designing scalable systems and supporting multi-cloud environments. The ideal candidate will have 10+ years of experience, strong software development... 
    Senior

    NVIDIA

    Santa Clara, CA
    13 hours ago
  • $85.5k - $149.8k

    ****@*****.*** Research Computingis seeking a Sr. HPC Systems Engineer who will design, build, and maintain advanced high-performance computing environments supporting Johns Hopkins University’s research mission. This position focuses on the reliable operation, configuration, and... 
    Senior
    Full time

    Johns Hopkins University

    Baltimore, MD
    13 hours ago
  • $165k - $230k

     ...the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. HIGH PERFORMANCE COMPUTING (HPC) SYSTEMS ENGINEER SpaceX is looking for an HPC Systems Engineer with strong knowledge and experience in a world class engineering organization... 
    Senior
    Permanent employment
    Temporary work
    Flexible hours
    Weekend work

    SpaceX

    Hawthorne, CA
    13 hours ago
  • Career Techniques in New York seeks an experienced infrastructure engineer to design, deploy, and scale large-scale GPU clusters for AI research. You will work across compute, storage, OS, and automation to support hundreds of petabytes and thousands of nodes. You will... 
    Senior

    Career Techniques

    New York, NY
    3 days ago
  • $137k - $176k

    The Solutions Group Llc is looking for a Systems Engineer to help design the next generation of high-performance compute infrastructure in Annapolis, Maryland. This role involves creating solutions that connect legacy systems with modern platforms, ensuring security and... 
    Senior

    The Solutions Group Llc

    Annapolis, MD
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior HPC Systems Engineer. Be the first to apply!