Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff HPC Systems Software Engineer

Full-time

Nscale

Role Description

We’re hiring a Staff HPC Systems Software Engineer to define the technical direction and evolution of a core HPC platform domain at Nscale.

In this role, you will operate beyond a single team, shaping how multiple teams build, automate, and run Slurm-based capabilities within Nscale’s wider cloud-native platform. You’ll work across engineering boundaries to bring coherence to architecture, interfaces, lifecycle models, and operational approaches, while partnering closely with teams working on platform tooling, infrastructure APIs, identity systems, and Kubernetes-adjacent systems.

This is a high-impact staff-level role for someone who combines deep hands-on software engineering with strong systems judgement. Your work will help ensure Nscale’s HPC services are robust, supportable, and maintainable, while creating leverage through shared patterns, reusable implementations, and clear technical direction across ambiguous, business-critical problem spaces.

What you'll be doing

  • Domain Architecture & Technical Direction
    • Own and evolve the technical direction for a defined HPC systems domain, such as Slurm platform architecture, scheduler integrations, cluster lifecycle, workload environments, or service automation.
    • Make architectural decisions that balance software quality, operational realities, customer needs, and long-term maintainability.
    • Define how proven Slurm implementations should be packaged, automated, and exposed as a service.
    • Resolve ambiguity around ownership, interfaces, lifecycle boundaries, and operating models across teams.
    • Act as the technical escalation point for the most complex issues within the domain.
  • Cross-Team Engineering Leverage
    • Establish shared patterns and standards for automation, service lifecycle management, observability, reliability, and supportability across the HPC platform.
    • Drive cross-team design for integrations between Slurm, Kubernetes-adjacent systems, infrastructure APIs, identity systems, and platform tooling.
    • Create reusable modules, automation, deployment patterns, and reference implementations that increase engineering leverage.
    • Identify and correct avoidable technical divergence, duplicated effort, and fragile operating models.
    • Ensure domain designs reflect the realities of GPU scheduling, HPC networking, performance isolation, and production operations.
  • Delivery, Reliability & Influence
    • Lead technically critical initiatives spanning 2–4 teams or a defined HPC platform area.
    • Unblock delivery by clarifying technical direction and reducing ambiguity in complex system design problems.
    • Contribute hands-on where needed to de-risk or accelerate critical work.
    • Influence engineering teams without formal authority through strong judgement, design clarity, and practical solutions.
    • Partner with adjacent cloud-native software engineers so HPC implementations build on shared platform patterns rather than separate ones.

KPIs

  • Technical direction across a defined HPC domain
  • Delivery of critical initiatives across 2–4 teams
  • Reduction in technical divergence and duplicated effort
  • Reliability and supportability of Slurm-based HPC services

Qualifications

  • Extensive experience designing and building production software and automation for HPC systems, especially Slurm-based environments.
  • Strong track record of writing maintainable, testable, and resilient software in Go, Python, or similar languages.
  • Proven ability to define technical direction across a domain spanning multiple teams or services.
  • Strong understanding of Slurm internals, scheduler behaviour, cluster lifecycle concerns, and operational trade-offs.
  • Strong practical understanding of GPU-backed infrastructure and HPC networking, including InfiniBand, RoCE, RDMA, and performance-sensitive workload characteristics.
  • Experience integrating HPC systems with cloud-native platforms, APIs, or service delivery models.
  • Experience creating engineering leverage through standards, reusable patterns, shared tooling, and architectural clarity.
  • Strong judgement in balancing short-term delivery with long-term platform health and supportability.
  • Strong written and verbal communication skills, with the ability to align multiple teams around a coherent technical direction.
  • Experience with other schedulers or batch systems such as Kueue is valuable.

Benefits

  • Highly competitive US compensation package (base + bonus + equity), with performance reviews every 12 months.
  • Join one of the fastest-growing AI infrastructure companies — your chance to directly shape how global AI capacity is planned and deployed.
  • Expect a dynamic progression plan tailored to your ambitions. Grow by leading critical cross-functional initiatives and shaping capital strategy — always with our full support.
  • Human-First Flexibility: We treat you as humans first. Our flexible workplace trusts Nscalers to deliver, giving you the autonomy to shape your day around life's moments.
Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Staff HPC Systems Software Engineer in Remote vacancy
  •  ...Windows Server operating systems, Windows Client...  ...approval. Implement software solutions for multiple...  ...and next-generation HPE HPC products. Ensure development...  ...test execution to test engineers at various global locations...  ...to less- experienced staff members. Provides... 
    Suggested
    Local area
    Remote work

    Net2Source

    San Jose, CA
    1 day ago
  • $184k - $287.5k

     ...lasting impact on the world.We are looking for a dedicated engineer for the Senior Systems Software Engineer role, focusing on GPU Performance at Scale. At...  ...and develop new, leading solutions. Engage with HPC, OS, CPU, GPU compute, and systems specialists to architect... 
    Suggested
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...NVIDIA is searching for a highly motivated, technical engineer to join the Tegra system-on-chip (SoC) software organization. You will work on key aspects of our...  ...Familiarity with CUDA programming and/or GPUs.Experience with HPC or large-scale computing environments.Your base... 
    Suggested
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $211k - $368k

     ...technology pathfinding, AI and system workload analysis, memory...  ...validation, and teamwork across engineering organizations and external...  ...issues across hardware, firmware, software, networking, power, thermal,...  ...AI, machine learning, HPC workloads, GPU platforms, or... 
    Suggested
    Full time
    Local area
    Immediate start
    Remote work

    Micron Technology

    San Jose, CA
    2 days ago
  • Calance is seeking an Infrastructure Deployment Engineer to support planning, physical deployment, commissioning, and documentation of infrastructure across HPC and data center environments. This role requires hands-on rack & stack, structured cabling, and coordination... 
    Suggested
    Remote job

    Calance

    New York, NY
    1 day ago
  •  ...quantitative trading firm that is continuing to scale its global HPC and data center infrastructure and is looking for an...  ...status, risks and resource requirements Working closely with engineering teams to resolve technical and operational blockers Using AI and... 
    For contractors

    Autonomai Recruitment

    Chicago, IL
    2 days ago
  • $102k - $125k

     ...convenience and exceptional service to our members. Job Title Host Systems Engineer I Position Details Status: Exempt Reports to: Mgr - IT...  ...documented procedures; escalating complex issues to senior staff when required. Monitor system performance, availability, and... 
    Remote job

    Golden 1 Credit Union

    Sacramento, CA
    1 day ago
  • Hydra Host is seeking a High Performance Compute Solutions Engineer to join their team. This role will report to the Co-Founder & CTO and focus on managing GPU clusters while collaborating with clients to optimize distributed computing resources. Ideal candidates will... 
    Remote job

    Hydra Host

    New York, NY
    1 day ago
  • $108.8k - $163.2k

     ...opportunities to work on revolutionary systems that impact people's lives around the world...  ...Grumman Space Systems, you will engineer the enduring icons of modern space exploration...  ..., IL is seeking an experienced Embedded Software Engineer for its Software & Controls Department... 
    Full time
    Remote work
    Relocation package
    Shift work

    Northrop Grumman

    Rolling Meadows, IL
    5 days ago
  • $152.1k - $190.1k

    Role Description The Staff Forward Deployed Solutions Engineer will work directly within a business domain (e.g....  ...unstructured data flows across the systems involved (CRM, ERP, ticketing, document...  ..., business professionals, software engineers and many other professionals... 
    Full time
    Immediate start

    Natera

    Remote
    a month ago
  • $184k - $287.5k

     ...brings together cutting‑edge hardware and software innovation to deliver industry‑leading...  ...workloads. We are a group of forward‑thinking engineers tackling some of the globe’s toughest...  ...of lives. We’re searching for a Senior Systems Software Engineer with deep expertise in... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    5 days ago
  • $145k - $185k

     ...'s how we do it. DRW is a place of high expectations, integrity, innovation and a willingness to challenge consensus.As a Systems Software Engineer on our Network Capture Services team, you will help build and operate the firm's network packet capture platform. The data... 
    Temporary work
    Remote work
    Flexible hours
    Night shift
    Weekend work

    DRW

    Chicago, IL
    3 days ago
  • $184k - $287.5k

     ...innovations are revolutionizing self-driving cars, machine learning, supercomputing, gaming, and visualization. As a Senior System Software Engineer on the NvSci team, you will play an integral role in crafting NVIDIA's leadership in AI.What you'll be doing:Build and... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

    NVIDIA is growing a senior engineering team focused on making our compute software stack first-class on NVIDIA CPU platforms. The team turns modern toolchains...  ...software components.We are looking for an experienced systems software engineer who can lead cross-stack... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

     ...technologies powering GeForce NOW are advancing the future of AR/VR, AI inference, and connected mobility.We're looking for a Senior Systems Software Engineer with strong C++ experience and familiarity with streaming technologies. You will help to improve the quality of... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

     ...tools to debug, profile and analyze the performance of their systems/applications using the low-level libraries that you helped to...  ...generation accelerated computing at datacenter scale.As a system software engineer in the Developer Tools group, you will be developing software... 
    Full time
    Remote work

    Nvidia

    Austin, TX
    5 days ago
  • $86.8k - $198k

    Space System Software EngineerThe Opportunity: As an embedded software engineer, you can resolve a problem with a complete end-to-end solution in a fast, agile environment. If you’re looking for the chance to not just develop software, but to help create a system that will... 
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    El Segundo, CA
    1 day ago
  • $197.4k - $271.2k

     ...offering from top to bottom, and as a Senior Engineer you’ll do everything from working on the...  ...to solve complex distributed systems problems at scale. You will build services...  ...in production with a solid grasp on good software engineering practices such as code reviews... 
    Local area
    Remote work

    Confluent

    New York, NY
    4 days ago
  • $120k - $130k

     ...Systems Software Engineer Company: Picarro Location: Santa Clara, CA (Onsite) Education: Bachelor’s Degree Required Position Overview Picarro is seeking a Systems Software Engineer to design, develop, and maintain robust software systems that support... 
    Full time
    Temporary work
    Summer holiday
    Worldwide
    Flexible hours

    Picarro, Inc

    Santa Clara, CA
    4 days ago
  • $120k - $250k

     ...run as efficiently as allowed by physics, bringing the world years ahead in AI quality and availability. MatX is seeking System Software Engineer to join our team as we create best-in-class silicon for high-performance and sustainable GenAI. Successful candidates for... 
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    Monday to Friday
    Flexible hours
    3 days per week

    MatX

    Mountain View, CA
    4 days ago
  • $100k - $145k

     ...management, mission-critical building systems, energy efficiency, and...  ...Stack Forging™ process, we engineer high-performance thermal components...  ...direct deployment into AI, HPC, and other demanding environments...  ..., technology, technical data, software, or other items subject to U.S... 
    Permanent employment
    Full time
    Remote work
    Relocation
    1 day per week

    Johnson Controls

    Burlington, MA
    2 days ago
  • $140k - $193k

     ...About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance...  ..., platform capabilities, and software orchestration stack so that you can confidently...  ...and partners Translate customer AI, HPC, and infrastructure requirements into... 
    Full time
    Flexible hours

    Nscale

    Remote
    2 days ago
  • $112.7k - $193.2k

     ...that help people live healthier lives and help make the health system work better for everyone. From advanced data analytics and AI...  ...together.As an Experienced z System ISV Infrastructure Software Engineer within Optum Technology's Z Systems Application Hosting team,... 
    Minimum wage
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    Weekend work

    UnitedHealth Group

    Eden Prairie, MN
    1 day ago
  • $107.5k - $204.5k

     ...leader in the design, manufacture and service of aircraft engines and auxiliary power systems and has been revolutionizing modern flight for over 100...  ...including: requirements analysis & definition, system & software architecture, algorithm design, and product software... 
    Permanent employment
    Full time
    Temporary work
    Work experience placement
    Work at office
    Remote work
    Relocation package
    Flexible hours

    RTX

    Jupiter, FL
    3 days ago
  •  ...Focused on developing and optimizing system software for next-generation O-RAN infrastructure, the full-time Senior System Software Engineer will work remotely or onsite to enhance performance and energy efficiency across SoC architecture, cellular systems, and firmware... 
    Full time
    Remote work

    Virtual Vocations Inc

    United States
    5 days ago
  • $75 - $80 per hour

     ...bridging research ideas with production-quality engineering solutions and driving innovation that secures vital systems through advanced analysis technologies. In this...  .... Strong professional experience developing software in C/C++. Hands-on experience with compiler... 
    Hourly pay
    Temporary work

    Skill Corp

    Redmond, WA
    4 days ago
  •  ...team member who will help us to improve our auth and storage systems, creating new microservices, improving existing ones, doing regression...  ...we need to see: ~ Degree in Computer Science, Computer Engineering, or closely related field ~5+ years of relevant experience,... 
    Work at office
    Remote work

    NVIDIA

    United States
    7 hours ago
  • Technology: HPE accepting resumes for Syss/Softw Engr II in Roseville, CA (Ref. #9777044). Designs limited enhancements, updates, & progg changes for portions & subsyss of syss softw, incl operating syss, compliers, networking, utilities, dbases, & Internet-rel tools. ...
    Remote work

    HPE

    Roseville, CA
    2 days ago
  • $132.4k - $251.6k

     ...than 100 years of experience and renowned engineering expertise to meet the needs of today’s...  ...DoDevelop antenna, radome, advanced sensor systems, and/or RCS ranges/chambersWork with interdisciplinary...  ...with high performance computing (HPC) environments and schedulers such as... 
    Temporary work
    Work experience placement
    Interim role
    Work at office
    Remote work
    Relocation
    Flexible hours

    Raytheon

    Tucson, AZ
    1 day ago
  • $86.8k - $198k

    Agentic Systems Full-Stack Software Engineer, SeniorThe Opportunity: Most engineering roles ask you to write software. This one asks you to design it, direct it, and stand behind it. As a full-stack software engineer, you can take a problem from vision to production-ready... 
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Fayetteville, NC
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff HPC Systems Software Engineer. Be the first to apply!