Staff HPC Systems Software Engineer
Nscale
Role Description
We’re hiring a Staff HPC Systems Software Engineer to define the technical direction and evolution of a core HPC platform domain at Nscale.
In this role, you will operate beyond a single team, shaping how multiple teams build, automate, and run Slurm-based capabilities within Nscale’s wider cloud-native platform. You’ll work across engineering boundaries to bring coherence to architecture, interfaces, lifecycle models, and operational approaches, while partnering closely with teams working on platform tooling, infrastructure APIs, identity systems, and Kubernetes-adjacent systems.
This is a high-impact staff-level role for someone who combines deep hands-on software engineering with strong systems judgement. Your work will help ensure Nscale’s HPC services are robust, supportable, and maintainable, while creating leverage through shared patterns, reusable implementations, and clear technical direction across ambiguous, business-critical problem spaces.
What you'll be doing
- Domain Architecture & Technical Direction
- Own and evolve the technical direction for a defined HPC systems domain, such as Slurm platform architecture, scheduler integrations, cluster lifecycle, workload environments, or service automation.
- Make architectural decisions that balance software quality, operational realities, customer needs, and long-term maintainability.
- Define how proven Slurm implementations should be packaged, automated, and exposed as a service.
- Resolve ambiguity around ownership, interfaces, lifecycle boundaries, and operating models across teams.
- Act as the technical escalation point for the most complex issues within the domain.
- Cross-Team Engineering Leverage
- Establish shared patterns and standards for automation, service lifecycle management, observability, reliability, and supportability across the HPC platform.
- Drive cross-team design for integrations between Slurm, Kubernetes-adjacent systems, infrastructure APIs, identity systems, and platform tooling.
- Create reusable modules, automation, deployment patterns, and reference implementations that increase engineering leverage.
- Identify and correct avoidable technical divergence, duplicated effort, and fragile operating models.
- Ensure domain designs reflect the realities of GPU scheduling, HPC networking, performance isolation, and production operations.
- Delivery, Reliability & Influence
- Lead technically critical initiatives spanning 2–4 teams or a defined HPC platform area.
- Unblock delivery by clarifying technical direction and reducing ambiguity in complex system design problems.
- Contribute hands-on where needed to de-risk or accelerate critical work.
- Influence engineering teams without formal authority through strong judgement, design clarity, and practical solutions.
- Partner with adjacent cloud-native software engineers so HPC implementations build on shared platform patterns rather than separate ones.
KPIs
- Technical direction across a defined HPC domain
- Delivery of critical initiatives across 2–4 teams
- Reduction in technical divergence and duplicated effort
- Reliability and supportability of Slurm-based HPC services
Qualifications
- Extensive experience designing and building production software and automation for HPC systems, especially Slurm-based environments.
- Strong track record of writing maintainable, testable, and resilient software in Go, Python, or similar languages.
- Proven ability to define technical direction across a domain spanning multiple teams or services.
- Strong understanding of Slurm internals, scheduler behaviour, cluster lifecycle concerns, and operational trade-offs.
- Strong practical understanding of GPU-backed infrastructure and HPC networking, including InfiniBand, RoCE, RDMA, and performance-sensitive workload characteristics.
- Experience integrating HPC systems with cloud-native platforms, APIs, or service delivery models.
- Experience creating engineering leverage through standards, reusable patterns, shared tooling, and architectural clarity.
- Strong judgement in balancing short-term delivery with long-term platform health and supportability.
- Strong written and verbal communication skills, with the ability to align multiple teams around a coherent technical direction.
- Experience with other schedulers or batch systems such as Kueue is valuable.
Benefits
- Highly competitive US compensation package (base + bonus + equity), with performance reviews every 12 months.
- Join one of the fastest-growing AI infrastructure companies — your chance to directly shape how global AI capacity is planned and deployed.
- Expect a dynamic progression plan tailored to your ambitions. Grow by leading critical cross-functional initiatives and shaping capital strategy — always with our full support.
- Human-First Flexibility: We treat you as humans first. Our flexible workplace trusts Nscalers to deliver, giving you the autonomy to shape your day around life's moments.
- ...Windows Server operating systems, Windows Client... ...approval. Implement software solutions for multiple... ...and next-generation HPE HPC products. Ensure development... ...test execution to test engineers at various global locations... ...to less- experienced staff members. Provides...SuggestedLocal areaRemote work
$184k - $287.5k
...lasting impact on the world.We are looking for a dedicated engineer for the Senior Systems Software Engineer role, focusing on GPU Performance at Scale. At... ...and develop new, leading solutions. Engage with HPC, OS, CPU, GPU compute, and systems specialists to architect...SuggestedFull timeRemote work$184k - $287.5k
...NVIDIA is searching for a highly motivated, technical engineer to join the Tegra system-on-chip (SoC) software organization. You will work on key aspects of our... ...Familiarity with CUDA programming and/or GPUs.Experience with HPC or large-scale computing environments.Your base...SuggestedFull timeRemote work$211k - $368k
...technology pathfinding, AI and system workload analysis, memory... ...validation, and teamwork across engineering organizations and external... ...issues across hardware, firmware, software, networking, power, thermal,... ...AI, machine learning, HPC workloads, GPU platforms, or...SuggestedFull timeLocal areaImmediate startRemote work- Calance is seeking an Infrastructure Deployment Engineer to support planning, physical deployment, commissioning, and documentation of infrastructure across HPC and data center environments. This role requires hands-on rack & stack, structured cabling, and coordination...SuggestedRemote job
- ...quantitative trading firm that is continuing to scale its global HPC and data center infrastructure and is looking for an... ...status, risks and resource requirements Working closely with engineering teams to resolve technical and operational blockers Using AI and...For contractors
$102k - $125k
...convenience and exceptional service to our members. Job Title Host Systems Engineer I Position Details Status: Exempt Reports to: Mgr - IT... ...documented procedures; escalating complex issues to senior staff when required. Monitor system performance, availability, and...Remote job- Hydra Host is seeking a High Performance Compute Solutions Engineer to join their team. This role will report to the Co-Founder & CTO and focus on managing GPU clusters while collaborating with clients to optimize distributed computing resources. Ideal candidates will...Remote job
$108.8k - $163.2k
...opportunities to work on revolutionary systems that impact people's lives around the world... ...Grumman Space Systems, you will engineer the enduring icons of modern space exploration... ..., IL is seeking an experienced Embedded Software Engineer for its Software & Controls Department...Full timeRemote workRelocation packageShift work$152.1k - $190.1k
Role Description The Staff Forward Deployed Solutions Engineer will work directly within a business domain (e.g.... ...unstructured data flows across the systems involved (CRM, ERP, ticketing, document... ..., business professionals, software engineers and many other professionals...Full timeImmediate start$184k - $287.5k
...brings together cutting‑edge hardware and software innovation to deliver industry‑leading... ...workloads. We are a group of forward‑thinking engineers tackling some of the globe’s toughest... ...of lives. We’re searching for a Senior Systems Software Engineer with deep expertise in...Full timeRemote work$145k - $185k
...'s how we do it. DRW is a place of high expectations, integrity, innovation and a willingness to challenge consensus.As a Systems Software Engineer on our Network Capture Services team, you will help build and operate the firm's network packet capture platform. The data...Temporary workRemote workFlexible hoursNight shiftWeekend work$184k - $287.5k
...innovations are revolutionizing self-driving cars, machine learning, supercomputing, gaming, and visualization. As a Senior System Software Engineer on the NvSci team, you will play an integral role in crafting NVIDIA's leadership in AI.What you'll be doing:Build and...Full timeRemote work$184k - $287.5k
NVIDIA is growing a senior engineering team focused on making our compute software stack first-class on NVIDIA CPU platforms. The team turns modern toolchains... ...software components.We are looking for an experienced systems software engineer who can lead cross-stack...Full timeRemote work$152k - $241.5k
...technologies powering GeForce NOW are advancing the future of AR/VR, AI inference, and connected mobility.We're looking for a Senior Systems Software Engineer with strong C++ experience and familiarity with streaming technologies. You will help to improve the quality of...Full timeRemote work$152k - $241.5k
...tools to debug, profile and analyze the performance of their systems/applications using the low-level libraries that you helped to... ...generation accelerated computing at datacenter scale.As a system software engineer in the Developer Tools group, you will be developing software...Full timeRemote work$86.8k - $198k
Space System Software EngineerThe Opportunity: As an embedded software engineer, you can resolve a problem with a complete end-to-end solution in a fast, agile environment. If you’re looking for the chance to not just develop software, but to help create a system that will...Full timeContract workPart timeWork at officeLocal areaRemote work$197.4k - $271.2k
...offering from top to bottom, and as a Senior Engineer you’ll do everything from working on the... ...to solve complex distributed systems problems at scale. You will build services... ...in production with a solid grasp on good software engineering practices such as code reviews...Local areaRemote work$120k - $130k
...Systems Software Engineer Company: Picarro Location: Santa Clara, CA (Onsite) Education: Bachelor’s Degree Required Position Overview Picarro is seeking a Systems Software Engineer to design, develop, and maintain robust software systems that support...Full timeTemporary workSummer holidayWorldwideFlexible hours$120k - $250k
...run as efficiently as allowed by physics, bringing the world years ahead in AI quality and availability. MatX is seeking System Software Engineer to join our team as we create best-in-class silicon for high-performance and sustainable GenAI. Successful candidates for...Full timeWork experience placementWork at officeLocal areaRemote workMonday to FridayFlexible hours3 days per week$100k - $145k
...management, mission-critical building systems, energy efficiency, and... ...Stack Forging™ process, we engineer high-performance thermal components... ...direct deployment into AI, HPC, and other demanding environments... ..., technology, technical data, software, or other items subject to U.S...Permanent employmentFull timeRemote workRelocation1 day per week$140k - $193k
...About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance... ..., platform capabilities, and software orchestration stack so that you can confidently... ...and partners Translate customer AI, HPC, and infrastructure requirements into...Full timeFlexible hours$112.7k - $193.2k
...that help people live healthier lives and help make the health system work better for everyone. From advanced data analytics and AI... ...together.As an Experienced z System ISV Infrastructure Software Engineer within Optum Technology's Z Systems Application Hosting team,...Minimum wageFull timeWork experience placementWork at officeLocal areaRemote workWeekend work$107.5k - $204.5k
...leader in the design, manufacture and service of aircraft engines and auxiliary power systems and has been revolutionizing modern flight for over 100... ...including: requirements analysis & definition, system & software architecture, algorithm design, and product software...Permanent employmentFull timeTemporary workWork experience placementWork at officeRemote workRelocation packageFlexible hours- ...Focused on developing and optimizing system software for next-generation O-RAN infrastructure, the full-time Senior System Software Engineer will work remotely or onsite to enhance performance and energy efficiency across SoC architecture, cellular systems, and firmware...Full timeRemote work
$75 - $80 per hour
...bridging research ideas with production-quality engineering solutions and driving innovation that secures vital systems through advanced analysis technologies. In this... .... Strong professional experience developing software in C/C++. Hands-on experience with compiler...Hourly payTemporary work- ...team member who will help us to improve our auth and storage systems, creating new microservices, improving existing ones, doing regression... ...we need to see: ~ Degree in Computer Science, Computer Engineering, or closely related field ~5+ years of relevant experience,...Work at officeRemote work
- Technology: HPE accepting resumes for Syss/Softw Engr II in Roseville, CA (Ref. #9777044). Designs limited enhancements, updates, & progg changes for portions & subsyss of syss softw, incl operating syss, compliers, networking, utilities, dbases, & Internet-rel tools. ...Remote work
$132.4k - $251.6k
...than 100 years of experience and renowned engineering expertise to meet the needs of today’s... ...DoDevelop antenna, radome, advanced sensor systems, and/or RCS ranges/chambersWork with interdisciplinary... ...with high performance computing (HPC) environments and schedulers such as...Temporary workWork experience placementInterim roleWork at officeRemote workRelocationFlexible hours$86.8k - $198k
Agentic Systems Full-Stack Software Engineer, SeniorThe Opportunity: Most engineering roles ask you to write software. This one asks you to design it, direct it, and stand behind it. As a full-stack software engineer, you can take a problem from vision to production-ready...Full timeContract workPart timeWork at officeLocal areaRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff HPC Systems Software Engineer. Be the first to apply!


