HPC/GPU Systems Engineer
$100k - $140kNscale
About Nscale
Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack — energy, data centres, GPU superclusters, orchestration, and AI services — delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world.
At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability, and delivering rapid response to customer issues. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future.
About the Role
Infrastructure Support Engineers are the delivery engine of Infrastructure Support (L2/L3), handling the day-to-day health of Nscale's GPU fleets — tickets, alerts, hardware faults, and customer issues — across GPU nodes, high-performance networks, Linux, and data centre operations. This is a hands-on technical role: you'll come in with a strong technical base and working knowledge of GPU infrastructure, and grow toward Senior through exposure to some of the most advanced AI infrastructure in the world.
You will:
- Own your tickets and tasks end-to-end, escalating early and appropriately when issues exceed your scope — with clean, evidence-rich handovers.
- Communicate technical detail clearly, specifically, and concisely — in tickets, to customers, and to colleagues. We treat communication quality as a core engineering skill, not a soft skill.
- Grasp new technical concepts and problems quickly; stay curious and know what questions to ask to get up to speed fast.
- Bring discipline and organisation: accurate records, structured troubleshooting, reliable follow-through.
- Seek feedback and invest in learning — this role is a deliberate pathway to Senior.
Experience required: 3-4+ years in infrastructure support or support engineering roles, including deep support/service desk experience in structured, customer-facing environments, with working knowledge of GPU infrastructure and hands-on hardware troubleshooting.
What You’ll Be Doing
- Join the Support duty rotation and handle day-to-day tickets and alerts, escalating early and appropriately. Collaborate with Engineering, with guidance, when incidents or changes require it.
- Perform GPU node triage and hardware troubleshooting: interpret nvidia-smi/DCGM output and system logs, isolate faults across GPU, NIC, and server hardware, carry out physical remediation (reseats, swap testing, component checks), and prepare clean evidence for vendor RMA.
- Run fabric and link diagnostics following established runbooks (mlxlink or equivalent), capture evidence accurately, and escalate with a handover that lets the next engineer continue without starting from scratch.
- Assist with storage and data-path investigations (mounts, connectivity, client-side symptoms) on high-performance platforms, gathering evidence for Senior or Engineering-led diagnosis.
- Follow established runbooks to resolve common issues; propose improvements and contribute incremental fixes with review.
- Accurately record, update, manage, and resolve tickets, keeping all parties informed with clear notes, next steps, and customer communications via the agreed channels.
- Participate in monitoring, troubleshooting, and triage. Capture logs and facts to enable efficient handover.
- Participate in changes under peer review, learning risk assessment and backout practices in live customer environments.
- Help maintain source-of-truth accuracy across DCIM, inventory, and asset records (NetBox or similar patterns).
- Identify opportunities for automation and contribute simple scripts and tooling improvements to optimise processes.
- Be the escalation point for onsite DC Operations staff; coordinate smart-hands tasks within your scope.
- Learn the Platform fundamentals so you can help customers get value from our services, asking for support when deeper expertise is needed.
- Share knowledge by documenting steps you've validated and contributing to training materials. Shadow Seniors during complex work to build capability.
- Take part in incident reviews as a contributor and help track preventative follow-ups in your scope.
- Deliver assigned tasks and project work to agreed quality and timelines. Flag blockers early and seek help when needed.
- Participate in on-call and out-of-hours work when scheduled and after onboarding. Travel to Nscale or customer locations to assist with deployments, troubleshooting, and operational tasks, and attend supplier training as required.
About You
- Experience. 3+ years in infrastructure support or support engineering, including support/service desk experience in structured, SLA-driven, customer-facing environments (cloud, data centre, or managed services).
- Communication. Clear written notes, concise updates, and reliable follow-through. Able to explain technical issues accurately to customers and colleagues, and produce handovers the next shift can act on immediately.
- GPU and hardware troubleshooting. Working knowledge of GPU infrastructure: hands-on with nvidia-smi or similar diagnostics, comfortable interpreting hardware error output and logs, and confident physically troubleshooting servers — reseating components, swap testing, working via BMC/out-of-band management — through to preparing RMA evidence. A strong technical base here is required, not a learning goal.
- Linux. Solid working knowledge: confident on the CLI with systemd, filesystems, permissions, and standard networking tools. Able to troubleshoot common issues independently and know when to escalate.
- Networking. Solid grasp of IP addressing, subnets, VLANs, routing, DNS, and firewalls. Awareness of high-performance east-west fabrics (RDMA/InfiniBand concepts) is a plus and a core growth area in this role.
- Ticketing and ITSM discipline. Experience working within structured support processes (ITIL or similar): prioritisation, escalation, SLA awareness, and accurate documentation.
- Observability foundations. Able to use dashboards and alerts to identify symptoms, gather evidence, and follow runbooks. Comfortable proposing simple alert or dashboard improvements with review.
- Scripting and automation basics. Comfortable reading and writing simple Bash or Python, and using Git for version control.
- Platform and DC fundamentals. Understanding of servers, networks, storage, and virtualisation concepts, ideally from a support or operations background.
- Growth mindset. Curious, dependable, and collaborative. You seek feedback, ask questions, and invest in learning to progress toward Senior.
- Adaptability. Able to work in a fast-moving environment with evolving processes, participate in on-call after onboarding, and travel when needed.
Nice to Have
These are growth areas, not prerequisites:
- High-performance fabrics and GPU-HPC: exposure to RDMA/InfiniBand, link-level diagnostics (mlxlink, ibdiagnet, or equivalent), NCCL-based troubleshooting, or NVLink concepts.
- High-performance storage: exposure to VAST or comparable AI-optimised storage platforms, Ceph, or NFS at scale, including basic storage–network troubleshooting.
- OpenStack and fleet operations tooling: familiarity with OpenStack troubleshooting flows, or fleet-scale tooling for provisioning and health (MAAS, NetBox, Redfish, or similar).
- Kubernetes: understanding of core concepts (nodes, pods, services, logs) and basic troubleshooting via runbooks. Helpful context for our platform, though not the core of this role.
- Automation and access tooling: experience with Ansible or Terraform, CI/CD participation (GitHub Actions or similar), or access and security tooling such as Teleport or Vault.
- Certifications: progress toward relevant Linux, networking, Kubernetes, cloud, or security certifications over time.
What We Can Offer You
At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something extraordinary, and we want you at the core.
- Highly competitive package (base + equity) with reviews every 12 months.
- Join the fastest-growing tech startup, your chance to push boundaries, collaborate with brilliant minds, and make your mark on cutting-edge AI. ✨
- Expect a dynamic progression plan tailored to your ambitions. Grow by trying new things, leading, challenging the status quo, and owning your impact, always with our full support.
- Human-First Flexibility: We treat you as humans first. Our flexible workplace trusts Nscalers to deliver, giving you the autonomy to shape your day around life's moments.
- Join our thriving remote-first team. Geography is no barrier to impact or connection. We build seamless virtual collaboration, empowering you, wherever you work.
Equal Opportunities Statement
We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.
If there’s anything we can do to accommodate your specific situation, please let us know.
The responsibilities outlined in this job description are not exhaustive and are intended to provide a general overview of the position. The employee may be required to perform additional duties, tasks, and responsibilities as assigned by management, consistent with the skills and qualifications required for the role.
The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.
Salary Range
$100,000—$140,000 USD
For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here. Nscale does not accept unsolicited candidate submissions from recruitment agencies.
$100k - $150k
...vertically integrated AI cloud engineered for AI. We own and operate the... ...stack — energy, data centres, GPU superclusters, orchestration,... ...2–3+ years hands-on with GPU, HPC, or large-scale data centre estates... ...job failures. ~ Linux systems engineering at scale....SuggestedFull timeRemote workFlexible hours$250k
...provider building a next-generation GPU platform designed for AI training,... ...a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments powering GPU... ...deliver highly available infrastructure systems Improve CI/CD pipelines,...SuggestedFull timeRemote work$230k
...research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more,... ...over unchecked growth.About the roleAs a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability...SuggestedWork at officeLocal areaFlexible hours$215k - $260k
...a Hardware Production / Sustaining Engineer to strengthen Crusoe’s Hardware Systems Engineering team and close critical... ...reliability across Crusoe Cloud’s GPU- and CPU-based infrastructure.You will... ...and how to leverage them in AI/HPC environments.Expertise supporting or...Suggested$180k - $240k
...About the Role We are looking for a Senior Systems Developer to lead the design, development... ...ongoing operation Mentor and guide engineers, raising the technical bar across the... ...environments Familiarity with server and GPU hardware architecture and system-level...SuggestedFull timeFlexible hours- Vision Systems Engineer Exploration Technology Group San Francisco, California, United States About this position etc, the Exploration Technology... ...suppression, track‑before‑detect — on embedded FPGA and GPU platforms. Performance Modeling & Test: Build end‑to‑end radiometric...Flexible hours
- ...households. The Role Our robots run real-time control, GPU inference, SLAM, camera pipelines, teleoperation, and... ...boards, MCUs, and the sensor architecture around them. As a Systems Software Engineer focused on the OS, you'll own the OS layer and the IPC layer...Remote work
- ...the world. The Production Engineering Team Examples of key problems... ...-scale compute customers (HPC, cloud, or AI labs) at a technical... .... You debug distributed systems methodically across layers you... ...follow up until they do. Bonus: GPU training workloads. InfiniBand...
- ...discover, and create.We’re seeking a Principal Machine Learning Systems Engineer (P60) to lead technical directions of GenAI Products &... ...HaveBackground in distributed systems, high-performance computing, or GPU optimization.Familiarity with search/GenAI evaluation metrics...Work at officeLocal area
- ...boundaries of what's possible in video generation.We're seeking a GPU Performance Engineer to squeeze every last FLOP from our H100 infrastructure and... ...and optimize GPU workloads using Nsight Systems, nvprof, and custom instrumentationWrite high-performance CUDA...
$174.5k - $240k
...join us as we power the shop local movement. If you believe in community, come join ours.About this roleGTM Engineering builds and operates the intelligent systems, integrations, and automations that power Faire's GTM revenue org. We're the AI and technical backbone of...Work experience placementWork at officeLocal areaRemote workMonday to FridayFlexible hours3 days per week$160k - $185k
...sensors for precision mapping, robotics, automotive, security systems, smart cities and various industrial solutions. We've transformed... ...need your help!The Role:We're seeking a proactive and skilled engineer who thrives in a dynamic, fast-paced environment to be an integral...Work experience placementLocal area$10 per hour
...challenges that impact business, society, and the environment? Come join us.The Opportunity: Flexport IT is looking for a Senior Systems Engineer (Identity & Access). In this role, you will design, implement, and administer our Identity and Access Management (IAM)...Flexible hours$182.9k - $228.6k
...hardware design, manufacturing, data processing, and software engineering, our office is a truly inspiring mix of experts from a variety... ...and The Netherlands.About the Role:Planet seeks a Senior Camera Systems Engineer to serve as the primary system architect and...Full timeContract workTemporary workFor contractorsWork at officeLocal areaRemote workHome office3 days per week$165k - $200k
...strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.About This RoleAs a Senior Systems Engineer, you’ll play a key role in building and optimizing Crusoe’s global technology infrastructure. This is a hands-on, on-site role...Temporary work- We are looking for a Senior Systems Engineer to strengthen and advance a secure, high-performing technology environment in San Francisco, California. This position plays a central role in support the infrastructure strategy across cloud services, on-premises systems, and...
- ...get stuck into. And that’s where you come in.Job Description:Hitachi Rail is looking for an enthusiastic self-motivated Senior System Engineer who thrives in a fast-paced environment. The successful candidate is comfortable performing a wide range of tasks from administrative...Full time
$152k
About the roleWe are seeking an experienced Senior Systems Engineer to own our macOS platform and drive endpoint management and device trust standards across Chime’s IT ecosystem.As a Senior Systems Engineer, you will drive multi-system initiatives across our IT domains...Full timeWork at officeLocal areaRemote workShift work$190k - $230k
...part of a high-performing team that believes in each other, come build with us at Crusoe.About This RoleWe’re seeking a Senior Systems Engineer to play a key role in executing Crusoe’s 2026 Enterprise AI Strategy. In this role, you will design and build agentic AI...Temporary work- ...get stuck into. And that’s where you come in.Job Description:Hitachi Rail is looking for an enthusiastic self-motivated Senior System Engineer who thrives in a fast-paced environment. The position is based Brisbane, Australia.An exciting opportunity has opened for a TMS...Full time
$156k - $195k
...Franklin Templeton. For more information, visit or follow us on LinkedIn.About the role:Ironclad is hiring a Senior Finance Systems & AI Engineer to own the data infrastructure, system integrations, and AI automation that power our Finance teams.This role sits within the...Full timeContract work- ...Adapt is hiring an engineer to own the computer underneath Adapt. This role focuses on systems that power customer workloads, sits in customer conversations, and ships user-facing features. You will work with a stack including Firecracker, gVisor, GCP, and Kubernetes....
- ...LangChain in San Francisco seeks a systems/database engineer to design, optimize, and harden our distributed storage engine. You’ll work on ingestion, storage layout, and query execution on Kubernetes-based services. The role requires 5+ years in systems engineering...
- ...infrastructure. Help design, implement, and monitor testnets Required Skills: Expert knowledge of peer-to-peer distributed system design and implementation (required) Ability to build and maintain high available infrastructure (required) Knowledge on how...
$139k - $204k
...Systems Engineer, Legal SystemsLivingston, NJ / New York, NY / Sunnyvale, CA / San Francisco, CA / Bellevue, WA CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators...Permanent employmentFull timeContract workCasual workWork at office- ...Figma is seeking a Content Support Engineer to own knowledge systems and architecture, ensuring scalable, accurate content for support and AI systems. You will design automated workflows, govern content templates, and define metrics to measure impact on CSAT and first...Remote work
$230k
...Building an AI agent that can securely take action inside enterprise systems is hard. The moment an agent accesses customer data, executes... ...a user, authorization, governance, and trust become the real engineering challenge. Arcade is the MCP runtime that gives agents the...Work at officeShift work- ...data, run data applications, so they can spend more time putting knowledge into action. We're looking for engineers who want to build the operating system for AI Data Applications and Workflows. About the role We're looking for experienced distributed systems...
$210k - $245k
...something different. The future of design at a16z isn't a bigger team cranking out more assets: it's a design function that scales through systems. Templates, tools, AI‑powered workflows, and self‑serve infrastructure that let anyone at the firm create high‑quality, on‑brand...$250k
...A Series A Funded start-up in California is seeking a Systems Engineer to design and optimize systems handling complex ML pipelines. The role involves building scalable infrastructure, developing CI/CD pipelines, and ensuring system performance. Key qualifications include...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to HPC/GPU Systems Engineer. Be the first to apply!
- healthcare systems engineer San Francisco, CA
- mission system engineer San Francisco, CA
- senior linux systems engineer San Francisco, CA
- senior staff systems engineer San Francisco, CA
- data systems engineer San Francisco, CA
- system engineer remote San Francisco, CA
- systems engineer San Francisco, CA
- space systems engineer San Francisco, CA
- software system engineer San Francisco, CA
- director systems engineering San Francisco, CA


