Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior HPC/GPU Systems Engineer

$100k - $150k
Full-time

Nscale

About Nscale

Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack — energy, data centres, GPU superclusters, orchestration, and AI services — delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world.

At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability, and delivering rapid response to customer issues. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future.

About the Role (Job Purpose)

Senior Infrastructure Support Engineers are the senior technical escalation point within Infrastructure Support, owning the health of Nscale's GPU fleets and the high-performance fabrics that connect them. This is a hands-on L2/L3 role operating at the intersection of GPU hardware, east-west networking, Linux, and data centre operations — acting as the operational bridge between Support, DC Operations, and Engineering.

You will:

  • Own complex, ambiguous problems end-to-end and make decisive calls in a results-driven environment, taking calculated risks where speed matters.
  • Communicate technical detail clearly, specifically, and concisely — to engineers, to customers, and to leadership. We treat communication quality as a core engineering skill, not a soft skill.
  • Influence without authority and build strong relationships with senior stakeholders across the business to get things done.
  • Grasp new technical concepts quickly, stay curious, and know which questions to ask to get up to speed fast.
  • Bring discipline and organisation: evidence-led investigations, accurate records, clean handovers.

Experience required: 6+ years in infrastructure, operations, or support engineering roles in production environments, including 2–3+ years hands-on with GPU, HPC, or large-scale data centre estates.

What You’ll be Doing (Responsibilities)

  • Join the Support duty rotation as a senior escalation point, collaborating with Infrastructure Engineering, CNPRE, Network Operations, and Product Engineering on incidents, investigations, and changes.
  • Diagnose and remediate GPU node faults across the full stack — driver, firmware, and hardware layers — from nvidia-smi/DCGM and XID/RAS analysis through BMC/Redfish and out-of-band management to physical fault isolation and vendor RMA.
  • Own east-west fabric health: run link-level diagnostics (mlxlink, ibdiagnet, or equivalent), isolate transceiver, optics, cabling, and switch-port faults, and validate topology across InfiniBand and RoCE/high-speed Ethernet fabrics.
  • Investigate data-path issues on high-performance storage platforms (e.g. VAST), including storage–network interactions across clients, mounts, VIPs, and routing.
  • Run structured, hypothesis-driven investigations; conduct root cause analysis for major incidents and drive long-term fixes to completion.
  • Author and execute changes in live customer environments with proper risk assessment, peer review, and backout plans.
  • Proactively improve dashboards, alerts, and runbooks to prevent repeat incidents; identify recurring patterns and convert them into problem records and automation.
  • Accurately record, update, and resolve tickets, keeping internal and external parties informed with clear customer-impact statements and evidence-rich notes that enable clean handover.
  • Design and implement automation scripts and small tools to reduce toil and human intervention.
  • Act as a key escalation point for the Support Organisation, taking ownership of strategic decisions where results matter.
  • Mentor and upskill mid-level engineers; contribute to knowledge sharing across Operations and Engineering, including training content, workshops, and PR reviews.
  • Lead by earning trust and speaking candidly. Disagree when appropriate and challenge the status quo; commit wholly to decisions once in motion.
  • Respond to critical incidents out of business hours and participate in on-call as required. Travel to Nscale or customer sites to provide onsite technical expertise.

About You (Skills / Qualifications Experience)

  • Experience.

6+ years in infrastructure, operations, or support engineering in production environments; 2–3+ years hands-on with GPU, HPC, or large-scale data centre estates, ideally in a customer-facing or escalation-driven capacity.

  • Communication.

Able to explain complex technical detail clearly, specifically, and concisely — in tickets, in incident updates, and face to face with customers and stakeholders at all levels. Strong written discipline: your notes let the next engineer pick up where you left off without starting from scratch.

  • GPU platforms (NVIDIA; AMD Instinct beneficial).

Practical, current experience with GPU drivers, firmware, and runtime stacks on AI training and inference clusters. Confident with nvidia-smi, DCGM, and XID/error interpretation; able to isolate faults across GPU, baseboard, NIC, and PCIe layers and drive them through diagnosis to RMA.

  • High-performance east-west fabrics.

Hands-on experience with RDMA fabrics — InfiniBand and/or RoCE — including link-layer diagnostics (mlxlink, ibdiagnet, or equivalent), transceiver and cabling fault isolation, and understanding of rail-optimised topologies, NVLink/NVSwitch, and NCCL-based performance troubleshooting on multi-node clusters.

  • HPC scheduling.

Slurm operations for large multi-GPU jobs — containers via Pyxis/Enroot, MPI, and diagnosing queue, topology, and job failures.

  • Linux systems engineering at scale.

Strong command of modern Linux distributions, kernel modules, systemd, networking stack, and filesystem tooling. Proven troubleshooting across compute, storage, and network layers in production.

  • Server hardware and control planes.

Comfortable with BMC/Redfish, firmware management, and bare-metal provisioning workflows (MAAS or similar) across large node fleets.

  • Networking fundamentals.

Solid grasp of L2/L3, routing, BGP, VLANs, VXLAN, firewalls, and load balancing, with a clear understanding of how east-west cluster traffic differs from north-south.

  • Observability and incident response.

Build and use alerting stacks and dashboards (Prometheus/Grafana or similar), interpret metrics and alerts, drive runbooks to resolution, and contribute to SLOs and post-incident reviews.

  • Change and risk judgment.

Experience authoring and executing changes in business-critical environments, including risk assessments, customer-impact analysis, and backout plans.

  • SRE-style operations.

Write and maintain runbooks, automate diagnostics, and reduce human intervention through scripts and small tools.

  • Automation and Git.

Scripting skills in Bash, Python, or equivalent for operational tooling and integrations; experience with infrastructure automation tools (Ansible, Terraform, or similar).

  • Data centre fundamentals.

Understanding of how data centres operate — servers, networks, storage, power, and cooling — ideally gained through an operational support background.

  • Leadership.

Disciplined, organised, and self-motivated, with the ability to mentor and motivate other engineers, take decisive action, and drive the team and wider organisation to improve.

  • Adaptability.

Able to adapt to customer-driven demands, including specialist support outside core hours and travel for onsite work.

Nice to Have

  • High-performance storage.

Hands-on experience with VAST or comparable AI-optimised storage platforms, or Ceph/parallel filesystems and NFS at scale (multipath, remoteports, nconnect), including diagnosing storage–network interaction and data-path performance issues.

  • OpenStack and fleet operations tooling.

OpenStack operations experience (Neutron, Cinder, error triage), plus familiarity with fleet-scale tooling for provisioning, health, and remediation across large GPU estates (MAAS, NetBox, Redfish-driven automation, or similar).

  • Kubernetes.

Operating and troubleshooting clusters, including GPU operator stacks and understanding how physical resources are abstracted up the stack. Helpful context for our platform, though not the core of this role.

  • Automation at scale.

Automated network configuration with safe, repeatable changes in business-critical environments; GitOps and CI/CD pipelines (GitHub Actions or similar); access and security tooling such as Teleport or Vault in production.

  • Certifications.

Relevant GPU/HPC, datacenter architecture, Linux, networking, Kubernetes, cloud, or security certifications (e.g. RHCSA/RHCE, CKA, NVIDIA-certified) are a plus.

What We Can Offer You

At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something extraordinary, and we want you at the core.

  • Highly competitive package, including base salary and equity, with reviews every 12 months.
  • Join one of the fastest-growing tech startups: your chance to push boundaries, collaborate with brilliant minds, and make your mark on cutting-edge AI.
  • Expect a dynamic progression plan tailored to your ambitions. Grow by trying new things, leading, challenging the status quo, and owning your impact, always with our full support.
  • Human-first flexibility. We treat you as humans first. Our flexible workplace trusts Nscalers to deliver, giving you the autonomy to shape your day around life's moments.
  • Join our thriving remote-first team. Geography is no barrier to impact or connection. We build seamless virtual collaboration, empowering you wherever you work.

Equal Opportunities Statement

At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enrich our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.

If there’s anything we can do to accommodate your specific situation, please let us know.

The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

Salary Range

$120,000—$170,000 USD

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

Nscale does not accept unsolicited candidate submissions from recruitment agencies.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior HPC/GPU Systems Engineer in San Francisco, CA vacancy
  • $100k - $140k

     ...vertically integrated AI cloud engineered for AI. We own and...  ...energy, data centres, GPU superclusters,...  ...infrastructure, and grow toward Senior through exposure to...  ...nvidia-smi/DCGM output and system logs, isolate faults...  ...performance fabrics and GPU-HPC: exposure to RDMA/... 
    Suggested
    Full time
    Immediate start
    Remote work
    Flexible hours
    Shift work

    Nscale

    San Francisco, CA
    2 days ago
  • $250k

     ...provider building a next-generation GPU platform designed for AI...  ...The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments powering...  ...highly available infrastructure systems Improve CI/CD pipelines,... 
    Senior
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  • $180k - $240k

     ...About the Role We are looking for a Senior Systems Developer to lead the design, development,...  ...ongoing operation Mentor and guide engineers, raising the technical bar across the team...  ...Familiarity with server and GPU hardware architecture and system-level optimization... 
    Senior
    Full time
    Flexible hours

    Nscale

    San Francisco, CA
    2 days ago
  • We are looking for a Senior Systems Engineer to strengthen and advance a secure, high-performing technology environment in San Francisco, California. This position plays a central role in support the infrastructure strategy across cloud services, on-premises systems, and... 
    Senior

    Robert Half

    San Francisco, CA
    1 day ago
  •  ...What We’re Looking For Strong experience building agent systems, LLM tool-calling pipelines, or orchestration frameworks Deep...  ...production, not just prototypes Work on complex, high-impact engineering workflows with real constraints High ownership and... 
    Senior

    Acceler8 Talent

    San Francisco, CA
    2 hours ago
  •  ...ServiceNow is seeking an experienced AI engineer to design, build, and operate production-grade agentic AI systems across our platform. You will focus on multi-agent orchestration, tool use, planning loops, memory, and safe failure recovery. You will push for robust... 
    Senior

    ServiceNow

    San Francisco, CA
    3 days ago
  • $200k - $300k

     ...A leading AI company based in San Francisco is seeking a Senior Neuro-Symbolic Systems Engineer to enhance AI systems interacting with the physical world. Candidates will design and implement complex representations, work with cross-functional teams, and bridge symbolic... 
    Senior

    Acceler8 Talent

    San Francisco, CA
    5 days ago
  •  ...to get stuck into. And that’s where you come in.Job Description:Hitachi Rail is looking for an enthusiastic self-motivated Senior System Engineer who thrives in a fast-paced environment. The successful candidate is comfortable performing a wide range of tasks from administrative... 
    Senior
    Full time

    Hitachi

    Brisbane, CA
    13 hours ago
  •  ...Adapt is hiring an engineer to own the computer underneath Adapt. This role focuses on systems that power customer workloads, sits in customer conversations, and ships user-facing features. You will work with a stack including Firecracker, gVisor, GCP, and Kubernetes.... 
    Senior

    ADAPT Corporation

    San Francisco, CA
    1 day ago
  •  ...to get stuck into. And that’s where you come in.Job Description:Hitachi Rail is looking for an enthusiastic self-motivated Senior System Engineer who thrives in a fast-paced environment. The position is based Brisbane, Australia.An exciting opportunity has opened for a... 
    Senior
    Full time

    Hitachi

    Brisbane, CA
    13 hours ago
  • $10 per hour

     ...challenges that impact business, society, and the environment? Come join us.The Opportunity: Flexport IT is looking for a Senior Systems Engineer (Identity & Access). In this role, you will design, implement, and administer our Identity and Access Management (IAM) solutions... 
    Senior
    Flexible hours

    Flexport

    San Francisco, CA
    1 day ago
  • $165k - $200k

     ...strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.About This RoleAs a Senior Systems Engineer, you’ll play a key role in building and optimizing Crusoe’s global technology infrastructure. This is a hands-on, on-site... 
    Senior
    Temporary work

    Crusoe

    San Francisco, CA
    4 days ago
  • $152k

    About the roleWe are seeking an experienced Senior Systems Engineer to own our macOS platform and drive endpoint management and device trust standards across Chime’s IT ecosystem.As a Senior Systems Engineer, you will drive multi-system initiatives across our IT domains... 
    Senior
    Full time
    Work at office
    Local area
    Remote work
    Shift work

    Chime

    San Francisco, CA
    11 hours ago
  • $190k - $230k

     ...be part of a high-performing team that believes in each other, come build with us at Crusoe.About This RoleWe’re seeking a Senior Systems Engineer to play a key role in executing Crusoe’s 2026 Enterprise AI Strategy. In this role, you will design and build agentic AI systems... 
    Senior
    Temporary work

    Crusoe

    San Francisco, CA
    4 days ago
  • $200k - $300k

     ...Base pay range: $200,000.00/yr - $300,000.00/yr Direct message the job poster from Acceler8 Talent Senior Neuro-Symbolic Systems Engineer - San Francisco, CA A company building AI systems that can interact with the physical world at scale – designing experiments, controlling... 
    Senior
    Full time
    Immediate start

    Acceler8 Talent

    San Francisco, CA
    13 hours ago
  • $160k - $200k

     ...Senior Systems Engineer Echo Neurotechnologies is an exciting new startup in the Brain-Computer Interface (BCI) space, driving innovation through advanced hardware engineering and AI solutions. Our mission is to deliver cutting-edge technologies that restore autonomy... 
    Senior

    Echo Neurotechnologies

    San Francisco, CA
    2 days ago
  • $182.9k - $228.6k

     ...hardware design, manufacturing, data processing, and software engineering, our office is a truly inspiring mix of experts from a...  ...Slovenia, and The Netherlands.About the Role:Planet seeks a Senior Camera Systems Engineer to serve as the primary system architect and technical... 
    Senior
    Full time
    Contract work
    Temporary work
    For contractors
    Work at office
    Local area
    Remote work
    Home office
    3 days per week

    Planet Labs

    San Francisco, CA
    13 hours ago
  • $156k - $195k

     ...BOND, and Franklin Templeton. For more information, visit or follow us on LinkedIn.About the role:Ironclad is hiring a Senior Finance Systems & AI Engineer to own the data infrastructure, system integrations, and AI automation that power our Finance teams.This role sits... 
    Senior
    Full time
    Contract work

    Ironclad Inc

    San Francisco, CA
    4 days ago
  • $135k - $145k

     ...Senior Systems EngineerLyra Technology Group is a private equity-backed holding company that invests in and operates industry leading technology...  ...term.People 1st IT is looking for an experienced Level 3 Engineer / Senior Systems Engineer to join our technical team. This is... 
    Senior
    Work at office
    Remote work
    Relocation

    Lyra Technology Group

    San Francisco, CA
    1 hour ago
  • $190k - $230k

     ...Senior Systems Engineer We are seeking a full-time Senior Systems Engineer to enhance the performance and efficiency of our deployments. You will be in charge of owning the entire end to end system level performance for our deployments. In this role, you will collaborate... 
    Senior
    Full time
    Immediate start

    Osaro, Inc.

    San Francisco, CA
    1 hour ago
  • $112k - $168k

     ...is responsible for providing technology systems, administration, and support to Klaviyos...  ...lifecycle management, building automation, and engineering solutions for both other internal teams...  ...department.About the role:As a Senior IT Systems Engineer on the IT Systems team... 
    Senior
    Work at office
    Shift work

    Klaviyo

    San Francisco, CA
    4 days ago
  •  ...We’re ruthlessly focused on business impact. We are a highly senior team composed of former pioneers from a variety of different...  ...even better.About the RoleWe are seeking a highly motivated Systems Engineer to help shape the architecture and integration of our next-generation... 
    Senior
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    13 hours ago
  • $160k - $180k

     ...been a “sleepy” industry for decades is now at the epicenter of sustaining the global economy. About the role: As a Senior Wireless RF Systems Engineer at Mytra, you will lead the hardware development of our next-generation wireless products. This role focuses on board... 
    Senior
    Work at office

    Mytra

    Brisbane, CA
    4 days ago
  • $165k - $200k

     ...build with deep respect for our end users, listening closely to their feedback and needs. About The Job Radar is hiring a Senior Systems Engineer with end-to-end ownership of system-level inventory features. You will lead the problem solving process, analyze test data,... 
    Senior
    Flexible hours

    RADAR

    San Francisco, CA
    2 days ago
  • $162.6k - $203.2k

     ...hardware design, manufacturing, data processing, and software engineering, our office is a truly inspiring mix of experts from a variety...  ...specifications between our customer and Planet’s teamsAssess system level budgets for engineering trades and feasibility studiesCollaborate... 
    Senior
    Full time
    Temporary work
    For contractors
    Work at office
    Local area
    Remote work
    Home office

    Planet Labs

    San Francisco, CA
    1 day ago
  •  ...and other priorities. We can hire people in any country where we have a legal entity. Responsibilities As a senior Machine Learning Systems Engineer on the Search Platform team, you will own and drive the design, development, and production deployment of machine... 
    Senior
    Work at office
    Local area

    Atlassian

    San Francisco, CA
    1 day ago
  • $179k - $218k

     ..."Silicon Reality" must be bridged.We are seeking a Senior Staff Data Center Operations Engineer, GPU Hardware Architecture to be the definitive technical...  ...production environment. Lead Root Cause Analysis (RCA) on systemic issues that span the boundary between hardware and... 
    Senior
    Temporary work

    Crusoe

    San Francisco, CA
    2 days ago
  • The Trade Desk is seeking a Senior Systems Engineer, focused on building high-impact analytical systems to support Sales and Client Services. This role involves designing, implementing, and maintaining systems that turn commercial data into actionable insights, improving... 
    Senior

    The Trade Desk

    San Francisco, CA
    4 days ago
  • $180k - $230k

    San Francisco, CAProduct Systems Engineering - Product Systems /Full time /On-siteWanna join the adventure?Loft Orbital builds a space infrastructure...  ...the lead-time and risk of a traditional space mission.As a Senior Product Systems Engineer at Loft, you will own the... 
    Senior
    Full time
    Temporary work

    Loft Orbital

    San Francisco, CA
    4 days ago
  • $151k - $230k

     ...requirements. - Be responsible for architectures and requirements sets for assigned system-level areas as part of Hardware Development.  - Work closely with stakeholders and domain engineers to transform product guidance into design through every stage of the product... 
    Senior
    Work at office

    Waabi

    San Francisco, CA
    a month ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior HPC/GPU Systems Engineer. Be the first to apply!