Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior HPC/GPU Systems Engineer

$120k - $170k
Full-time

Nscale

About Nscale

Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack — energy, data centres, GPU superclusters, orchestration, and AI services — delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world.

At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability, and delivering rapid response to customer issues. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future.

About the Role (Job Purpose)

Senior Infrastructure Support Engineers are the senior technical escalation point within Infrastructure Support, owning the health of Nscale's GPU fleets and the high-performance fabrics that connect them. This is a hands-on L2/L3 role operating at the intersection of GPU hardware, east-west networking, Linux, and data centre operations — acting as the operational bridge between Support, DC Operations, and Engineering.

You will:

  • Own complex, ambiguous problems end-to-end and make decisive calls in a results-driven environment, taking calculated risks where speed matters.
  • Communicate technical detail clearly, specifically, and concisely — to engineers, to customers, and to leadership. We treat communication quality as a core engineering skill, not a soft skill.
  • Influence without authority and build strong relationships with senior stakeholders across the business to get things done.
  • Grasp new technical concepts quickly, stay curious, and know which questions to ask to get up to speed fast.
  • Bring discipline and organisation: evidence-led investigations, accurate records, clean handovers.

Experience required: 6+ years in infrastructure, operations, or support engineering roles in production environments, including 2–3+ years hands-on with GPU, HPC, or large-scale data centre estates.

What You’ll be Doing (Responsibilities)

  • Join the Support duty rotation as a senior escalation point, collaborating with Infrastructure Engineering, CNPRE, Network Operations, and Product Engineering on incidents, investigations, and changes.
  • Diagnose and remediate GPU node faults across the full stack — driver, firmware, and hardware layers — from nvidia-smi/DCGM and XID/RAS analysis through BMC/Redfish and out-of-band management to physical fault isolation and vendor RMA.
  • Own east-west fabric health: run link-level diagnostics (mlxlink, ibdiagnet, or equivalent), isolate transceiver, optics, cabling, and switch-port faults, and validate topology across InfiniBand and RoCE/high-speed Ethernet fabrics.
  • Investigate data-path issues on high-performance storage platforms (e.g. VAST), including storage–network interactions across clients, mounts, VIPs, and routing.
  • Run structured, hypothesis-driven investigations; conduct root cause analysis for major incidents and drive long-term fixes to completion.
  • Author and execute changes in live customer environments with proper risk assessment, peer review, and backout plans.
  • Proactively improve dashboards, alerts, and runbooks to prevent repeat incidents; identify recurring patterns and convert them into problem records and automation.
  • Accurately record, update, and resolve tickets, keeping internal and external parties informed with clear customer-impact statements and evidence-rich notes that enable clean handover.
  • Design and implement automation scripts and small tools to reduce toil and human intervention.
  • Act as a key escalation point for the Support Organisation, taking ownership of strategic decisions where results matter.
  • Mentor and upskill mid-level engineers; contribute to knowledge sharing across Operations and Engineering, including training content, workshops, and PR reviews.
  • Lead by earning trust and speaking candidly. Disagree when appropriate and challenge the status quo; commit wholly to decisions once in motion.
  • Respond to critical incidents out of business hours and participate in on-call as required. Travel to Nscale or customer sites to provide onsite technical expertise.

About You (Skills / Qualifications Experience)

  • Experience.

6+ years in infrastructure, operations, or support engineering in production environments; 2–3+ years hands-on with GPU, HPC, or large-scale data centre estates, ideally in a customer-facing or escalation-driven capacity.

  • Communication.

Able to explain complex technical detail clearly, specifically, and concisely — in tickets, in incident updates, and face to face with customers and stakeholders at all levels. Strong written discipline: your notes let the next engineer pick up where you left off without starting from scratch.

  • GPU platforms (NVIDIA; AMD Instinct beneficial).

Practical, current experience with GPU drivers, firmware, and runtime stacks on AI training and inference clusters. Confident with nvidia-smi, DCGM, and XID/error interpretation; able to isolate faults across GPU, baseboard, NIC, and PCIe layers and drive them through diagnosis to RMA.

  • High-performance east-west fabrics.

Hands-on experience with RDMA fabrics — InfiniBand and/or RoCE — including link-layer diagnostics (mlxlink, ibdiagnet, or equivalent), transceiver and cabling fault isolation, and understanding of rail-optimised topologies, NVLink/NVSwitch, and NCCL-based performance troubleshooting on multi-node clusters.

  • HPC scheduling.

Slurm operations for large multi-GPU jobs — containers via Pyxis/Enroot, MPI, and diagnosing queue, topology, and job failures.

  • Linux systems engineering at scale.

Strong command of modern Linux distributions, kernel modules, systemd, networking stack, and filesystem tooling. Proven troubleshooting across compute, storage, and network layers in production.

  • Server hardware and control planes.

Comfortable with BMC/Redfish, firmware management, and bare-metal provisioning workflows (MAAS or similar) across large node fleets.

  • Networking fundamentals.

Solid grasp of L2/L3, routing, BGP, VLANs, VXLAN, firewalls, and load balancing, with a clear understanding of how east-west cluster traffic differs from north-south.

  • Observability and incident response.

Build and use alerting stacks and dashboards (Prometheus/Grafana or similar), interpret metrics and alerts, drive runbooks to resolution, and contribute to SLOs and post-incident reviews.

  • Change and risk judgment.

Experience authoring and executing changes in business-critical environments, including risk assessments, customer-impact analysis, and backout plans.

  • SRE-style operations.

Write and maintain runbooks, automate diagnostics, and reduce human intervention through scripts and small tools.

  • Automation and Git.

Scripting skills in Bash, Python, or equivalent for operational tooling and integrations; experience with infrastructure automation tools (Ansible, Terraform, or similar).

  • Data centre fundamentals.

Understanding of how data centres operate — servers, networks, storage, power, and cooling — ideally gained through an operational support background.

  • Leadership.

Disciplined, organised, and self-motivated, with the ability to mentor and motivate other engineers, take decisive action, and drive the team and wider organisation to improve.

  • Adaptability.

Able to adapt to customer-driven demands, including specialist support outside core hours and travel for onsite work.

Nice to Have

  • High-performance storage.

Hands-on experience with VAST or comparable AI-optimised storage platforms, or Ceph/parallel filesystems and NFS at scale (multipath, remoteports, nconnect), including diagnosing storage–network interaction and data-path performance issues.

  • OpenStack and fleet operations tooling.

OpenStack operations experience (Neutron, Cinder, error triage), plus familiarity with fleet-scale tooling for provisioning, health, and remediation across large GPU estates (MAAS, NetBox, Redfish-driven automation, or similar).

  • Kubernetes.

Operating and troubleshooting clusters, including GPU operator stacks and understanding how physical resources are abstracted up the stack. Helpful context for our platform, though not the core of this role.

  • Automation at scale.

Automated network configuration with safe, repeatable changes in business-critical environments; GitOps and CI/CD pipelines (GitHub Actions or similar); access and security tooling such as Teleport or Vault in production.

  • Certifications.

Relevant GPU/HPC, datacenter architecture, Linux, networking, Kubernetes, cloud, or security certifications (e.g. RHCSA/RHCE, CKA, NVIDIA-certified) are a plus.

What We Can Offer You

At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something extraordinary, and we want you at the core.

  • Highly competitive package, including base salary and equity, with reviews every 12 months.
  • Join one of the fastest-growing tech startups: your chance to push boundaries, collaborate with brilliant minds, and make your mark on cutting-edge AI.
  • Expect a dynamic progression plan tailored to your ambitions. Grow by trying new things, leading, challenging the status quo, and owning your impact, always with our full support.
  • Human-first flexibility. We treat you as humans first. Our flexible workplace trusts Nscalers to deliver, giving you the autonomy to shape your day around life's moments.
  • Join our thriving remote-first team. Geography is no barrier to impact or connection. We build seamless virtual collaboration, empowering you wherever you work.

Equal Opportunities Statement

At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enrich our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.

If there’s anything we can do to accommodate your specific situation, please let us know.

The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

Salary Range

$120,000—$170,000 USD

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior HPC/GPU Systems Engineer in Seattle, WA vacancy
  • Zoox is seeking a Senior Software Engineer focused on High Performance Computing in Seattle. This role involves enhancing Zoox's HPC infrastructure for machine learning workflows. Candidates...  ...in designing distributed storage systems and proficiency in languages like Python... 
    Senior

    jobs.frontdoordefense.com - Jobboard

    Seattle, WA
    15 hours ago
  • $153k - $204k

     ...Senior Systems Engineer, Test Frameworks & Validation PlatformCoreWeave is The Essential Cloud for AI™. Built for...  ...before it reaches one of the largest GPU fleets in the world, and we're extending that framework into HPC verification, Slurm-on-Kubernetes, and further... 
    Senior
    Permanent employment
    Full time
    Casual work
    Live in
    Work at office

    CoreWeave

    Bellevue, WA
    22 hours ago
  • $182k - $242k

     ...at What You'll Do: The Systems Engineering team owns the Linux kernel and...  ...one of the largest GPU fleets in the world. When something...  .... About the role: As a Senior Software Engineer on the Systems...  ...Partner with hardware, platform, HPC, and Fleet teams to ensure... 
    Senior
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Bellevue, WA
    13 days ago
  • $165k - $242k

    CoreWeave in Bellevue, WA is seeking an experienced engineer to lead design reviews and optimize distributed systems. The role requires 5-8 years in cloud services, strong skills in Python or Go, and hands-on experience with Kubernetes. Key responsibilities include defining... 
    Senior
    Flexible hours

    CoreWeave

    Bellevue, WA
    1 day ago
  • $180k - $240k

     ...Senior Systems EngineerHouston; New York; San Francisco; Seattle; USAbout the RoleWe are looking...  ...and ongoing operationMentor and guide engineers, raising the technical bar across the teamDiagnose...  ...environmentsFamiliarity with server and GPU hardware architecture and system-level... 
    Senior
    Flexible hours

    Nscale

    Seattle, WA
    3 days ago
  • $135.2k - $306.4k

    Oracle hardware platform development engineering is seeking a highly driven GPU/CPU Platform System Engineer at the Principal Engineer level. The GPU System Engineer will work within development engineering with a small team of talented engineers who lead the development... 
    Temporary work
    Work experience placement
    Remote work
    Flexible hours

    Oracle Corporation

    Seattle, WA
    2 days ago
  • CoreWeave is seeking a Senior Software Engineer on the Systems Engineering team to own and evolve our test framework — a Kubernetes-native, in-house system...  ...’ll harden the framework's core, broaden coverage into HPC and Slurm-on-Kubernetes, and keep CI fast and... 
    Senior

    CoreWeave

    Bellevue, WA
    1 day ago
  • A global technology company is seeking a Senior Engineer to optimize core data infrastructure and ensure high-performance computing within their advertising platform. In this role, you will drive innovations and engage in all software development lifecycle aspects while... 
    Senior

    The Trade Desk, Inc.

    Bellevue, WA
    15 hours ago
  •  ...Trade Desk, located in Bellevue, WA, is searching for Software Engineers to own all aspects of data-focused product development. You will...  ..., emphasizing high-performance computing and distributed systems. With 6+ years of experience in software development and strong... 
    Senior

    Thetradedesk

    Bellevue, WA
    15 hours ago
  • ByteDance in Seattle is seeking a Senior Research Engineer/Scientist for Storage for LLM to design and maintain a high-performance KV cache layer...  ...eviction policies, memory management, and explore open‑source KV stores or GPU‑aware caching techniques. #J-18808-Ljbffr ByteDance
    Senior

    ByteDance

    Seattle, WA
    15 hours ago
  • $100k - $150k

     ...GPU Systems Engineer - Remote    Bright Vision Technologies is a technology consulting and software development company delivering cloud...  ...performance CUDA kernels for compute-intensive workloads across AI and HPC use cases.  Profile and optimize GPU code using tools such... 
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    Bellevue, WA
    2 days ago
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its...  ...level performance projection tooling, and agentic optimization systems that improve GPU kernels at the assembly layer. Our team... 
    Senior
    Full time

    Nvidia

    Seattle, WA
    3 days ago
  • $152k - $241.5k

     ...Come join the team and see how you can make a lasting impact on the world.NVIDIA is hiring a Senior Compiler Engineer to join our team driving the next generation of GPU systems programming. We are redefining how developers write high-performance GPU software by... 
    Senior
    Full time
    Remote work

    Nvidia

    Seattle, WA
    4 days ago
  • Crusoe Cloud is revolutionizing HPC by offering sustainable, low-cost GPU compute power. As a Senior Cloud Support Engineer, you will be the primary technical support contact, helping customers leverage Crusoe Cloud to achieve their research goals and accelerate development... 
    Senior

    crusoe

    Bellevue, WA
    1 day ago
  • $155.9k - $233.9k

     ...name synonymous with entertainment excellence and creativity.Senior Systems EngineerBellevue, WAAre you on a mission to be a part of...  ...but not limited to, the following areas System Administration\Engineering on Linux and Windows Server and DesktopHigh performance networking... 
    Senior
    Immediate start
    3 days per week

    Sony Interactive Entertainment America

    Bellevue, WA
    15 hours ago
  • $110.33k

     ...community to enable innovation, learning, discovery, and service.IT Infrastructure provides knowledgeable system design and administration, software engineering, and operational support for academic systems, administrative systems, and research computing. The Computing... 
    Senior
    Full time
    Temporary work
    Work at office
    Shift work
    2 days per week

    University of Washington

    Seattle, WA
    1 day ago
  • $146.2k - $197.8k

    Senior Mission Systems EngineerCompany:The Boeing CompanyBoeing Defense, Space & Security (BDS) Mobility, Surveillance & Bombers (MS&B) is seeking a Senior Mission Systems Engineer with a focus on avionics and electrical systems to join our P-8A Augmented Fleet Support... 
    Senior
    Permanent employment
    Full time
    Temporary work
    Interim role
    Visa sponsorship
    Work visa
    Relocation package
    Flexible hours
    Shift work

    Boeing

    Tukwila, WA
    3 days ago
  • $272k - $431.25k

    NVIDIA data center systems have become core to NVIDIA's rapidly growing...  ...fully optimized NVIDIA AI and HPC software stack.We’re seeking a...  ...paths for transfers between GPU memory and storage, bypassing...  ...Computer Science, Electrical Engineering, or a related field.12+ overall... 
    Senior
    Full time

    Nvidia

    Seattle, WA
    1 day ago
  •  ...Blue Origin in Seattle is seeking a qualified engineer to perform fluid and thermal analysis on Lunar vehicle pump components. Candidates...  ...cycle, ultimately contributing to sustainable lunar transport systems. Benefits include comprehensive insurance plans, paid time off,... 
    Senior

    Blue Origin

    Seattle, WA
    21 hours ago
  • $184k - $287.5k

     ...customers in building AI/ML and HPC software solutions at scale....  ...performance issues for both AI and systems performance.What we need to...  ...MS/PhD in Electrical/Computer Engineering, Computer Science, Physics, or...  ....Hands-on experience with GPU systems in general including but... 
    Senior
    Full time

    Nvidia

    Seattle, WA
    3 days ago
  • TerraPower is looking for an Azure Systems Engineer to optimize and manage its Azure Government resources in Bellevue, Washington. This role...  ...least 5 years of Azure experience, along with capabilities in HPC and cloud resource management. TerraPower offers competitive compensation... 

    Dormont Manufacturing Co

    Bellevue, WA
    15 hours ago
  •  ...leading cloud service provider in Seattle seeks a Technical Leader for AWS platforms. You will solve complex architectural problems, own systems, and implement automation solutions driving AI transformation. The ideal candidate has extensive programming experience, strong... 
    Senior

    Amazon

    Seattle, WA
    21 hours ago
  • $152k - $241.5k

     ...computing, known for inventing the GPU and driving breakthroughs in...  ...generative AI to autonomous systems, and we continue to shape the...  ...tools that enable researchers and engineers to develop the next generation...  ...are looking for a strong AI & HPC Observability Engineer to... 
    Senior
    Full time

    Nvidia

    Seattle, WA
    1 day ago
  •  ...Sr Systems EngineerLocal (In the Office One Day a Week) 10 months, possible extensionTop 3 must-have hard skills:Scripting 5+Cloud 5+Security 5+Disqualifiers?:No App Dev developers!Not willing to commute to the officeTechnology requirements?Scripting experienceExperience... 
    Senior
    Work at office
    1 day per week

    Software Technology Inc

    Seattle, WA
    2 days ago
  •  ...embedded CPUs, touching millions of packets per second. You will implement data structures, optimize code, and work closely with senior engineers across EC2 and AWS to deliver new features for the platform. Ideal candidates have 5+ years in C/C++ (or Rust) and hands-on... 
    Senior

    Amazon

    Seattle, WA
    1 day ago
  • CoreWeave is seeking a Senior Engineer for its Benchmarking & Performance team to own kernel authoring and optimization for LLM inference. You...  ...-functional partners. Experience with CUDA, C++, Python, and GPU architectures is required; familiarity with vLLM, TensorRT-LLM... 
    Senior

    CoreWeave

    Bellevue, WA
    3 days ago
  • $120k - $200k

     ...Last Energy seeks a Senior Systems Engineer to lead the integration of our power plant systems and ensure seamless coordination between various engineering disciplines such as electrical, mechanical and structural. This candidate will optimize the product features, system... 
    Senior
    Full time

    Last Energy

    Seattle, WA
    21 hours ago
  • $160k - $185k

     ...Core Systems EngineerSkydance Games, a division of Paramount, a Skydance Corporation, is building the future of franchise-driven interactive...  ...extraordinary with us.Skydance is looking for a core systems engineer skilled in working with game development teams to build key... 
    Senior
    Currently hiring
    Remote work

    Skydance

    Seattle, WA
    15 hours ago
  • Adobe is seeking a senior, hands‑on engineer to own and evolve the cross‑platform GPU rendering platform at the heart of Premiere and After Effects. You will lead major initiatives, drive architecture decisions, and mentor teams across Windows and Mac environments. You... 
    Senior

    Adobe

    Seattle, WA
    3 days ago
  • $164k - $313.3k

    The Opportunity Photoshop ART is seeking a Senior Machine Learning (ML) Systems & Efficiency Engineer to join our R&D team focused on delivering practical, production...  ...teams to influence model design decisions, improve GPU utilization, and build scalable, cost‑aware ML... 
    Senior
    Local area

    Adobe Inc.

    Seattle, WA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior HPC/GPU Systems Engineer. Be the first to apply!