Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

HPC/GPU Systems Engineer

$100k - $140k
Full-time

Nscale

About Nscale

Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack — energy, data centres, GPU superclusters, orchestration, and AI services — delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world.

At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability, and delivering rapid response to customer issues. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future.

About the Role

Infrastructure Support Engineers are the delivery engine of Infrastructure Support (L2/L3), handling the day-to-day health of Nscale's GPU fleets — tickets, alerts, hardware faults, and customer issues — across GPU nodes, high-performance networks, Linux, and data centre operations. This is a hands-on technical role: you'll come in with a strong technical base and working knowledge of GPU infrastructure, and grow toward Senior through exposure to some of the most advanced AI infrastructure in the world.

You will:

  • Own your tickets and tasks end-to-end, escalating early and appropriately when issues exceed your scope — with clean, evidence-rich handovers.
  • Communicate technical detail clearly, specifically, and concisely — in tickets, to customers, and to colleagues. We treat communication quality as a core engineering skill, not a soft skill.
  • Grasp new technical concepts and problems quickly; stay curious and know what questions to ask to get up to speed fast.
  • Bring discipline and organisation: accurate records, structured troubleshooting, reliable follow-through.
  • Seek feedback and invest in learning — this role is a deliberate pathway to Senior.

Experience required: 3-4+ years in infrastructure support or support engineering roles, including deep support/service desk experience in structured, customer-facing environments, with working knowledge of GPU infrastructure and hands-on hardware troubleshooting.

What You’ll Be Doing

  • Join the Support duty rotation and handle day-to-day tickets and alerts, escalating early and appropriately. Collaborate with Engineering, with guidance, when incidents or changes require it.
  • Perform GPU node triage and hardware troubleshooting: interpret nvidia-smi/DCGM output and system logs, isolate faults across GPU, NIC, and server hardware, carry out physical remediation (reseats, swap testing, component checks), and prepare clean evidence for vendor RMA.
  • Run fabric and link diagnostics following established runbooks (mlxlink or equivalent), capture evidence accurately, and escalate with a handover that lets the next engineer continue without starting from scratch.
  • Assist with storage and data-path investigations (mounts, connectivity, client-side symptoms) on high-performance platforms, gathering evidence for Senior or Engineering-led diagnosis.
  • Follow established runbooks to resolve common issues; propose improvements and contribute incremental fixes with review.
  • Accurately record, update, manage, and resolve tickets, keeping all parties informed with clear notes, next steps, and customer communications via the agreed channels.
  • Participate in monitoring, troubleshooting, and triage. Capture logs and facts to enable efficient handover.
  • Participate in changes under peer review, learning risk assessment and backout practices in live customer environments.
  • Help maintain source-of-truth accuracy across DCIM, inventory, and asset records (NetBox or similar patterns).
  • Identify opportunities for automation and contribute simple scripts and tooling improvements to optimise processes.
  • Be the escalation point for onsite DC Operations staff; coordinate smart-hands tasks within your scope.
  • Learn the Platform fundamentals so you can help customers get value from our services, asking for support when deeper expertise is needed.
  • Share knowledge by documenting steps you've validated and contributing to training materials. Shadow Seniors during complex work to build capability.
  • Take part in incident reviews as a contributor and help track preventative follow-ups in your scope.
  • Deliver assigned tasks and project work to agreed quality and timelines. Flag blockers early and seek help when needed.
  • Participate in on-call and out-of-hours work when scheduled and after onboarding. Travel to Nscale or customer locations to assist with deployments, troubleshooting, and operational tasks, and attend supplier training as required.

About You

  • Experience. 3+ years in infrastructure support or support engineering, including support/service desk experience in structured, SLA-driven, customer-facing environments (cloud, data centre, or managed services).
  • Communication. Clear written notes, concise updates, and reliable follow-through. Able to explain technical issues accurately to customers and colleagues, and produce handovers the next shift can act on immediately.
  • GPU and hardware troubleshooting. Working knowledge of GPU infrastructure: hands-on with nvidia-smi or similar diagnostics, comfortable interpreting hardware error output and logs, and confident physically troubleshooting servers — reseating components, swap testing, working via BMC/out-of-band management — through to preparing RMA evidence. A strong technical base here is required, not a learning goal.
  • Linux. Solid working knowledge: confident on the CLI with systemd, filesystems, permissions, and standard networking tools. Able to troubleshoot common issues independently and know when to escalate.
  • Networking. Solid grasp of IP addressing, subnets, VLANs, routing, DNS, and firewalls. Awareness of high-performance east-west fabrics (RDMA/InfiniBand concepts) is a plus and a core growth area in this role.
  • Ticketing and ITSM discipline. Experience working within structured support processes (ITIL or similar): prioritisation, escalation, SLA awareness, and accurate documentation.
  • Observability foundations. Able to use dashboards and alerts to identify symptoms, gather evidence, and follow runbooks. Comfortable proposing simple alert or dashboard improvements with review.
  • Scripting and automation basics. Comfortable reading and writing simple Bash or Python, and using Git for version control.
  • Platform and DC fundamentals. Understanding of servers, networks, storage, and virtualisation concepts, ideally from a support or operations background.
  • Growth mindset. Curious, dependable, and collaborative. You seek feedback, ask questions, and invest in learning to progress toward Senior.
  • Adaptability. Able to work in a fast-moving environment with evolving processes, participate in on-call after onboarding, and travel when needed.

Nice to Have

These are growth areas, not prerequisites:

  • High-performance fabrics and GPU-HPC: exposure to RDMA/InfiniBand, link-level diagnostics (mlxlink, ibdiagnet, or equivalent), NCCL-based troubleshooting, or NVLink concepts.
  • High-performance storage: exposure to VAST or comparable AI-optimised storage platforms, Ceph, or NFS at scale, including basic storage–network troubleshooting.
  • OpenStack and fleet operations tooling: familiarity with OpenStack troubleshooting flows, or fleet-scale tooling for provisioning and health (MAAS, NetBox, Redfish, or similar).
  • Kubernetes: understanding of core concepts (nodes, pods, services, logs) and basic troubleshooting via runbooks. Helpful context for our platform, though not the core of this role.
  • Automation and access tooling: experience with Ansible or Terraform, CI/CD participation (GitHub Actions or similar), or access and security tooling such as Teleport or Vault.
  • Certifications: progress toward relevant Linux, networking, Kubernetes, cloud, or security certifications over time.

What We Can Offer You

At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something extraordinary, and we want you at the core.

  • Highly competitive package (base + equity) with reviews every 12 months.
  • Join the fastest-growing tech startup, your chance to push boundaries, collaborate with brilliant minds, and make your mark on cutting-edge AI. ✨
  • Expect a dynamic progression plan tailored to your ambitions. Grow by trying new things, leading, challenging the status quo, and owning your impact, always with our full support.
  • Human-First Flexibility: We treat you as humans first. Our flexible workplace trusts Nscalers to deliver, giving you the autonomy to shape your day around life's moments.
  • Join our thriving remote-first team. Geography is no barrier to impact or connection. We build seamless virtual collaboration, empowering you, wherever you work.

Equal Opportunities Statement

We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.

If there’s anything we can do to accommodate your specific situation, please let us know.

The responsibilities outlined in this job description are not exhaustive and are intended to provide a general overview of the position. The employee may be required to perform additional duties, tasks, and responsibilities as assigned by management, consistent with the skills and qualifications required for the role.

The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

Salary Range

$100,000—$140,000 USD

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

Nscale does not accept unsolicited candidate submissions from recruitment agencies.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the HPC/GPU Systems Engineer in Houston, TX vacancy
  • $100k - $150k

     ...vertically integrated AI cloud engineered for AI. We own and operate the...  ...stack — energy, data centres, GPU superclusters, orchestration,...  ...2–3+ years hands-on with GPU, HPC, or large-scale data centre estates...  ...job failures. ~ Linux systems engineering at scale.... 
    Suggested
    Full time
    Remote work
    Flexible hours

    Nscale

    Houston, TX
    2 days ago
  • $54 per hour

     ...Description: A minimum of 5 years’ experience working in a large HPC enterprise environment comprising thousands of servers, large...  ...installation, configuration and management of Linux based operating systems, preferably using RHEL, CentOS, Rocky Linux. Experience with IBM... 
    Suggested
    Hourly pay
    Contract work
    Work experience placement
    Local area
    Remote work
    Weekend work

    US Tech Solutions

    Houston, TX
    2 days ago
  • $180k - $240k

     ...About the Role We are looking for a Senior Systems Developer to lead the design, development...  ...ongoing operation Mentor and guide engineers, raising the technical bar across the...  ...environments Familiarity with server and GPU hardware architecture and system-level... 
    Suggested
    Full time
    Flexible hours

    Nscale

    Houston, TX
    2 days ago
  • $105.5k - $243k

    HPC and AI Performance EngineerThis role has been designed as 'Hybrid...  ...HPC benchmarks on HPE systems to generate performance results...  ...Computer Science, Mathematics, Engineering, Physics, Chemistry, Environmental...  ...as CUDA, OpenACC, and related GPU programming... 
    Suggested
    Full time
    Work experience placement
    Work at office
    Local area
    Immediate start
    2 days per week

    Hewlett Packard Enterprise

    Houston, TX
    11 hours ago
  •  ...infrastructure that powers AI platforms, GPU-accelerated workloads, large-...  ...computing solutions, aligning system architecture and deployment...  ...CUDA along with LLM inference engines (TensorRT-LLM), production...  ...1,000+ GPU clusters for AI, HPC, and agentic AI workloads with... 
    Suggested
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    Houston, TX
    2 days ago
  • $130k - $200k

     ...Site Reliability Engineer Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient...  ...SRE role for someone who wants to own systems, not just watch them. You'll take real...  ...workloads, or high-performance computing (HPC). • Familiarity with high-performance... 
    Shift work

    Nscale

    Houston, TX
    4 days ago
  •  ...'s workload, design the cloud HPC architecture, prove the performance...  ...an HPC/AI Technical Solution Engineer, you will operate in the zone...  ..., cloud transformation, AI/GPU infrastructure, and customer business...  .... Distributed/parallel file systems and storage fundamentals.... 
    Full time
    Work at office
    Worldwide

    CGG

    Houston, TX
    11 hours ago
  • Are you passionate about solving complex problems? Join our dynamic team as a Systems Engineer and become a vital contributor to spaceflight equipment development! As an integral part of our organization, you'll play a crucial role in ensuring the successful definition... 
    Work at office

    Oceaneering International

    Houston, TX
    3 days ago
  • $86.9k - $198k

    Human Spaceflight Systems EngineerThe Opportunity:Are you looking for an opportunity to combine your technical skills with big picture...  ...makes you an integral part of delivering a customer focused engineering solution. As a systems engineer on our team, you have the chance... 
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Houston, TX
    11 hours ago
  •  ...architecture patterns for industrial data services across edge systems, control system data, cloud services, partner platforms,...  ...Vendor and Partner CoordinationWork with vendors and partner engineering teams to define integration responsibilities, data handoffs, technical... 
    Local area
    Remote work

    National Oilwell Varco

    Houston, TX
    2 days ago
  •  ...culture and contribute to our core mission which is enhancing our customer's experience.Position SummaryThe Originations Business Systems Engineer I, is responsible for the independent design, configuration, and implementation of solutions within the enterprise... 
    Visa sponsorship
    Work visa
    Monday to Friday
    Weekend work

    First Investors

    Houston, TX
    3 days ago
  •  ...sustainable future. For more information, please visit fluenceenergy.com.Job Description:Role SummaryFluence is seeking a Senior Systems Engineering Manager, Operability to lead the global Hypercare organization responsible for post-deployment technical stabilization, fleet... 
    Permanent employment
    Full time
    Visa sponsorship
    Work visa

    Fluence Energy

    Houston, TX
    1 day ago
  •  ...jobWe are seeking an experienced DevOps / Platform Engineer with deep expertise in AWS services, Terraform, Python, and HPC infrastructure. This role will work closely with...  ...autoscaling groups, AWS Batch and both CPU and GPU compute resourcesSet up monitors and logs for... 

    EPAM Systems

    Houston, TX
    9 hours ago
  • $115k - $135k

     ...Innovation) Team. This position is a hands-on cloud and platform engineering role with a strong focus on AI-enabled platforms, GitHub-based...  ...workloads, primarily within Microsoft Azure.The CCSI Systems Senior Engineer is a full-time role located in Southfield, MI,... 
    Full time
    Work at office
    Remote work
    Relocation package
    Monday to Friday

    AlixPartners

    Houston, TX
    3 days ago
  •  ...FEATURED JOBS Systems Engineer EXPERIENCE REQUIRED: 10 to 15 Years NUMBER OF POSITIONS: 01 DEPARTMENT: Information Technology REPORTS TO: Vice President of IT LOCATION: Houston, TX ROLE OVERVIEW: The Systems Engineer will participate in the delivery... 
    Work at office
    Local area
    Immediate start
    Remote work
    Night shift

    AIS, LLC

    Houston, TX
    5 days ago
  •  ...Jones Lang LaSalle Incorporated seeks a Lead Operating Engineer for on-site control of mechanical, HVAC, electrical and plumbing systems in Houston, TX. The role requires 3-5 years of relevant experience, with leadership responsibilities and a focus on preventative maintenance... 

    Jones Lang LaSalle Incorporated

    Houston, TX
    2 days ago
  •  ...and hardwareApplication migration to different Windows OS platforms (2012, 2016 2019, 2022)Analyze, design, and document platforms/systems to meet enterprise requirementsInstall, customize, maintain, test, and troubleshoot operating systems and other systems softwareConfigure... 
    Work at office

    My3Tech Inc

    Houston, TX
    17 hours ago
  •  ...primarily on the United States inland and Intracoastal Waterway systems. The partnership’s assets include approximately 50,000 miles of...  ...your opportunities for success.The Staff Systems Planning Engineer at Enterprise Products is responsible for performing advanced hydraulic... 

    Enterprise Products

    Houston, TX
    4 days ago
  •  ...plan operations. We offer care management programs for asthma, diabetes, and high-risk pregnancy. An affiliate of the Harris Health System (Harris Health), Community is financially self-sufficient and receives no financial support from Harris Health or from Harris... 
    Contract work
    Work experience placement
    Flexible hours
    Night shift
    Weekend work

    Harris Health System

    Houston, TX
    2 days ago
  •  ...Sr. System Engineer Client: Japanese IT Company Working Location: Houston, TX Employment Type: Full-time Salary: Up to 140K / per year (DOE) Benefit: Full Benefits Visa Support: Possible Language: English and Japanese Key Responsibilities Provide IT infrastructure... 
    Full time
    Visa sponsorship

    Cinter LLC

    Houston, TX
    2 days ago
  •  ...Senior Systems Engineer for NASA's Orion Program ARES Corporation is seeking a Senior Systems Engineer to support NASA's Orion Program and broader NASA human spaceflight initiatives. This position is a key contributor to Cross Program Ground Support Equipment (GSE)... 
    For contractors
    Work at office
    Flexible hours

    ARES

    Houston, TX
    2 days ago
  •  ...Job Description Job Description Title: HPC Platform Engineer Location: Houston, TX (Hybrid, First 6 Months Onsite)...  ...infrastructure Configure and optimize parallel file systems (Lustre preferred) Support GPU and accelerated computing environments Monitor... 
    Contract work

    Summa

    Houston, TX
    28 days ago
  • $124k - $280k

     ...Description & SummaryAt PwC, our people in data and analytics engineering focus on leveraging advanced technologies and techniques to design...  ...will involve designing and optimising algorithms, models, and systems to enable intelligent decision-making and automation.Growing... 
    Full time
    H1b

    PwC

    Houston, TX
    4 days ago
  • $76k - $155.7k

     ...Electromagnetic Compatibility (EMC) / Electrostatic Discharge (ESD) Systems EngineerJob Category: EngineeringTime Type: Full timeMinimum...  ...(EMI) / Electromagnetic Compatibility (EMC) Systems Engineer with an Electrostatic Discharge (ESD) specialty to join the EMI... 
    Permanent employment
    Contract work
    Work experience placement
    Flexible hours

    CACI International

    Houston, TX
    1 day ago
  • The Senior Systems Engineer - Predictive Fleet Intelligence is the engineering authority responsible for defining how fleet health is monitored, how developing equipment failures are detected, and how engineering knowledge is transformed into scalable digital solutions... 
    Remote work
    Night shift
    Weekend work

    Patterson-UTI

    Houston, TX
    1 day ago
  • $110k - $165k

     ...future of our communities. This is a Lead Cloud & Infrastructure Engineering position at the Vice President level, which is part of the job...  ...infrastructure and ensuring the seamless operation of IT systems to support business needs effectively.Morgan Stanley is an industry... 
    Temporary work
    Local area
    Remote work
    Flexible hours

    Morgan Stanley

    Houston, TX
    11 hours ago
  • Position SummaryNOV Tuboscope is seeking an NDT Systems Engineer to support the development, testing, implementation, and continuous improvement of advanced inspection technologies for tubular products. This role will work closely with engineering, operations, and manufacturing... 

    National Oilwell Varco

    Houston, TX
    11 hours ago
  •  ...office in Houston, Texas. Our summer 2027 internships start mid to end of May and will run through the beginning of August.Financial Systems InternLocation: Houston, TX (Hybrid schedule, 4 days in office)About TMHCCHelp us insure it. Tokio Marine HCC is a leading global... 
    Summer work
    Internship
    Summer internship
    Work at office
    Local area

    Tokio Marine HCC

    Houston, TX
    1 day ago
  •  ...gathering, processing, treating, transmission, storing, and exporting natural gas and liquids. You will assist with process simulations, system modeling, flow measurements, pipe sizing, equipment functions, and field operations across pipelines, storage, and terminals. The... 
    Internship
    Summer internship

    Energy Transfer Equity

    Houston, TX
    3 days ago
  • $79.6k - $156.7k

    Systems Engineer, ATCS Position Description Are you passionate about human space exploration, understanding the origins of the universe, and working with a passionate and diverse team to make a difference? If you are, we need you! We have an exciting opportunity... 
    Work at office
    Local area
    Houston, TX
    14 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to HPC/GPU Systems Engineer. Be the first to apply!