Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Infrastructure Support Engineer

Nscale

Senior Infrastructure Support Engineer

Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack — energy, data centres, GPU superclusters, orchestration, and AI services — delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world.

At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability, and delivering rapid response to customer issues. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future.

About the Role (Job Purpose)

Senior Infrastructure Support Engineers are the senior technical escalation point within Infrastructure Support, owning the health of Nscale's GPU fleets and the high-performance fabrics that connect them. This is a hands-on L2/L3 role operating at the intersection of GPU hardware, east-west networking, Linux, and data centre operations — acting as the operational bridge between Support, DC Operations, and Engineering.

You will:

  • Own complex, ambiguous problems end-to-end and make decisive calls in a results-driven environment, taking calculated risks where speed matters.
  • Communicate technical detail clearly, specifically, and concisely — to engineers, to customers, and to leadership. We treat communication quality as a core engineering skill, not a soft skill.
  • Influence without authority and build strong relationships with senior stakeholders across the business to get things done.
  • Grasp new technical concepts quickly, stay curious, and know which questions to ask to get up to speed fast.
  • Bring discipline and organisation: evidence-led investigations, accurate records, clean handovers.
What You'll be Doing (Responsibilities)
  • Join the Support duty rotation as a senior escalation point, collaborating with Infrastructure Engineering, CNPRE, Network Operations, and Product Engineering on incidents, investigations, and changes.
  • Diagnose and remediate GPU node faults across the full stack — driver, firmware, and hardware layers — from nvidia-smi/DCGM and XID/RAS analysis through BMC/Redfish and out-of-band management to physical fault isolation and vendor RMA.
  • Own east-west fabric health: run link-level diagnostics (mlxlink, ibdiagnet, or equivalent), isolate transceiver, optics, cabling, and switch-port faults, and validate topology across InfiniBand and RoCE/high-speed Ethernet fabrics.
  • Investigate data-path issues on high-performance storage platforms (e.g. VAST), including storage–network interactions across clients, mounts, VIPs, and routing.
  • Run structured, hypothesis-driven investigations; conduct root cause analysis for major incidents and drive long-term fixes to completion.
  • Author and execute changes in live customer environments with proper risk assessment, peer review, and backout plans.
  • Proactively improve dashboards, alerts, and runbooks to prevent repeat incidents; identify recurring patterns and convert them into problem records and automation.
  • Accurately record, update, and resolve tickets, keeping internal and external parties informed with clear customer-impact statements and evidence-rich notes that enable clean handover.
  • Design and implement automation scripts and small tools to reduce toil and human intervention.
  • Act as a key escalation point for the Support Organisation, taking ownership of strategic decisions where results matter.
  • Mentor and upskill mid-level engineers; contribute to knowledge sharing across Operations and Engineering, including training content, workshops, and PR reviews.
  • Lead by earning trust and speaking candidly. Disagree when appropriate and challenge the status quo; commit wholly to decisions once in motion.
  • Respond to critical incidents out of business hours and participate in on-call as required. Travel to Nscale or customer sites to provide onsite technical expertise.
About You (Skills / Qualifications Experience)
  • 6+ years in infrastructure, operations, or support engineering in production environments; 2–3+ years hands-on with GPU, HPC, or large-scale data centre estates, ideally in a customer-facing or escalation-driven capacity.
  • Able to explain complex technical detail clearly, specifically, and concisely — in tickets, in incident updates, and face to face with customers and stakeholders at all levels. Strong written discipline: your notes let the next engineer pick up where you left off without starting from scratch.
  • Practical, current experience with GPU drivers, firmware, and runtime stacks on AI training and inference clusters. Confident with nvidia-smi, DCGM, and XID/error interpretation; able to isolate faults across GPU, baseboard, NIC, and PCIe layers and drive them through diagnosis to RMA.
  • Hands-on experience with RDMA fabrics — InfiniBand and/or RoCE — including link-layer diagnostics (mlxlink, ibdiagnet, or equivalent), transceiver and cabling fault isolation, and understanding of rail-optimised topologies, NVLink/NVSwitch, and NCCL-based performance troubleshooting on multi-node clusters.
  • Slurm operations for large multi-GPU jobs — containers via Pyxis/Enroot, MPI, and diagnosing queue, topology, and job failures.
  • Strong command of modern Linux distributions, kernel modules, systemd, networking stack, and filesystem tooling. Proven troubleshooting across compute, storage, and network layers in production.
  • Comfortable with BMC/Redfish, firmware management, and bare-metal provisioning workflows (MAAS or similar) across large node fleets.
  • Solid grasp of L2/L3, routing, BGP, VLANs, VXLAN, firewalls, and load balancing, with a clear understanding of how east-west cluster traffic differs from north-south.
  • Build and use alerting stacks and dashboards (Prometheus/Grafana or similar), interpret metrics and alerts, drive runbooks to resolution, and contribute to SLOs and post-incident reviews.
  • Experience authoring and executing changes in business-critical environments, including risk assessments, customer-impact analysis, and backout plans.
  • Write and maintain runbooks, automate diagnostics, and reduce human intervention through scripts and small tools.
  • Scripting skills in Bash, Python, or equivalent for operational tooling and integrations; experience with infrastructure automation tools (Ansible, Terraform, or similar).
  • Understanding of how data centres operate — servers, networks, storage, power, and cooling — ideally gained through an operational support background.
  • Disciplined, organised, and self-motivated, with the ability to mentor and motivate other engineers, take decisive action, and drive the team and wider organisation to improve.
  • Able to adapt to customer-driven demands, including specialist support outside core hours and travel for onsite work.
Nice to Have
  • Hands-on experience with VAST or comparable AI-optimised storage platforms, or Ceph/parallel filesystems and NFS at scale (multipath, remoteports, nconnect), including diagnosing storage–network interaction and data-path performance issues.
  • OpenStack operations experience (Neutron, Cinder, error triage), plus familiarity with fleet-scale tooling for provisioning, health, and remediation across large GPU estates (MAAS, NetBox, Redfish-driven automation, or similar).
  • Operating and troubleshooting clusters, including GPU operator stacks and understanding how physical resources are abstracted up the stack. Helpful context for our platform, though not the core of this role.
  • Automated network configuration with safe, repeatable changes in business-critical environments; GitOps and CI/CD pipelines (GitHub Actions or similar); access and security tooling such as Teleport or Vault in production.
  • Relevant GPU/HPC, datacenter architecture, Linux, networking, Kubernetes, cloud, or security certifications (e.g. RHCSA/RHCE, CKA, NVIDIA-certified) are a plus.
What We Can Offer You

At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something extraordinary, and we want you at the core.

  • Highly competitive package, including base salary and equity, with reviews every 12 months.
  • Join one of the fastest-growing tech startups: your
Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Senior Infrastructure Support Engineer in United States vacancy
  •  ...Senior Infrastructure Support Engineer The Senior Infrastructure Support Engineer role is responsible for the operations of secure and highly available eDiscovery platforms, servers, and networks. They install, maintain, upgrade, and continuously improve the client... 
    Senior
    Currently hiring
    Remote work

    GEORGE JON

    United States
    4 days ago
  • $95 - $125 per hour

     ...Job Summary: Our client is seeking a Security Infrastructure Support Senior Security Engineer to join their team! This position is located in Bethesda, Maryland. Duties: Design, deploy, and maintain enterprise IT security systems across hybrid... 
    Senior
    Local area

    KellyMitchell Group

    Bethesda, MD
    4 days ago
  • Nscale seeks a Senior Infrastructure Support Engineer to own the health of GPU fleets and high‑performance fabrics. You will operate across GPU hardware, Linux, and data centre operations, bridging Support, DC Operations, and Engineering. You’ll diagnose complex issues... 
    Senior
    Remote work

    Nscale

    San Francisco, CA
    4 days ago
  • $100k - $120k

    Todyl in Atlanta, Georgia, is seeking a Technical Support Engineer III to deliver critical technical support for networking and security products. The ideal candidate will have over 4 years of experience, strong troubleshooting skills, and the ability to communicate technical... 
    Senior

    Todyl

    Atlanta, GA
    1 day ago
  • Nokia in Coppell, TX is seeking a Customer Care Support Specialist to serve as the technical interface for customers deploying Nokia Fixed Networks Access products. You will provide world-class post-sales leadership, act as subject matter expert, and build strong customer... 
    Senior

    Nokia

    Coppell, TX
    2 days ago
  • Nokia seeks a Customer Care Support Specialist to act as a trusted technical interface for customers deploying Nokia Fixed Networks. You will provide world-class post-sales technical leadership, becoming the subject matter expert on Nokia Fixed Networks products and solutions... 
    Senior

    Nokia Corp.

    Dallas, TX
    3 days ago
  • Fortinet is seeking a Support Specialist to enhance Customer Success and Support. This role requires strong troubleshooting skills and a minimum of 3 years of relevant experience in technical support or system administration. The candidate will work directly with customers... 
    Senior

    Zoomcar

    Frisco, TX
    3 days ago
  • $130k - $155k

     ...and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each...  ...offering sustainable, low-cost GPU compute power. As a Senior Cloud Support Engineer, you'll play a crucial role in empowering our customers... 
    Senior
    Full time
    Temporary work

    Crusoe

    Denver, CO
    4 days ago
  • Fortinet is seeking a Technical Support Specialist in Atlanta, GA to provide exceptional customer service and technical support for Fortinet's products, including FortiAnalyzer and FortiManager. The ideal candidate should have over 4 years of experience in a technical support... 
    Senior

    Zoomcar

    Atlanta, GA
    3 days ago
  •  ...Join Our Compute Product Support Group Are you passionate about...  ...with Site Reliability Engineering and Product teams for in-depth...  ...effective mitigation. As a Senior Cloud Support Engineer, you...  ...Cloud Computing products and infrastructure. Demonstrate advanced troubleshooting... 
    Senior
    Remote work

    Akamai

    United States
    1 day ago
  • Prosum is urgently seeking a Tier 2 Support Engineer located in El Segundo, California. The successful candidate will configure and support systems, troubleshoot network issues, and maintain application security. This role involves significant client interaction and may... 
    Senior

    Prosum

    El Segundo, CA
    1 day ago
  • $120k - $207k

    We are seeking a highly motivated Senior HPC Support Engineer - Ethernet / AI Infrastructure to be engaging with one of our prestige customers onsite and remote, passionate about data center and networking technologies, to provide comprehensive solutions for sophisticated... 
    Senior
    Full time
    Work experience placement
    Remote work

    Nvidia

    Texas
    18 hours ago
  •  ...Job Description Position Overview: We are seeking a Senior Cloud Support Engineer to join our dynamic team. In this role, you will...  ...environments. The ideal candidate has a strong background in cloud infrastructure, networking, and security technologies, with a preferred... 
    Senior
    Temporary work
    Work at office

    Futurex

    Bulverde, TX
    more than 2 months ago
  •  ...Senior Cloud Support Engineer (CSE) Snowflake seeks a Senior Cloud Support Engineers who combine technical expertise, customer empathy, and an AI-first mindset. You'll leverage and refine AI tools to accelerate troubleshooting, improve knowledge bases, and reduce time... 
    Senior
    Remote work
    Shift work
    Weekend work

    Streamlit

    United States
    1 day ago
  • A leading IT solutions provider is seeking a Support Telecom/VoIP/Cloud Engineer in California. This role involves delivering Tier 2 and Tier 3...  ...responsible for troubleshooting issues, managing VoIP infrastructure, and collaborating with teams to enhance communication... 
    Senior

    ProTelesis Corporation

    San Diego, CA
    4 days ago
  • CoreWeave is seeking a Technical Support Engineer - III to empower customers utilizing their advanced Kubernetes-powered HPC cloud infrastructure. This role involves mentoring team members, providing high-touch support, and collaborating with cross-functional teams to... 
    Senior

    CoreWeave

    Livingston, NJ
    4 days ago
  • Neverware seeks a motivated individual for its Support team to provide technical assistance to US and European education and enterprise customers. In this role, you will resolve customer issues through calls and emails, ensuring a seamless CloudReady experience. The ideal... 
    Senior

    Neverware

    New York, NY
    4 days ago
  • Esolvit Inc in Dallas, Texas is seeking an experienced IT Infrastructure Specialist with strong expertise in Microsoft Windows 2012, Active Directory, and Azure Cloud. The role demands a minimum of 4 years of experience in server management, cloud integration, and problem... 
    Senior

    Esolvit Inc

    Dallas, TX
    1 day ago
  • Salesforce is seeking an experienced technical support engineer to diagnose and resolve complex issues across our software products. You...  ...and release governance, while coordinating with R&D and Infrastructure on escalated cases. The role emphasizes ownership of critical... 
    Senior

    Salesforce.Com Inc

    Austin, TX
    4 days ago
  • A leading e-mobility company is seeking a Technical Customer Support Engineer to provide support for B2B customer issues and queries. The role involves adapting to frequent context switching while managing multiple projects. Required qualifications include 3+ years of... 
    Senior

    Driivz

    Austin, TX
    18 hours ago
  • Lightedge Solutions is seeking an experienced Support Engineer to collaborate across networks, storage, virtualization, and disaster recovery disciplines. You will lead escalations, partner with platform engineers, and help maintain high availability and customer satisfaction... 
    Senior

    Lightedge

    Kansas City, MO
    1 day ago
  • Crusoe Cloud is revolutionizing HPC by offering sustainable, low-cost GPU compute power. As a Senior Cloud Support Engineer, you will be the primary technical support contact, helping customers leverage Crusoe Cloud to achieve their research goals and accelerate development... 
    Senior

    crusoe

    Bellevue, WA
    2 days ago
  • Motorola Solutions is seeking an experienced Technical Support professional for the RapidDeploy cloud SaaS platform. You will provide day-to-day customer support, resolve issues, and drive CSAT across diverse user groups. You will work with telemetry tools to troubleshoot... 
    Senior
    Relocation

    Motorola Solutions

    Chicago, IL
    2 days ago
  • Lightedge in Austin is seeking an experienced Support Engineer focusing on collaboration across technologies and departments. In this role, you will address complex customer issues, work with cutting-edge tech, and contribute to innovation and customer support excellence... 
    Senior

    Lightedge

    Austin, TX
    1 day ago
  • Hammerspace is seeking a highly skilled Customer Support Engineer for the West Coast and Asia, to ensure success and retention of our customer base. The role focuses on solving complex L3/L4 issues with storage, filesystem, and networking expertise, while maintaining strong... 
    Senior
    Night shift

    Triwill Group

    New York, NY
    3 days ago
  • Axon in San Francisco is hiring a Technical Support Engineer (Tier-2+) to support customers and internal teams on the Axon 911 platform, ensuring reliability, performance, and operational excellence across a growing SaaS environment. You will collaborate with Engineering... 
    Senior
    Work at office

    Prepared

    San Francisco, CA
    4 days ago
  • Crusoe Cloud seeks a Senior Cloud Support Engineer to empower customers using sustainable GPU compute power. You’ll be the main technical contact, diagnose issues, and work across SRE, Networking, and Storage teams to ensure service reliability. The role emphasizes customer... 
    Senior

    Crusoe

    Sunnyvale, CA
    1 day ago
  • Crusoe Cloud is hiring a Senior Cloud Support Engineer to empower customers using Crusoe's sustainable GPU compute platform. You will be the primary technical contact, resolving issues across AI/ML workloads, physics simulations, and computational biology, while driving... 
    Senior
    Remote work

    Crusoe

    Dallas, TX
    2 days ago
  • Crusoe Cloud is revolutionizing high‑performance computing with sustainable, low‑cost GPU power. As a Senior Cloud Support Engineer, you will be the primary contact for technical support, helping customers leverage Crusoe Cloud to accelerate AI research and development... 
    Senior

    Crusoe

    Dallas, TX
    3 days ago
  • Crusoe Cloud is revolutionizing high-performance computing with sustainable GPU compute power. As a Cloud Support Engineer, you will be the primary contact for technical support, helping customers leverage Crusoe Cloud to achieve their goals and accelerate research and... 
    Senior
    Remote work

    Crusoe

    Sunnyvale, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Infrastructure Support Engineer. Be the first to apply!