Head of Infrastructure Support
Nscale
Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack — energy, data centres, GPU superclusters, orchestration, and AI services — delivering high-performance infrastructure to AI-native companies, enterprises and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world.
At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability, and delivering rapid response to customer issues. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future.
About the Role
The Head of Infrastructure Support owns Infrastructure Support for their region — the team, the function, and its impact on customers. Reporting directly to the VP of Support and operating alongside counterpart Heads of Infrastructure Support across EMEA, the US, and APAC, you are accountable for the success of regional support outcomes: service performance, escalation quality, customer experience, and the health of the GPU estates your team supports.
The regional Infrastructure Support engineers report directly to you, and you own their management end to end — hiring, 1:1s, performance reviews, development planning and documented performance management through to outcome. Their performance, growth, and results are your responsibility.
You will grow your regional Infrastructure Support team during a period of rapid company scaling, embed a consistent operating model with your counterparts in the other regions to deliver true follow-the-sun coverage, and act as the organisational accountability layer for your region — ensuring that strategic and tactical work spanning Support and Operations lands with clear owners and gets driven to completion. As the function matures, you will own a global capability area on behalf of all regions and shape your team's structure — including developing team leads — as headcount grows.
You remain technically credible: close enough to GPU infrastructure, high-performance fabrics, and Linux operations to lead complex incident response, challenge technical decisions on their merits, and earn the respect of Senior Engineers — while spending the majority of your time leading.
Experience required:
8+ years in infrastructure, operations, or support engineering in production environments, including 4+ years of direct line management of engineers in an operational support function, with demonstrable ownership of performance management. Significant exposure to GPU, HPC or large-scale data centre estates.
What You'll Be Doing
Regional Ownership & Accountability
- Own the success of Infrastructure Support for your region: service outcomes, customer impact, and team performance sit with you.
- Own regional service performance against defined KPIs — SLA adherence, MTTR, first-response time, backlog health, and CSAT — with accurate reporting to the VP of Support and senior leadership.
- Identify regional risks — capacity, capability, coverage, or customer — early, and either resolve them or make clear recommendations with escalation.
- Own regional capacity modelling and headcount planning: forecast support demand against fleet growth and customer onboarding, and make the business case for investment to the VP of Support.
- Act as regional accountability layer during rapid growth: when cross-functional work spanning Support, DC Operations, deployment, firmware and Engineering lacks a clear owner, make sure it gets one and gets done.
- Partner with the Heads of Infrastructure Support in the other regions — across EMEA, the US and APAC — to run a single global function: consistent standards, processes and quality, with true follow-the-sun handover between regions.
- Own a global capability area on behalf of all regions — such as escalation management standards, the knowledge and runbook system, or the tooling and automation roadmap — working with other Heads of Infrastructure Support defining the standard every regional Support team operates to.
- Own day-to-day people management for your regional Infrastructure Support team: regular 1:1s, performance reviews, development planning and documented performance management—including underperformance—through to outcome.
- Hire and grow the team: define role requirements, run structured interviews, and build a bench of engineers who meet Nscale's technical and communication bar.
- Design your team's structure as the region scales, appointing and developing team leads and building second-line management capability as headcount grows.
- Set and monitor individual and team objectives, driving accountability and continuous improvement.
- Design and own shift planning, rota coverage and on-call scheduling for the region, ensuring sustainable 24/7 support in coordination with the global coverage model.
- Identify skills gaps and drive upskilling through training, mentoring and knowledge sharing across teams.
- Ensure roles, responsibilities and expectations are clearly understood and consistently applied.
Service & Operational Performance
- Own ticket queue health for the region: accurate prioritisation, timely resolution and clean escalation flow from frontline triage into L2/L3.
- Monitor team productivity and workload trends, addressing bottlenecks before they become service risks.
- Ensure adherence to ITIL-aligned processes across incident, request, change and problem management.
- Improve dashboards, alerting and runbooks to reduce repeat incidents and drive right-first-time resolution.
- Maintain consistent standards, processes and documentation across regional teams; ensure compliance with audit, security and operational requirements.
Incident, Escalation & Stakeholder Leadership
- Act as the senior regional escalation point for complex or high-impact incidents, including customer-facing escalations, participating in regional on-call as needed.
- Lead post-incident reviews, identify recurring patterns and ensure follow-up actions are tracked and delivered — converting incidents into problem records and durable fixes.
- Represent Infrastructure Support to regional customers and internal senior stakeholders; communicate clearly, candidly and concisely at every level from engineer to executive.
- Contribute to readiness and support planning for new services, data centre deployments and customer onboarding in the region.
Technical Leadership & Contribution
- Work alongside Senior Engineers on complex incidents, technical improvements and operational tooling — close enough to the work to lead it credibly.
- Maintain hands‑on fluency across GPU infrastructure (drivers, firmware, hardware fault isolation, RMA workflows), Linux at scale and east-west high-performance fabrics (InfiniBand/RoCE diagnostics and fault isolation).
- Guide investigation quality: evidence-led diagnosis, structured troubleshooting and handovers that stand up to scrutiny.
- Contribute to scripting and automation direction reduce toil across the regional operation.
- Travel to Nscale or customer sites when needed to lead onsite support activity.
About You
5+ years of direct line management of engineers in an operational support environment, with end-to-end ownership of performance management: reviews, development plans and documented underperformance processes through to outcome. You can describe your management framework and point to engineers you've grown.
- Operational ownership.
Experience owning team workload, prioritisation and service delivery against SLAs, with accountability for the numbers — and experience explaining those numbers to senior leadership.
- Function building.
Experience hiring, scaling or standing up support/operations capability in a fast-moving environment, including capacity modelling and headcount planning against demand; comfortable operating where processes are still evolving and helping define them without slowing delivery.
- Communication.
Excellent written and verbal communication — clear, specific, and concise at every level, from ticket notes to executive updates to difficult customer conversations. We treat communication quality as a core leadership skill and assess it directly.
- Decisiveness and accountability.
A bias for decisive action and calculated risk in ambiguous situations; you take ownership of outcomes, speak candidly, disagree when appropriate and commit fully once decisions are made.
- Technical foundation — 8+ years across:
- Linux systems engineering in production, with proven troubleshooting across compute, storage and network layers.
- GPU infrastructure. Working knowledge of GPU platforms (NVIDIA; AMD beneficial) — driver/firmware stacks, hardware diagnostics (nvidia-smi, DCGM), fault isolation and RMA workflows on AI or HPC estates.
- High-performance east‑west fabrics. Understanding of RDMA over InfiniBand and/or RoCE, link‑level diagnostics, and how fabric health drives cluster performance; able to lead and challenge fabric‑related incident response.
- HPC scheduling and orchestration. Hands‑on Slurm operation for multi‑GPU workloads, including containerised execution via Pyxis/Enroot, MPI‑based communication and deep diagnostics of queue health, network topology and job‑level failures.
- Networking fundamentals. L2/L3, routing, VLANs, load balancing and how east‑west cluster traffic differs from north‑south.
- Data centre operations. Solid understanding of servers, networking, storage, power, virtualization and IT in an operational support context, including working with onsite DC Operations and smart‑hands teams.
- Observability and incident response. Interpreting metrics and alerts, driving incidents to resolution and leading post‑incident improvement.
- Automation. Scripting (Bash, Python or similar) and familiarity with IaC tools (Ansible, Terraform or similar).
- Process literacy.
Strong understanding of ITIL‑aligned incident, problem and change management and of SRE practices — runbooks, toil reduction, continuous improvement.
Comfortable with out‑of‑hours escalations, regional on‑call participation and travel for onsite leadership.
Nice to Have
NCCL‑based performance troubleshooting, NVLink/NVSwitch, Slurm‑scheduled multi‑GPU workloads or rack‑scale systems.
- High‑performance storage.
Exposure to VAST or comparable AI‑optimised storage platforms, Ceph, or NFS at scale.
- OpenStack and fleet operations tooling.
OpenStack operations, or fleet‑scale provisioning and health tooling (MAAS, NetBox, Redfish‑driven automation or similar).
Operating or supporting clusters and GPU operator stacks. Helpful context for our platform, though not the core of this role.
- Multi‑region or follow‑the‑sun operations.
Experience running or coordinating support across regions and time zones.
ITIL certification or relevant Linux/networking/cloud certifications.
For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.
#J-18808-Ljbffr$180k - $277k
...provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise... ...bolsters technical capabilities and directly supports strategic business outcomes, including... .... Role Overview We are seeking a Head of Infrastructure Operations to lead the...SuggestedContract workFor contractorsFlexible hours$302k - $335k
...platforms that power innovation across a complex organization? As AI Infrastructure Director , you’ll design, manage, and optimize the firm’s AI... ..., paid time off, and retirement. We also offer personal support and tailored learning and development opportunities all...SuggestedContract workWork at officeWorldwideFlexible hours$71.3k
...more $71,300.00 yearly Full-time As a Senior Technology Sales Support Specialist, you will play a pivotal role in designing and... ...manage, maintain, and improve the company's existing technology infrastructure.HillDay is a growing Houston-based professional servi... Show...SuggestedFull timePart timeWork at officeRemote workWork from homeRelocation- ...version, large‑scale virtualized environments). Strong in Linux support and system engineering Deep expertise in computer storage,... ...Insight Global is assisting a client in identifying a hands‑on Infrastructure Operations Manager to lead a highly technical team...Suggested
- ...A growing energy infrastructure company is seeking a Controller to lead accounting operations and financial reporting in a multi-entity... ...oversee the close processes, strengthen internal controls, and support financial planning initiatives. Candidates should have a Bachelor...Suggested
$110k - $130k
...Prevost is looking for a Service Network Support Manager based in Houston, Texas, to enhance customer satisfaction and profitability across the service center network in the US and Canada. You will work closely with the Sr. Director and lead initiatives to drive financial...$110k - $130k
...deploying scalable network solutions and upgrading existing infrastructure. Manage hybrid connectivity, including VPCs, Transit Gateways... ...within [AWS/Azure] environments. Provide Tier 2/3 technical support for complex network outages and performance issues, performing...Work experience placement- Role description Job Title: Network Engineer Location: Houston, TX (Onsite) Job Description: 7 to 11 years Experience in deployment configuration of Cisco switches Cisco routers F5 Cisco ISE Cisco WLCs and Cisco Access Point Strong knowledge and Hands on...
- ...management abilities Roles & Responsibilities Network Design & Implementation • Design, configure, and maintain LAN/WAN infrastructure, including routers, switches, firewalls, and wireless systems. • Plan and execute network upgrades, migrations, and lifecycle...
- ...and the ability to manage complex network environments while supporting modernization and transformation initiatives. Responsibilities... ...Lead network hardware modernization initiatives to ensure infrastructure alignment with current technologies Plan and execute...Local area
- ...expertise in Cloudflare Web Application Firewall (WAF) to join our infrastructure team. This is a highly specialized role that sits at the... ..., administration, and optimization of Cloudflare services supporting enterprise web applications while partnering closely with networking...
- Artech is the 10th Largest IT Staffing Company in the US, according to Staffing Industry Analysts' 2012 annual report. Artech provides technical expertise to fill gaps in clients' immediate skill-sets availability, deliver emerging technology skill-sets, refresh existing...Work experience placementImmediate start
$80k - $100k
...Employment Model: Fully onsite at client location Job Summary We are seeking an Infrastructure Network Engineer to serve as a dedicated onsite resource supporting a single client location. This role is heavily networking-focused and responsible for day-to-...- ...Network Engineer The Network Engineer is primarily responsible for supporting HPC data center networking, enterprise security infrastructure, and distributed WAN environments across our global regions. Job Responsibilities Support deployment, troubleshooting, and maintenance...
- ...automation System hardening according to Cisco guides Able to document as-built design (L1 – L3) and provide KT/hand-off to support team Able to create full set of test-cases and validate & document network failure conditions Able to manage compute systems...Local areaFlexible hours
- ...accelerates the transition to clean, renewable energy. Conduit Power is seeking a Network Engineer to support the communications and connectivity infrastructure that enables reliable operation of distributed generation assets, field technology, monitoring systems,...Long term contractWork at officeRemote workMonday to ThursdayFlexible hours
- ...world. Job Overview The Network Engineer will directly support the Zscaler Platform Manager in executing product strategy, technical... ...- Experience with virtual-machine management and cloud-infrastructure deployment Advanced Technical Capabilities Deep...Visa sponsorship
- ...project goals. Evaluate the contractual scope of work. Plan, organize and staff key project positions through regional department heads, subordinate project managers, general superintendents, etc. Establish project objectives, policies, procedures and performance...Contract workFor contractorsFor subcontractor
- ...Infrastructure Project Manager Responsibilities Develops, clarifies and manages the scope of the project, defines deliverables and achieves targeted outcomes Manages and delivers IT infrastructure and build-out projects Responsible for tracking progress against target...
- ...Workforce Management (Pay Policies, Rules Engine, Time & Attendance, Scheduling) with SQL/SSIS, payroll integrations, and experience supporting or migrating to Oracle Fusion Cloud (Time & Labor). Job Description: The Sumplr Workforce Systems Analyst...Contract work
- ...library system located in Harris County, Texas. This technical support role focuses on intermediate-level systems, network, and... ...This role provides intermediate-level technical support for infrastructure, networking, systems administration, enterprise applications,...Full timeRemote workMonday to FridayAfternoon shift
$67.7k - $90.27k
...enterprises, governments, and communities. At Lumen, you’ll work on infrastructure customers rely on today and build for what’s next, where... ...us today. The Role The Network Inventory GIS Engineer supports the organization’s GIS network inventory. This role is...Full timeTemporary workRemote workWork from home$59.15k - $106.93k
...seeking a highly skilled and experienced Network Engineer II to support the Advanced Enterprise Global Information Technology... ...configuration, and troubleshooting of the NASA enterprise network infrastructure with a focus on Wi‑Fi technologies. You will leverage your understanding...Work at office- ...site organization. The Network Engineer will collaborate with infrastructure and application teams to ensure optimal network performance,... ...and hybrid environments. Key Responsibilities Implements and supports firewall/DMZ security infrastructure, including traffic flow...Work experience placement
- ...Solutions, Network Security, Data Center & Network Technologies, IT Support Services, Network Cabling, Phone Systems & Physical Security.... ...3+ yearsof experience working with Enterprise-Grade Network Infrastructure and other technologies in a corporate environment. Operations...Work experience placementLocal areaFlexible hours
- Overview We are seeking a hands-on Network Engineer to support a large-scale switch and wireless refresh project across North America... ...on design and BOM development for switching and wireless infrastructure We are a company committed to creating diverse and inclusive...Work at officeLocal areaRemote workWeekend work
- ...Engineer with hands-on experience in Starlink deployments, network infrastructure, and enterprise connectivity solutions. The ideal candidate... ..., provisioning, configuration, troubleshooting, and ongoing support across LAN/WAN environments. Roles & Responsibilities Deploy,...
- ...Network Engineer to join our team in Houston, TX. This role is integral to designing, implementing, and optimizing our network infrastructure to support organizational growth and ensure reliable connectivity. The ideal candidate will bring advanced expertise in routing,...
- ...high‑quality services, Aldridge is committed to helping its clients optimize their technology infrastructure and achieve their business goals. Position Overview The Support Engineer will be a pivotal team member based in Dallas, TX, responsible for handling some of our...Local areaRemote work
$72.12 - $120.19 per hour
...Infrastructure Deployment Engineer This is a full-time exempt (40+ hours/week), remote and travel-heavy contract role located in the... ...with engineering teams, contractors, and facility operations to support successful datacenter turn-ups. The position drives...Hourly payFull timeContract workFor contractorsRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Head of Infrastructure Support. Be the first to apply!
- head of infrastructure Houston, TX
- infrastructure manager Houston, TX
- director of infrastructure Houston, TX
- information technology support Houston, TX
- onsite support Houston, TX
- recovery support Houston, TX
- purchasing support Houston, TX
- linux support Houston, TX
- remote support Houston, TX
- support operator Houston, TX



