Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Head of Infrastructure Support

Nscale

About Nscale Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack — energy, data centres, GPU superclusters, orchestration, and AI services — delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world. About The Role The Head of Infrastructure Support owns Infrastructure Support for their region — the team, the function, and its impact on customers. Reporting directly to the VP of Support and operating alongside counterpart Heads of Infrastructure Support across EMEA, the US, and APAC, you are accountable for the success of regional support outcomes: service performance, escalation quality, customer experience, and the health of the GPU estates your team supports. The regional Infrastructure Support engineers report directly to you, and you own their management end to end — hiring, 1:1s, performance reviews, development planning, and documented performance management through to outcome. Their performance, growth, and results are your responsibility. You will grow your regional Infrastructure Support team during a period of rapid company scaling, embed a consistent operating model with your counterparts in the other regions to deliver true follow-the-sun coverage, and act as the organisational accountability layer for your region — ensuring that strategic and tactical work spanning Support and Operations lands with clear owners and gets driven to completion. As the function maturing, you will own a global capability area on behalf of all regions and shape your team's structure — including developing team leads — as headcount grows. You remain technically credible: close enough to GPU infrastructure, high-performance fabrics, and Linux operations to lead complex incident response, challenge technical decisions on their merits, and earn the respect of Senior Engineers — while spending the majority of your time leading. Experience Required 8+ years in infrastructure, operations, or support engineering in production environments, including 4+ years of direct line management of engineers in an operational support function, with demonstrable ownership of performance management. Significant exposure to GPU, HPC, or large-scale data centre estates. What You'll Be Doing Regional Ownership & Accountability Own the success of Infrastructure Support for your region: service outcomes, customer impact, and team performance sit with you. Own regional service performance against defined KPIs — SLA adherence, MTTR, first-response time, backlog health, and CSAT — with accurate reporting to the VP of Support and senior leadership. Identify regional risks — capacity, capability, coverage, or customer — early, and either resolve them or escalate them with a clear recommendation. Own regional capacity modelling and headcount planning: forecast support demand against fleet growth and customer onboarding, and make the business case for investment to the VP of Support. Act as the regional accountability layer during rapid growth: when cross-functional work spanning Support, DC Operations, deployment, firmware, and Engineering lacks a clear owner, make sure it gets one and gets done. Partner with the Heads of Infrastructure Support in the other regions — across EMEA, the US, and APAC — to run a single global function: consistent standards, processes, and quality, with true follow-the-sun handover between regions. Own a global capability area on behalf of all regions — such as escalation management standards, the knowledge and runbook system, or the tooling and automation roadmap — working with other Heads of Infrastructure Support defining the standard every regional Support team operates to. People Leadership & Team Building Own day-to-day people management for your regional Infrastructure Support team: regular 1:1s, performance reviews, development planning, and documented performance management — including underperformance — through to outcome. Hire and grow the team: define role requirements, run structured interviews, and build a bench of engineers who meet Nscale's technical and communication bar. Design your team's structure as the region scales, appointing and developing team leads and building second-line management capability as headcount grows. Set and monitor individual and team objectives, driving accountability and continuous improvement. Design and own shift planning, rota coverage, and on-call scheduling for the region, ensuring sustainable 24/7 support in coordination with the global coverage model. Identify skills gaps and drive upskilling through training, mentoring, and knowledge sharing across teams. Ensure roles, responsibilities, and expectations are clearly understood and consistently applied. Service & Operational Performance Own ticket queue health for the region: accurate prioritisation, timely resolution, and clean escalation flow from frontline triage into L2/L3. Monitor team productivity and workload trends, addressing bottlenecks before they become service risks. Ensure adherence to ITIL-aligned processes across incident, request, change, and problem management. Improve dashboards, alerting, and runbooks to reduce repeat incidents and drive right-first-time resolution. Maintain consistent standards, processes, and documentation across regional teams; ensure compliance with audit, security, and operational requirements. Incident, Escalation & Stakeholder Leadership Act as the senior regional escalation point for complex or high-impact incidents, including customer-facing escalations, participating in regional on-call as required. Lead post-incident reviews, identify recurring patterns, and ensure follow-up actions are tracked and delivered — converting incidents into problem records and durable fixes. Represent Infrastructure Support to regional customers and internal senior stakeholders; communicate clearly, candidly, and concisely at every level from engineer to executive. Contribute to readiness and support planning for new services, data centre deployments, and customer onboarding in the region. Technical Leadership & Contribution Work alongside Senior Engineers on complex incidents, technical improvements, and operational tooling — close enough to the work to lead it credibly. Maintain hands‑on fluency across GPU infrastructure (drivers, firmware, hardware fault isolation, RMA workflows), Linux at scale, and east‑west high-performance fabrics (InfiniBand/RoCE diagnostics and fault isolation). Guide investigation quality: evidence‑led diagnosis, structured troubleshooting, and handovers that stand up to scrutiny. Contribute to scripting and automation direction to reduce toil across the regional operation. Travel to Nscale or customer sites when needed to lead onsite support activity. About You Leadership experience. 5+ years of direct line management of engineers in an operational support environment, with end-to-end ownership of performance management: reviews, development plans, and documented underperformance processes through to outcome. You can describe your management framework and point to engineers you've grown. Operational ownership. Experience owning team workload, prioritisation, and service delivery against SLAs, with accountability for the numbers — and experience explaining those numbers to senior leadership. Function building. Experience hiring, scaling, or standing up support/operations capability in a fast-moving environment, including capacity modelling and headcount planning against demand; comfortable operating where processes are still evolving and helping define them without slowing delivery. Communication. Excellent written and verbal communication — clear, specific, and concise at every level, from ticket notes to executive updates to difficult customer conversations. We treat communication quality as a core leadership skill and assess it directly. Decisiveness and accountability. A bias for decisive action and calculated risk in ambiguous situations; you take ownership of outcomes, speak candidly, disagree when appropriate, and commit fully once decisions are made. Technical foundation — 8+ years across: Linux systems engineering in production, GPU infrastructure, high-performance east-wester fabrics, HPC scheduling and orchestration, networking fundamentals, data centre operations, observability and incident response, automation, process literacy. Strong understanding of ITIL-aligned incident, problem, and change management, and of SRE practices — runbooks, toil reduction, and continuous improvement. Adaptability. Comfortable with out-of-hours escalations, regional on-call participation, and travel for onsite leadership. Nice to Have Deeper GPU/HPC exposure: NCCL-based performance troubleshooting, NVLink/NVSwitch, Slurm-scheduled multi-GPU workloads, or rack-scale systems. High-performance storage: Exposure to VAST or comparable AI-optimised storage platforms, Ceph, or NFS at scale. OpenStack and fleet operations tooling: OpenStack operations, or fleet-scale provisioning and health tooling (MAAS, NetBox, Redfish-driven automation, or similar). Kubernetes: Operating or supporting clusters and GPU operator stacks. Helpful context for our platform, though not the core of this role. Multi-region or follow-the-sun operations: Experience running or coordinating support across regions and time zones. Formal qualifications: ITIL certification, or relevant Linux/networking/cloud certifications. For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here. #J-18808-Ljbffr Nscale

Vacancy posted 12 hours ago
Similar jobs that could be interesting for youBased on the Head of Infrastructure Support in San Francisco, CA vacancy
  • The San Francisco Compute Company is hiring a hands-on Customer Support Lead to own day-to-day support and success. You will shape the team, tooling, and metrics that keep customers live, productive, and loyal. It’s a builder’s seat where you’ll scale support as we grow... 
    Suggested

    Doist

    San Francisco, CA
    12 hours ago
  • The San Francisco Compute Company is seeking a hands-on Customer Support Lead to own day-to-day support and customer success. You will build the team, tooling, and metrics that keep customers live, productive, and loyal, with rapid problem resolution. This is a builder’... 
    Suggested
    3 days per week

    The San Francisco Compute Company

    San Francisco, CA
    12 hours ago
  •  ...connect their financial accounts to the apps they use every day. Our Support team plays a critical role—delivering fast, high-quality help...  ...issues and helps customers get more from our products. As Head of Support, you will own our global support strategy and outcomes... 
    Suggested
    Work experience placement
    Local area

    Plaid Inc

    San Francisco, CA
    1 day ago
  • $180k - $250k

     ...Equity Role Overview This is a founding leadership role for the Support function. You will join an initial team of three and scale it...  ...Knowledge base — Create scalable documentation and knowledge infrastructure. Operating cadence — Build proactive systems that prevent... 
    Suggested
    Full time
    Immediate start

    Recruitment Room - Global

    San Francisco, CA
    4 days ago
  • Recruitment Room - Global is seeking a founding Head of Support in San Francisco to scale an initial team of three to eight to ten specialists by year-end. You will directly handle key accounts and escalations while building the processes and systems that define hospital... 
    Suggested
    Immediate start

    Recruitment Room - Global

    San Francisco, CA
    23 hours ago
  • $180k - $240k

    About Pylon Pylon is building the future of B2B customer support. We believe support should be a strategic advantage: deeply technical,...  ...support organization should set the standard. We are looking for a Head of Support to build a world‑class customer experience today... 
    Work at office

    Pylon

    San Francisco, CA
    1 day ago
  • Nscale in San Francisco is seeking a Head of Infrastructure Support to lead regional operations across the US, EMEA, and APAC. You will own regional service outcomes, customer impact, and the health of GPU estates your team supports, reporting to the VP of Support. You... 
    Shift work

    Nscale

    San Francisco, CA
    4 days ago
  • Pylon is building the future of B2B customer support. We are looking for a Head of Support to build a world‑class customer experience today while defining what an AI‑native support organization should look like for the next several years. You will lead our global Support... 

    Pylon

    San Francisco, CA
    23 hours ago
  •  ...is looking for an experienced Engineering leader to lead our Infrastructure, SRE and Enterprise Governance Engineering teams. Anyscale...  ...cluster launcher, cloud providers (AWS/GCP/Azure/etc.), Kubernetes support, cluster autoscaling, control plane, data plane, reliability,... 

    Cerebras

    San Francisco, CA
    17 hours ago
  •  ...contracted MW. What we need from you You have personally sourced and closed MW-scale colocation agreements at a compute or infrastructure company -- CoreWeave, Crusoe, Lambda, Nebius, Applied Digital, a hyperscaler siting team, or a colo/data center developer... 
    Contract work
    Remote work

    General Compute

    San Francisco, CA
    17 hours ago
  • An AI infrastructure startup located in San Francisco is seeking a Head of Product to define and drive their product vision. The role entails ownership of product...  ...customer discovery, and ensuring the product roadmap supports competitive advantages. The ideal candidate will... 

    Hamilton Barnes Associates Limited

    San Francisco, CA
    12 hours ago
  •  ...ensure rapid energization of data center capacity. The role requires deep knowledge of North American power markets, a strong rolodex of operators and brokers, and a hands-on approach to early-stage infrastructure work, including site #J-18808-Ljbffr General Compute Inc.

    General Compute Inc.

    San Francisco, CA
    4 days ago
  • Casca is seeking a Director of Customer Support to build and lead the customer support function from the ground up, ensuring high-quality support for banking customers on the Casca platform. This role involves hiring and training a team, managing escalations, and integrating... 

    Casca

    San Francisco, CA
    1 day ago
  •  ...be sent from @Rippling.com addresses.About The RoleRippling Infrastructure operates as a scaled engineering organization responsible for...  ...that power every product and platform across the company. We support nearly 1,000 engineers globally and deliver the foundational... 

    Rippling

    San Francisco, CA
    2 days ago
  • $400k

    Join a seed-stage AI infrastructure company building large-scale training and inference platforms...  ...complex. They are now seeking a Head of Product to own and shape the product...  ...Identify patterns across sales cycles, support issues, and product usage data Convert... 
    Immediate start

    Hamilton Barnes Associates Limited

    San Francisco, CA
    12 hours ago
  • Baseten in San Francisco seeks a Head of Data Center Infrastructure to own the strategy, design, and execution of our first-generation GPU data centers, including power distribution, cooling, rack design, and low-voltage systems. This high-ownership, hands-on leadership... 
    For contractors

    Neura Market

    San Francisco, CA
    2 days ago
  •  ...financial world.Role DescriptionAs the Director of Corporate Infrastructure, you will drive efforts to oversee the design, implementation...  ...operations to improve user experience and performance, and also supporting the multi-terabit backbone network that interconnects edge... 
    Remote work

    SoFi

    San Francisco, CA
    3 days ago
  • Plaid Inc in San Francisco is seeking a Head of Support to own global support strategy. You will lead a distributed team focused on delivering high-quality customer support while managing critical incidents and aligning with cross-functional teams. The ideal candidate has... 

    Plaid Inc

    San Francisco, CA
    1 day ago
  •  ...Gumloop is seeking a Head of Marketing to guide the function through its next growth phase. You will set the marketing strategy, align priorities, and own how Gumloop is positioned and brought to market. This role leads a multidisciplinary marketing team and partners... 

    Gumloop

    San Francisco, CA
    17 hours ago
  • Blacksmith is seeking a Capacity Planning Manager to ensure we can scale our infrastructure in San Francisco. You'll own the capacity planning process, build forecasting models, and manage supplier relationships to ensure availability as demand grows. Ideal candidates... 

    Blacksmith

    San Francisco, CA
    2 days ago
  •  ...A leading digital marketplace bank in the U.S. is seeking a Sr. Director for Risk Infrastructure in San Francisco. The role requires 15+ years of experience in fintech, expertise in decision engines, and strong leadership skills. You will be responsible for driving strategic... 

    LendingClub

    San Francisco, CA
    17 hours ago
  • $293k - $385k

     ...first-party data centers, strategic partners, and industrial infrastructure environments. We focus on converting power, land, hardware, and...  ...execution into reliable compute capacity that can support frontier AI training and inference workloads.This team operates... 
    Work at office
    Local area
    Flexible hours

    OpenAI

    San Francisco, CA
    12 hours ago
  • $229k - $285k

     ...sector.Visit gomotive.com to learn more.About the Role: We are looking for a seasoned engineering leader who will lead the Core Infrastructure and Operations Engineering. You will be responsible for managing and mentoring a multi-functional team across SRE, Operations,... 
    Temporary work
    Work experience placement

    Motive Technologies

    San Francisco, CA
    2 days ago
  • Prime Intellect, Inc. seeks a Compute Strategy Lead to oversee GPU sourcing, economics, and contracts. This role involves shaping the industry by negotiating substantial agreements and collaborating with research teams on compute requirements. Key qualifications include...
    Remote job
    Flexible hours

    Prime Intellect, Inc.

    San Francisco, CA
    2 days ago
  • $109k - $186k

     ...technical decisions throughout the design and deployment phases, and support our customers through network turn-up and validation. Your...  ...to scale and meet the growing demand for reliable internet infrastructure.You will have an impact byDesigning performant and resilient... 

    Meter

    San Francisco, CA
    2 days ago
  • Nurix Therapeutics is seeking an operating leader for IT infrastructure and operations in Brisbane, CA. You will own the function end to end: infrastructure, cybersecurity operations, end user services, and IT service management as one integrated scope. You partner with... 

    Nurix Therapeutics

    Brisbane, CA
    3 days ago
  • Robertson & Sumner Ltd is seeking a Vice President of Sales for AI Infrastructure on the US West Coast. This executive-level role involves direct ownership of strategic expansion across leading AI companies and offers significant autonomy in shaping sales strategies. The... 

    Robertson & Sumner Ltd

    San Francisco, CA
    1 day ago
  • $160k - $260k

     ...businesses. Powered by our unique combination of proprietary infrastructure and software, we empower over 250,000 businesses worldwide -...  ...s build what’s next.About the teamCorporate IT drives our IT support, IT engineering and business engineering functions at Airwallex... 
    Temporary work
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours
    Weekend work

    Airwallex

    San Francisco, CA
    1 day ago
  •  ...consulting, enterprise architecture Experience in Cisco DC networks like Cisco Nexus Switching, network security, internet connectivity/infrastructure design and DC network migration is must.DC network design, implementation and migration compromising Cisco Nexus (ACI & NXOS),... 
    Flexible hours

    Sonsoft

    San Francisco, CA
    4 days ago
  •  ...issues to restore service quickly.Troubleshoot connectivity, performance, and wireless issues across the leaf / spine (Clos) fabric.Support core network services - DNS, DHCP, and similar - running on open-source platforms.Maintain documentation, runbooks, and network... 

    Eclaro International

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Head of Infrastructure Support. Be the first to apply!