Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Infrastructure Support Engineer

$120k - $170k

Nscale

About Nscale Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack—energy, data centres, GPU superclusters, orchestration, and AI services—delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world. About The Role (Job Purpose) Senior Infrastructure Support Engineers are the senior technical escalation point within Infrastructure Support, owning the health of Nscale's GPU fleets and the high-performance fabrics that connect them. This is a hands‑on L2/L3 role operating at the intersection of GPU hardware, east‑west networking, Linux, and data centre operations—acting as the operational bridge between Support, DC Operations, and Engineering. You Will Own complex, ambiguous problems end-to-end and make decisive calls in a results‑driven environment, taking calculated risks where speed matters. Communicate technical detail clearly, specifically, and concisely— to engineers, to customers, and to leadership. We treat communication quality as a core engineering skill, not a soft skill. Influence without authority and build strong relationships with senior stakeholders across the business to get things done. Grasp new technical concepts quickly, stay curious, and know which questions to ask to get up to speed fast. Bring discipline and organisation: evidence‑led investigations, accurate records, clean handovers. What You’ll Be Doing (Responsibilities) Join the Support duty rotation as a senior escalation point, collaborating with Infrastructure Engineering, CNPRE, Network Operations, and Product Engineering on incidents, investigations, and changes. Diagnose and remediate GPU node faults across the full stack—driver, firmware, and hardware layers—from nvidia-smi/DCGM and XID/RAS analysis through BMC/Redfish and out-of-band management to physical fault isolation and vendor RMA. Own east‑west fabric health: run link-level diagnostics (mlxlink, ibdiagnet or equivalent), isolate transceiver, optics, cabling, and switch‑port faults, and validate topology across InfiniBand and RoCE/high-speed Ethernet fabrics. Investigate data-path issues on high‑performance storage platforms (e.g. VAST), including storage–network interactions across clients, mounts, VIPs, and routing. Run structured, hypothesis-driven investigations; conduct root cause analysis for major incidents and drive long-term fixes to completion. Author and execute changes in live customer environments with proper risk assessment, peer review, and backout plans. Proactively improve dashboards, alerts, and runbooks to prevent repeat incidents; identify recurring patterns and convert them into problem records and automation. Accurately record, update, and resolve tickets, keeping internal and external parties informed with clear customer-impact statements and evidence-rich notes that enable clean handover. Design and implement automation scripts and small tools to reduce toil and human intervention. Act as a key escalation point for the Support Organisation, taking ownership of strategic decisions where results matter. Mentor and upskill mid-level engineers; contribute to knowledge sharing across Operations and Engineering, including training content, workshops, and PR reviews. Lead by earning trust and speaking candidly. Disagree when appropriate and challenge the status quo; commit wholly to decisions once in motion. Respond to critical incidents out of business hours and participate in on-call as required. Travel to Nscale or customer sites to provide onsite technical expertise. About You (Skills / Qualifications Experience) Experience. 6+ years in infrastructure, operations, or support engineering in production environments; 2–3+ years hands‑on with GPU, HPC, or large-scale data centre estates, ideally in a customer-facing or escalation-driven capacity. Communication. Able to explain complex technical detail clearly, specifically, and concisely— in tickets, in incident updates, and face to face with customers and stakeholders at all levels. Strong written discipline: your notes let the next engineer pick up where you left off without starting from scratch. GPU platforms (NVIDIA; AMD Instinct beneficial). Practical, current experience with GPU drivers, firmware, and runtime stacks on AI training and inference clusters. Confident with nvidia-smi, DCGM, and XID/error interpretation; able to isolate faults across GPU, baseboard, NIC, and PCIe layers and drive them through diagnosis to RMA. High-performance east-west fabrics. Hands‑on experience with RDMA fabrics—InfiniBand and/or RoCE—including link-layer diagnostics (mlxlink, ibdiagnet or equivalent), transceiver and cabling fault isolation, and understanding of rail‑optimised topologies, NVLink/NVSwitch, and NCCL-based performance troubleshooting on multi-node clusters. HPC scheduling. Slurm operations for large multi-GPU jobs—containers via Pyxis/Enroot, MPI, and diagnosing queue, topology, and job failures. Linux systems engineering at scale. Strong command of modern Linux distributions, kernel modules, systemd, networking stack, and filesystem tooling. Proven troubleshooting across compute, storage, and network layers in production. Server hardware and control planes. Comfortable with BMC/Redfish, firmware management, and bare-metal provisioning workflows (MAAS or similar) across large node fleets. Networking fundamentals. Solid grasp of L2/L3, routing, BGP, VLANs, VXLAN, firewalls, and load balancing, with a clear understanding of how east-west cluster traffic differs from north-south. Observability and incident response. Build and use alerting stacks and dashboards (Prometheus/Grafana or similar), interpret metrics and alerts, drive runbooks to resolution, and contribute to SLOs and post-incident reviews. Change and risk judgment. Experience authoring and executing changes in business-critical environments, including risk assessments, customer-impact analysis, and backout plans. SRE-style operations. Write and maintain runbooks, automate diagnostics, and reduce human intervention through scripts and small tools. Automation and Git. Scripting skills in Bash, Python or equivalent for operational tooling and integrations; experience with infrastructure automation tools (Ansible, Terraform or similar). Data centre fundamentals. Understanding of how data centres operate—servers, networks, storage, power, and cooling—ideally gained through an operational support background. Leadership. Disciplined, organised, and self-motivated, with the ability to mentor and motivate other engineers, take decisive action, and drive the team and wider organisation to improve. Adaptability. Able to adapt to customer-driven demands, including specialist support outside core hours and travel for onsite work. Nice to Have High-performance storage. Hands‑on experience with VAST or comparable AI-optimised storage platforms, or Ceph/parallel filesystems and NFS at scale (multipath, remoteports, nconnect), including diagnosing storage–network interaction and data-path performance issues. OpenStack and fleet operations tooling. OpenStack operations experience (Neutron, Cinder, error triage), plus familiarity with fleet-scale tooling for provisioning, health, and remediation across large GPU estates (MAAS, NetBox, Redfish-driven automation or similar). Kubernetes. Operating and troubleshooting clusters, including GPU operator stacks and understanding how physical resources are abstracted up the stack. Helpful context for our platform, though not the core of this role. Automation at scale. Automated network configuration with safe, repeatable changes in business-critical environments; GitOps and CI/CD pipelines (GitHub Actions or similar); access and security tooling such as Teleport or Vault in production. Certifications. Relevant GPU/HPC, datacenter architecture, Linux, networking, Kubernetes, cloud, or security certifications (e.g. RHCSA/RHCE, CKA, NVIDIA-certified) are a plus. What We Can Offer You Highly competitive package, including base salary and equity, with reviews every 12 months. Join one of the fastest-growing tech startups: your chance to push boundaries, collaborate with brilliant minds, and make your mark on cutting-edge AI. Expect a dynamic progression plan tailored to your ambitions. Grow by trying new things, leading, challenging the status quo, and owning your impact, always with our full support. Human-first flexibility. We treat you as humans first. Our flexible workplace trusts Nscalers to deliver, giving you the autonomy to shape your day around life's moments. Join our thriving remote-first team. Geography is no barrier to impact or connection. We build seamless virtual collaboration, empowering you wherever you work. Equal Opportunities Statement At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enrich our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds. If there’s anything we can do to accommodate your specific situation, please let us know. Salary Range: $120,000 USD – $170,000 USD #J-18808-Ljbffr Nscale

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Senior Infrastructure Support Engineer in Seattle, WA vacancy
  • Crusoe Cloud is revolutionizing HPC by offering sustainable, low-cost GPU compute power. As a Senior Cloud Support Engineer, you will be the primary technical support contact, helping customers leverage Crusoe Cloud to achieve their research goals and accelerate development... 
    Senior

    Crusoe

    Bellevue, WA
    1 day ago
  • Crusoe Cloud is hiring a Cloud Support Engineer to empower customers with sustainable, low-cost GPU compute power. You’ll be the primary technical support contact, handling incidents, triage, and escalations while collaborating with SRE, Networking, and Storage teams.... 
    Senior
    Full time
    Remote work

    Crusoe

    Bellevue, WA
    1 day ago
  • Axon is seeking a Technical Support Engineer (Tier 2+) to support customers and internal teams for the Axon 911 platform. You will troubleshoot issues, maintain reliability, and partner with Engineering to deliver high‑quality customer experiences. The role emphasizes... 
    Senior
    Work at office

    Axon

    Seattle, WA
    3 days ago
  •  ...Labor & Industries (L&I) seeks an IT System Administrator - Senior Specialist to own statewide servers and infrastructure. You’ll ensure 24/7 availability, deploy physical, virtual, and cloud resources, and support Service Desk, DevOps, data management, and cloud... 
    Senior

    Washington State Department of Labor & Industries

    Seattle, WA
    10 hours ago
  • $145k - $175k

     ...energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each...  ...sustainable, low‑cost GPU compute power. As a Cloud Support Engineer, you'll play a crucial role in empowering our customers to... 
    Senior
    Full time
    Temporary work
    Remote work
    Shift work
    Weekend work

    Crusoe

    Bellevue, WA
    2 days ago
  •  ...DivergeIT in Bellevue, WA is seeking a Dedicated Support Engineer to serve as a trusted technical anchor for a key client environment. You will be the go-to troubleshooter, handling on-site and remote support, diagnosing issues and delivering measurable value. The role... 
    Senior
    Remote work

    DivergeIT

    Bellevue, WA
    2 days ago
  • $90.04k - $126.05k

    Salal Credit Union is hiring an IT Infrastructure Support Engineer to help keep our technology environment reliable, secure, and ready to support employees and members every day. In this role, you'll support the networks, systems, cloud services, and infrastructure that... 
    For contractors
    Work at office
    Local area

    Salal Credit Union

    Seattle, WA
    2 days ago
  • $79.2k - $209.5k

    We're looking for a Senior Software Engineer to help deliver the next generation of high-performance...  ...file storage powering Oracle Cloud Infrastructure (OCI). In this role, you'll design and...  ...develop distributed storage systems that support some of the world's most demanding... 
    Senior
    Temporary work
    Flexible hours

    Oracle Corporation

    Seattle, WA
    1 day ago
  •  ...Armada is hiring a Senior Software Engineer to design and implement on-premise CaaS and GPUaaS, and to help build a cloud-integrated marketplace on bare metal Kubernetes clusters. You will contribute to architecture, implement core services, and drive performance and... 
    Senior

    Armada

    Bellevue, WA
    2 days ago
  • $125k - $170k

     ...com to learn more. About the Role: As a Senior IT Engineer at Motive, you will help design, build...  ..., and improve scalable corporate IT infrastructure and services for both office-based and...  ...productivity, and extend IT’s ability to support the business. Build and maintain... 
    Senior
    Temporary work
    Work experience placement
    Work at office
    Remote work

    Doist

    Seattle, WA
    1 day ago
  •  ...JPMorgan Chase & Co. in Seattle seeks a Senior Principal Software Engineer to lead design and delivery of the JPMC-VPC platform, a software‑defined networking initiative aimed at modernizing the firm's network for resilience and scale. You will guide multiple agile... 
    Senior

    Jobleads-US

    Seattle, WA
    1 day ago
  • $135.2k - $306.4k

     ...distributed systems, networking, and AI infrastructure, driving architecture, design,...  ...optimization across software components that support thousands of GPUs and high-bandwidth...  ...architecture across multiple teams, mentor senior engineers, and help shape the roadmap for Oracle... 
    Senior
    Temporary work
    Flexible hours

    Oracle Corporation

    Seattle, WA
    3 days ago
  • $168k - $270.25k

     ...you’ll be immersed in a diverse, supportive environment where everyone is inspired...  ...Experience (NVEX) Solutions Engineering team is looking for a senior Computer or Software Engineer. This...  ...-X that link GPUs and AI compute infrastructure. Candidates must have a software... 
    Senior
    Full time
    Weekend work

    Nvidia

    Seattle, WA
    1 day ago
  •  ...profile and get back to you at the earliest. Role Role: IT Support Engineer (network and device migration) Onsite - Seattle WA 98104...  ...initiative. This project encompasses the migration of network infrastructure across 45 sites and the transition of over 7,000 endpoint... 
    Work at office
    Local area

    Net2Source (N2S)

    Seattle, WA
    1 day ago
  • A staffing solutions company is seeking an IT Support Engineer to assist with a critical network and device migration project in Seattle...  ...issues. Candidates should have experience in IT support and be familiar with network infrastructure. #J-18808-Ljbffr Net2Source (N2S)
    Local area

    Net2Source (N2S)

    Seattle, WA
    1 day ago
  • $99.4k - $150.3k

    Salesforce.com, inc. is looking for a Technical Support Specialist in Seattle, WA, to manage support requests and provide technical resources for customers. This role requires strong technical understanding and excellent communication skills. The ideal candidate will have... 
    Senior

    Salesforce.Com Inc

    Seattle, WA
    3 days ago
  • $142k - $220.5k

    Job DescriptionWe are seeking a Senior Network Engineer to join a high-performing network engineering...  ...of enterprise-grade network infrastructure. In this role you will serve as a subject...  ...proud to offer a variety of benefits to support employees and their families,... 
    Senior
    Full time
    Night shift

    Nordstrom

    Seattle, WA
    10 hours ago
  •  ...individual can thrive.Job Description: Senior Network EngineerPosition Title: Senior Network...  ...the RoleWe are seeking a Senior Network Engineer who brings deep technical expertise,...  ...openly, communicates respectfully, supports others, and contributes to a positive, inclusive... 
    Senior
    Full time
    Local area

    F5 Networks

    Seattle, WA
    10 hours ago
  •  ...A leading tech recruitment firm is seeking a Network Engineer in Seattle for a 6-month contract. The role involves delivering wireless...  ...network planning and implementation services, providing operational support for complex deployments, and leading design projects.... 
    Senior
    Contract work

    Randstad

    Seattle, WA
    2 days ago
  • $99.4k - $150.3k

    Salesforce is seeking a Technical Support Specialist in Seattle to join our Technical Support team. You will manage support requests and work closely with our Success Architect team, ensuring we provide top-notch support. This role is crucial for maintaining customer satisfaction... 
    Senior

    Salesforce

    Seattle, WA
    3 days ago
  • $91k - $152k

     ...problems? Do you have a deep passion and desire to engineer and operate the world's largest cloud computing infrastructure to build a better world for future generations?...  ...who provide troubleshooting and operations support, and innovate to automate operational tasks.About... 
    Flexible hours
    Shift work
    Night shift
    Weekend work

    AmazonWebServices

    Seattle, WA
    1 day ago
  • A global technology company is seeking a Senior Software Engineer - Substrate in Seattle to design and build managed Kubernetes product offerings...  ...have strong programming skills in Go and experience with infrastructure automation tools like Terraform. Successful candidates... 
    Senior
    Flexible hours

    Palantir Technologies

    Seattle, WA
    1 day ago
  • $60 - $70 per hour

    Cloud Support EngineerLocation: Seattle, WA (Hybrid, 3 Days Onsite)Seeking a Cloud Operations Engineer to support and maintain AWS infrastructure and cloud platforms supporting 100+ production applications. This role combines cloud operations, platform administration, incident... 
    Contract work
    Temporary work

    TEKsystems

    Seattle, WA
    10 hours ago
  •  ...Development EngineerAs a Principal Network Development Engineer within Oracle Cloud Infrastructure (OCI) Network Reliability Engineering (NRE), you will...  ...operational best practices; participating in a rotational support model, including on-call responsibilities and incident... 
    Senior

    Oracle

    Seattle, WA
    2 days ago
  • $130k - $195k

     ...ubiquitous. We build the foundation for agent engineering in the real world, helping developers...  ...The Role We’re hiring a Technical Support Engineer to lead our customer support experience...  ...technical users, from AI engineers to infrastructure architects. You’ll be on the front... 
    Senior
    Remote work
    Flexible hours

    LangChain

    Seattle, WA
    10 hours ago
  • $126.2k - $264.1k

    What You’ll DoAs a Senior Principal Network Reliability Engineer, you will:Provide technical leadership for the...  ...management of OCI’s global network infrastructure.Define and drive the long-term...  ...complex engineering initiatives supporting AI infrastructure, new OCI Regions... 
    Senior
    Temporary work
    Flexible hours

    Oracle Corporation

    Seattle, WA
    3 days ago
  •  ...CloudLambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range...  ...be part of day2 operations and on-call rotation for Network Engineering teamYouHave 10+ years of experience in IT and networking... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Corporation

    Bellevue, WA
    2 days ago
  • $75k - $215k

     ...Senior EngineerGEICO is seeking an experienced Engineer with a passion for building high-performance, low maintenance...  ...issuesProvide 24x7 after hours on-call support and support for off-hour...  ...SQLFamiliarity with full stack network infrastructure functions including, but not... 
    Senior
    Hourly pay
    Work experience placement
    Shift work

    GEICO

    Seattle, WA
    1 day ago
  • A leading tech firm is looking for a Senior Principal Engineer to lead platform architecture and strategic planning. In this role, you will define the technical vision across multiple engineering organizations, ensuring a seamless integration of product systems. Ideal candidates... 
    Senior
    Remote job

    Docker

    Seattle, WA
    2 days ago
  • $155.4k - $210.2k

     ...reliable, scalable, low-cost infrastructure platform in the cloud that...  ...commercial and technically savvy Senior Backbone Network Developer...  ...and backbone network engineering. This role requires collaboration...  ...keep the cloud running. We support all AWS data centers and all... 
    Senior
    Worldwide
    Flexible hours

    Amazon

    Seattle, WA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Infrastructure Support Engineer. Be the first to apply!