Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Infrastructure Support Engineer

$120k - $170k

Nscale

About Nscale Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack—energy, data centres, GPU superclusters, orchestration, and AI services—delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world. About The Role (Job Purpose) Senior Infrastructure Support Engineers are the senior technical escalation point within Infrastructure Support, owning the health of Nscale's GPU fleets and the high-performance fabrics that connect them. This is a hands‑on L2/L3 role operating at the intersection of GPU hardware, east‑west networking, Linux, and data centre operations—acting as the operational bridge between Support, DC Operations, and Engineering. You Will Own complex, ambiguous problems end-to-end and make decisive calls in a results‑driven environment, taking calculated risks where speed matters. Communicate technical detail clearly, specifically, and concisely— to engineers, to customers, and to leadership. We treat communication quality as a core engineering skill, not a soft skill. Influence without authority and build strong relationships with senior stakeholders across the business to get things done. Grasp new technical concepts quickly, stay curious, and know which questions to ask to get up to speed fast. Bring discipline and organisation: evidence‑led investigations, accurate records, clean handovers. What You’ll Be Doing (Responsibilities) Join the Support duty rotation as a senior escalation point, collaborating with Infrastructure Engineering, CNPRE, Network Operations, and Product Engineering on incidents, investigations, and changes. Diagnose and remediate GPU node faults across the full stack—driver, firmware, and hardware layers—from nvidia-smi/DCGM and XID/RAS analysis through BMC/Redfish and out-of-band management to physical fault isolation and vendor RMA. Own east‑west fabric health: run link-level diagnostics (mlxlink, ibdiagnet or equivalent), isolate transceiver, optics, cabling, and switch‑port faults, and validate topology across InfiniBand and RoCE/high-speed Ethernet fabrics. Investigate data-path issues on high‑performance storage platforms (e.g. VAST), including storage–network interactions across clients, mounts, VIPs, and routing. Run structured, hypothesis-driven investigations; conduct root cause analysis for major incidents and drive long-term fixes to completion. Author and execute changes in live customer environments with proper risk assessment, peer review, and backout plans. Proactively improve dashboards, alerts, and runbooks to prevent repeat incidents; identify recurring patterns and convert them into problem records and automation. Accurately record, update, and resolve tickets, keeping internal and external parties informed with clear customer-impact statements and evidence-rich notes that enable clean handover. Design and implement automation scripts and small tools to reduce toil and human intervention. Act as a key escalation point for the Support Organisation, taking ownership of strategic decisions where results matter. Mentor and upskill mid-level engineers; contribute to knowledge sharing across Operations and Engineering, including training content, workshops, and PR reviews. Lead by earning trust and speaking candidly. Disagree when appropriate and challenge the status quo; commit wholly to decisions once in motion. Respond to critical incidents out of business hours and participate in on-call as required. Travel to Nscale or customer sites to provide onsite technical expertise. About You (Skills / Qualifications Experience) Experience. 6+ years in infrastructure, operations, or support engineering in production environments; 2–3+ years hands‑on with GPU, HPC, or large-scale data centre estates, ideally in a customer-facing or escalation-driven capacity. Communication. Able to explain complex technical detail clearly, specifically, and concisely— in tickets, in incident updates, and face to face with customers and stakeholders at all levels. Strong written discipline: your notes let the next engineer pick up where you left off without starting from scratch. GPU platforms (NVIDIA; AMD Instinct beneficial). Practical, current experience with GPU drivers, firmware, and runtime stacks on AI training and inference clusters. Confident with nvidia-smi, DCGM, and XID/error interpretation; able to isolate faults across GPU, baseboard, NIC, and PCIe layers and drive them through diagnosis to RMA. High-performance east-west fabrics. Hands‑on experience with RDMA fabrics—InfiniBand and/or RoCE—including link-layer diagnostics (mlxlink, ibdiagnet or equivalent), transceiver and cabling fault isolation, and understanding of rail‑optimised topologies, NVLink/NVSwitch, and NCCL-based performance troubleshooting on multi-node clusters. HPC scheduling. Slurm operations for large multi-GPU jobs—containers via Pyxis/Enroot, MPI, and diagnosing queue, topology, and job failures. Linux systems engineering at scale. Strong command of modern Linux distributions, kernel modules, systemd, networking stack, and filesystem tooling. Proven troubleshooting across compute, storage, and network layers in production. Server hardware and control planes. Comfortable with BMC/Redfish, firmware management, and bare-metal provisioning workflows (MAAS or similar) across large node fleets. Networking fundamentals. Solid grasp of L2/L3, routing, BGP, VLANs, VXLAN, firewalls, and load balancing, with a clear understanding of how east-west cluster traffic differs from north-south. Observability and incident response. Build and use alerting stacks and dashboards (Prometheus/Grafana or similar), interpret metrics and alerts, drive runbooks to resolution, and contribute to SLOs and post-incident reviews. Change and risk judgment. Experience authoring and executing changes in business-critical environments, including risk assessments, customer-impact analysis, and backout plans. SRE-style operations. Write and maintain runbooks, automate diagnostics, and reduce human intervention through scripts and small tools. Automation and Git. Scripting skills in Bash, Python or equivalent for operational tooling and integrations; experience with infrastructure automation tools (Ansible, Terraform or similar). Data centre fundamentals. Understanding of how data centres operate—servers, networks, storage, power, and cooling—ideally gained through an operational support background. Leadership. Disciplined, organised, and self-motivated, with the ability to mentor and motivate other engineers, take decisive action, and drive the team and wider organisation to improve. Adaptability. Able to adapt to customer-driven demands, including specialist support outside core hours and travel for onsite work. Nice to Have High-performance storage. Hands‑on experience with VAST or comparable AI-optimised storage platforms, or Ceph/parallel filesystems and NFS at scale (multipath, remoteports, nconnect), including diagnosing storage–network interaction and data-path performance issues. OpenStack and fleet operations tooling. OpenStack operations experience (Neutron, Cinder, error triage), plus familiarity with fleet-scale tooling for provisioning, health, and remediation across large GPU estates (MAAS, NetBox, Redfish-driven automation or similar). Kubernetes. Operating and troubleshooting clusters, including GPU operator stacks and understanding how physical resources are abstracted up the stack. Helpful context for our platform, though not the core of this role. Automation at scale. Automated network configuration with safe, repeatable changes in business-critical environments; GitOps and CI/CD pipelines (GitHub Actions or similar); access and security tooling such as Teleport or Vault in production. Certifications. Relevant GPU/HPC, datacenter architecture, Linux, networking, Kubernetes, cloud, or security certifications (e.g. RHCSA/RHCE, CKA, NVIDIA-certified) are a plus. What We Can Offer You Highly competitive package, including base salary and equity, with reviews every 12 months. Join one of the fastest-growing tech startups: your chance to push boundaries, collaborate with brilliant minds, and make your mark on cutting-edge AI. Expect a dynamic progression plan tailored to your ambitions. Grow by trying new things, leading, challenging the status quo, and owning your impact, always with our full support. Human-first flexibility. We treat you as humans first. Our flexible workplace trusts Nscalers to deliver, giving you the autonomy to shape your day around life's moments. Join our thriving remote-first team. Geography is no barrier to impact or connection. We build seamless virtual collaboration, empowering you wherever you work. Equal Opportunities Statement At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enrich our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds. If there’s anything we can do to accommodate your specific situation, please let us know. Salary Range: $120,000 USD – $170,000 USD #J-18808-Ljbffr Nscale

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior Infrastructure Support Engineer in Seattle, WA vacancy
  • $75k - $215k

     ...we want you to feel valued, supported, and proud to work here. That...  ...is seeking an experienced Engineer with a passion for building...  ...Position Description Our Senior Engineer is a key member of...  ...Familiarity with full stack network infrastructure functions including, but not... 
    Senior
    Hourly pay
    Full time
    Work experience placement
    Local area
    Shift work

    GEICO

    Seattle, WA
    3 days ago
  • $100k - $140k

     ...Nscale Nscale is the GPU cloud engineered for AI. We provide cost‑effective, high‑performance infrastructure for AI start‑ups and large...  ...technical capabilities and directly supports strategic business outcomes,...  ...training materials. Shadow seniors during complex work to build... 
    Suggested
    Flexible hours

    Nscale

    Seattle, WA
    3 days ago
  • $130k - $195k

     ...ubiquitous. We build the foundation for agent engineering in the real world, helping developers...  ...The Role We’re hiring a Technical Support Engineer to lead our customer support experience...  ...technical users, from AI engineers to infrastructure architects. You’ll be on the front... 
    Senior
    Remote work
    Flexible hours

    LangChain

    Seattle, WA
    4 days ago
  •  ...for people like you. As a Senior Principal Architect at...  ....g., incident investigation support and knowledge capture) with...  ...or certification on software engineering concepts and 10+ years applied...  ...developing and operating large‑scale Infrastructure and Platform as a Service,... 
    Senior

    JPMorgan Chase & Co.

    Seattle, WA
    4 days ago
  •  ...introduction of new Steam and other gaming hardware. Supported by experienced mentors from across the industry, we are...  ...that! The Opportunity KOMODO is seeking a Senior Networking Engineer to join the team working on Nanos, the codename for a competitive... 
    Senior
    Summer holiday

    Komodo Co., Ltd.

    Bellevue, WA
    3 days ago
  • A leading tech recruitment firm is seeking a Network Engineer in Seattle for a 6-month contract. The role involves delivering wireless...  ...network planning and implementation services, providing operational support for complex deployments, and leading design projects.... 
    Senior
    Contract work

    Randstad

    Seattle, WA
    3 days ago
  • $80.9k - $122.3k

     ...future of Salesforce. This Government Cloud Weekend Nightforce Support Engineer role provides the highest level of overnight support and...  .... Liaise and work closely with the Salesforce R&D and Infrastructure teams on escalated technical issues and product roadmap changes... 
    Work experience placement
    Night shift
    Weekend work

    Salesforce.Com Inc

    Bellevue, WA
    2 days ago
  • $160k - $185k

     ...seeking a highly skilled and motivated Sr. Infrastructure Engineer to join our Hardware Engineering Dev...  ...: Incident Management & Support: Assist in incident response efforts...  ...quickly, working under the guidance of more senior engineers. Help document incidents,... 
    Senior
    Full time
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours

    Coreweave

    Bellevue, WA
    1 day ago
  •  ...GEICO is looking for a talented Engineer in Seattle to enhance our network infrastructure and implement security measures in alignment with our tech transformation goals. You will lead the strategy and execution of technical roadmaps, ensuring network performance and security... 
    Senior

    GEICO

    Seattle, WA
    1 day ago
  •  ...qualifications, experience, interview performance, and location. In addition to a competitive salary, we offer a variety of benefits to support your needs, including: ~ Medical, dental, and vision insurance - 100% paid for by CoreWeave ~ Company-paid Life Insurance  ~... 
    Senior
    Full time
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours

    Coreweave

    Bellevue, WA
    1 day ago
  • $135k - $200k

     ...missing children, and more. The Role As a Senior Software Engineer on Network Infrastructure you will be joining a team whose mission is to make...  ...• 10 paid holidays throughout the calendar year • Supportive leave of absence program including time off for military... 
    Senior
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package

    Palantir Technologies

    Seattle, WA
    1 day ago
  • $125k - $145k

     ...and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each...  ...offering sustainable, low-cost GPU compute power. As a Senior Cloud Support Engineer, you'll play a crucial role in empowering our customers... 
    Full time
    Temporary work

    Crusoe

    Seattle, WA
    4 days ago
  •  ...SecurityScorecard Inc. in Seattle is seeking an experienced Senior Engineer to lead technical initiatives and build scalable systems. You will work across the stack, optimizing performance and collaborating with various teams to drive platform evolution. The ideal candidate... 
    Senior

    SecurityScorecard

    Seattle, WA
    3 hours ago
  • $300k

     ...Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to...  ...robust abstractions that work across providers, and make smart infrastructure decisions that keep us cost-effective at massive scale. Your... 
    Senior
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    Seattle, WA
    1 day ago
  •  ...About the Platform Team The Platform Engineering team is the invisible engine that powers...  ...Finance. Our mission is to build the core infrastructure, tools, and shared services that allow...  ...mission‑critical systems that directly support close to 1,000 engineers building... 
    Senior
    Work at office
    3 days per week

    Rippling

    Seattle, WA
    2 days ago
  • $193.6k - $414.4k

     ...At Oracle Cloud Infrastructure (OCI), we build the future of the cloud for Enterprises...  ...company in the world. This Senior Director of Network Engineering will be the business leader and service...  ...and troubleshooting processes to support our initiatives, ensuring the successful... 
    Senior
    Temporary work
    Part time
    Flexible hours

    Oracle Corporation

    Seattle, WA
    2 days ago
  • $320k

     ...Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to...  ...beneficial AI systems. About the role We're building the infrastructure that lets people talk to Claude—real-time, bidirectional... 
    Senior
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    Seattle, WA
    1 day ago
  • $168k - $252k

     ..., define processes, and develop frameworks to allow Anduril’s engineers and operators to execute at all stages of the software development...  ...other security practitioners and the software platform team, supporting efforts to improve Anduril’s security posture while delivering... 
    Senior
    Full time
    Work experience placement
    Local area
    Relocation package

    Anduril Industries

    Seattle, WA
    1 day ago
  • The Allen Institute for Artificial Intelligence in Seattle is seeking a Senior Engineer to design foundational architecture for AI research agents. The ideal candidate has 8+ years of experience, strong Python skills, and a background in developing production services and... 
    Senior

    The Allen Institute for Artificial Intelligence

    Seattle, WA
    5 days ago
  • Indeed is hiring a Software Engineer III in Seattle, WA, to drive development of platform services connecting AI products to language...  ...in a related field and substantial experience in cloud infrastructure and distributed systems. The position offers competitive salary... 
    Senior

    Indeed

    Seattle, WA
    5 days ago
  • $145.4k - $174.8k

     ...innovative and cost-effective engineering solutions for data centers,...  ...you! ***This is a physical infrastructure design role — not software...  ...seeking an RCDD-certified Senior ICT Infrastructure Design...  ...Adjacent Systems Integration Support the design and coordination... 
    Senior
    For contractors
    For subcontractor
    Work at office
    Local area

    Meade

    Seattle, WA
    4 days ago
  •  ...technology company is seeking an experienced Sr .NET Full Stack Developer with strong expertise in Content Management Systems (CMS) to support enterprise-grade web applications for a Microsoft client. Your role will involve designing, developing, and maintaining... 
    Senior
    Remote job

    Galactic Minds INC

    Bellevue, WA
    4 days ago
  •  ...revolutionizing the lending landscape. SoFi is seeking enthusiastic Senior Software Engineers who are ready to lead the development of key advancements...  .... Experience with public cloud compute, storage, and infrastructure. Experience with Kafka, Docker, Kubernetes, and Spring... 
    Senior
    Full time
    Work experience placement
    Remote work

    SoFi

    Seattle, WA
    4 days ago
  • $182k - $222k

     ...Employees, researchers, customers, and partners Win Together by fostering empowerment, inclusion, respect, and accountability. Senior Software Engineer, Platform Location: Seattle, WA; Austin, TX; Boston, MA; Washington, DC 1 day onsite per week Position Summary At... 
    Senior
    Apprenticeship
    Work at office
    Local area
    Remote work
    Flexible hours
    Shift work
    1 day per week

    HackerOne

    Renton, WA
    3 days ago
  • $28 - $37 per hour

     ...community. Our services include both: On-demand (as-needed) IT support Managed IT for small business clients (1-100 employees)...  ...humans - you'll fit right in. The Role We're hiring a Senior IT Technician with 5+ years of hands-on experience across hardware... 
    Senior
    Hourly pay
    Full time
    Work at office
    Immediate start

    NerdsToGo Inc

    Bellevue, WA
    3 days ago
  • $83.41k - $168.59k

     ...career in Advisory. KPMG is currently seeking a Senior Associate, Infrastructure Project Advisory (Construction/Engineering) in Infrastructure and Projects Advisory for...  ..., roles, responsibilities, reporting, and supporting information technology Conduct project reviews... 
    Senior
    Contract work
    H1b
    Local area

    Stryker Corporation

    Seattle, WA
    2 days ago
  •  ...systems that power our products, enable our engineers, and keep our platform infrastructure reliable as we grow. As a Senior Software Engineer on the Platform Team, you will...  ...services and the AWS infrastructure that supports them Develop and scale application backends in... 
    Senior

    COMPA

    Seattle, WA
    4 days ago
  • $85k - $105k

     ...Senior IT Technician (SR Tech) Consistently recognized as a best workplace, and for...  ...are looking to be a part of an open, supportive team and receive exciting challenges that...  ..., mobile devices, applications, infrastructure components, cloud services, and core business... 
    Senior
    Temporary work
    Work at office
    Immediate start

    BNBuilders

    Seattle, WA
    3 days ago
  • $140k - $160k

     ...Europe, Africa, and Asia, Standish supports investment managers and their investors...  ...Standish Management is seeking a Senior IT Solutions Engineer to join the IT organization. This is...  ...needed Collaborate with the MSP on infrastructure, security operations, and platform services... 
    Senior
    Work at office

    Standish Management, LLC

    Seattle, WA
    5 days ago
  • $102.5k - $187.9k

    Location: Anywhere in Country Job Overview Senior Technology Analyst - Oracle Services Supply Chain sub-practice. Supports large, complex full lifecycle initiatives in Oracle applications and technology, advising clients on architecture and implementation of core applications... 
    Senior
    Flexible hours

    EY

    Seattle, WA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Infrastructure Support Engineer. Be the first to apply!