Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Infrastructure Engineer, AI Cluster Performance & Validation

Nscale

Principal Infrastructure Engineer, AI Cluster Performance & Validation Houston; New York; San Francisco; Seattle Overview As a Principal Infrastructure Engineer, AI Cluster Performance & Validation, you will be a critical member of the AI Infrastructure Operations team, responsible for ensuring the acceptance, performance, and scalability of our cutting-edge AI and High-Performance Computing (HPC) environments. Leveraging software engineering and testing principles, you will focus on building and maintaining the control plane, tooling, and automation that supports performance and validation testing of large-scale AI clusters. Your work will directly translate into higher system availability, compute optimization and reduced operational costs. Key Responsibilities Own the technical definition of "healthy at scale." Set the architecture, roadmap, and acceptance criteria by which multi-thousand-GPU clusters are declared production-ready, and establish the performance bar (collective bandwidth, job goodput, model FLOPs utilization) that every cluster must clear before and after customer handover. Establish technology and product direction in collaboration with other tech leads, managers, and senior leadership. Run and instrument real AI workloads as a diagnostic instrument. Stand up and execute distributed training and inference jobs — open-source and customer-representative models — across thousands of accelerators to validate cluster behavior under genuine load rather than synthetic proxies alone, and translate what those runs reveal into fleet-wide fixes. Lead deep diagnosis of large-scale cluster failures and performance regressions , isolating root cause across the full stack: GPU and NIC firmware, PCIe/NVLink topology and NUMA placement, InfiniBand/RoCE fabric health, congestion control and routing, storage and data-loader throughput, scheduler placement, and framework/communication-library behavior. Serve as the final escalation point for the hardest slow-job and stalled-job investigations. Design and build the validation and burn-in systems that qualify nodes, racks, and full pods at scale — NCCL/RCCL collective sweeps, HPL/HPCG and MLPerf-style benchmarks, thermal and power soak tests, straggler and flapping-link detection — and automate them so that qualification is a repeatable pipeline, not a manual campaign. Drive cluster optimization end to end , tuning fabric configuration (adaptive routing, QoS and congestion control, SHARP in-network reduction, rail and topology-aware placement), collective communication libraries and algorithm selection, GPUDirect RDMA and storage paths, and host-level settings (huge pages, IRQ affinity, CPU governors, MIG and driver configuration) to convert raw hardware into delivered throughput. Partner with Infrastructure, Platform, SRE, and customer-facing teams to translate operational and customer performance needs into durable engineering solutions, and to feed diagnostic signal back into provisioning, remediation, and capacity workflows. Build production-grade Python systems and performance tooling for automated triage, telemetry correlation, and regression detection, leveraging AI tools to accelerate delivery. Assess impact to the team's software and validation stack from new hardware product programs, and explore AI-driven process improvement and automation. Establish engineering standards for reliability, observability, benchmarking methodology, and operational excellence across all services, and raise the diagnostic capability of the wider organization through mentorship, runbooks, and post-incident technical write-ups. Required Qualifications Education: Bachelor's or higher degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience. Experience: 10+ years of relevant experience building, operating, or debugging large-scale compute infrastructure, including significant time at staff or principal level owning cross-team technical direction. AI Workload Expertise: Hands-on experience running real AI compute jobs at scale — pre-training, fine-tuning, or large-scale inference of open-source or proprietary models — including practical familiarity with distributed training strategies (data, tensor, pipeline, and expert parallelism) and frameworks such as PyTorch, Megatron-LM, DeepSpeed, or equivalent. Cluster Validation: Demonstrated experience validating and accepting large clusters (thousands of GPUs) for performance and reliability, with a working command of benchmark methodology and the ability to defend a number to both engineers and customers. Performance Debugging: Proven ability to diagnose distributed performance problems — stragglers, collective stalls, link flaps, thermal throttling, silent data corruption, ECC and Xid errors, noisy-neighbor and storage-bound bottlenecks — using tools such as NCCL debug tracing, Nsight Systems/Compute, PyTorch Profiler, perf, and fabric telemetry. Networking: Deep understanding of high-performance fabrics — InfiniBand and/or RoCEv2, RDMA, GPUDirect, adaptive routing, congestion control, and rail-optimized topologies — and of networking fundamentals (TCP/IP, BGP). Systems & Programming: Deep Linux systems expertise (kernel tunables, NUMA, PCIe, IRQ and memory behavior) and strong production Python, plus experience with C/C++ or Go and with configuration management tooling (e.g., Ansible, Terraform). Schedulers: Experience operating and debugging AI workloads under SLURM and/or Kubernetes at scale. Preferred Qualifications Master's degree or PhD in Engineering, Computer Science, or a related technical field. Experience bringing up and qualifying a greenfield GPU supercluster from first rack to production traffic, including firmware, driver, and topology standardization across a heterogeneous fleet. Direct experience with NVIDIA GPU platforms (H200/GB200/GB300-class), NVLink and NVSwitch domains, DCGM, SHARP, UFM, and the NVIDIA software stack; or equivalent depth on AMD Instinct and ROCm/RCCL. Published or presented benchmark, scaling, or post-mortem work — MLPerf submissions, scaling studies, or public technical write-ups on large-cluster behavior. Experience with advanced observability and monitoring systems (Prometheus, Grafana, OpenTelemetry) applied to high-cardinality GPU and fabric telemetry, including automated anomaly and regression detection. Experience with high-throughput parallel storage (Lustre, GPFS, WEKA, VAST) and with diagnosing data-pipeline-bound training jobs. Familiarity with cloud-native technologies (Kubernetes, Docker), infrastructure-as-code principles, and integration with infrastructure tooling such as DCIMs, NetBox, and bare metal APIs (MAAS, Ironic, IPMI, Redfish). Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements). Familiarity with SLOs/metrics measurement and logs/telemetry/metrics integration with tools for enhanced operator experience. For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice:Here. #J-18808-Ljbffr Nscale

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Principal Infrastructure Engineer, AI Cluster Performance & Validation in Seattle, WA vacancy
  • $270k - $330k

     ...Principal Network Engineer, AI Infrastructure & High-Performance Networking About Nscale Nscale is the GPU cloud engineered...  ...responsible for the design, validation, and ongoing operation of all...  ...bare-metal provisioning and cluster management systems. Own... 
    Performance
    Full time
    Flexible hours

    Nscale

    Seattle, WA
    1 day ago
  • $135.2k - $306.4k

     ...distributed systems, networking, and AI infrastructure, driving architecture, design, implementation, and performance optimization across software...  ...teams, mentor senior engineers, and help shape the roadmap...  ...communication across GPU clusters.Drive the evolution of collective... 
    Performance
    Temporary work
    Flexible hours

    Oracle Corporation

    Seattle, WA
    2 days ago
  • $121.5k - $306.4k

    Oracle Cloud Infrastructure (OCI) is building some of...  ...most advanced GPU clusters to power the next generation of AI. The Strategic Customer Engineering (SCE) Core Infrastructure...  ...reliability, performance, and scalability for...  ...development, testing, validation, and deployment... 
    Performance
    Temporary work
    Flexible hours

    Oracle Corporation

    Seattle, WA
    1 day ago
  • $102.3k - $209.5k

     ...global leaders in the RDMA cluster networking domain and...  ..., accelerated High-Performance Compute (HPC),...  ...tailored specifically for AI, ML, HPC workloads.We...  ...scale global Oracle Cloud Infrastructure (OCI). Primarily focused...  ...in CS or related engineering field with 6+ years of... 
    Performance
    Temporary work
    Immediate start
    Flexible hours

    Oracle Corporation

    Seattle, WA
    2 days ago
  • $84.9k - $209.5k

    As a Principal Core Infrastructure Engineer (AI/ML Forward Deployed Infrastructure Engineer), you will play a critical role in designing, implementing...  ...will be crucial in optimizing our infrastructure for performance, reliability, and cost-effectiveness. In this hands-... 
    Performance
    Temporary work
    Flexible hours
    Shift work

    Oracle Corporation

    Seattle, WA
    1 day ago
  • $193.6k - $414.4k

    At Oracle Cloud Infrastructure (OCI), we build the future...  ...Director of Network Engineering will be the business...  ...organizational planning, performance management, and have...  ...qualification and validation processes to ensure optimal...  ...care. And with AI embedded across our products... 
    Performance
    Temporary work
    Flexible hours

    Oracle Corporation

    Seattle, WA
    4 days ago
  • $260k - $320k

    Director of Data Engineering &...  ...develop high-performing Data Engineering...  ...serving the firm’s AI and analytics...  ...defined by the Principal Data Architect...  ...and structural validation are established...  ...tier design, job cluster and SQL...  ...gateway and VM infrastructure coordination,... 
    Performance
    Permanent employment
    Full time
    Contract work
    Temporary work
    For contractors
    Work at office
    Immediate start
    Flexible hours
    Weekend work

    Cooley

    Seattle, WA
    7 hours ago
  • $153k - $204k

     ...Essential Cloud for AI™. Built for pioneers...  ...CoreWeave combines superior infrastructure performance with deep technical...  ...'ll Do The Systems Engineering team owns the host...  ...broaden what it can validate, and keep CI fast and...  ...experience. HPC or large-cluster experience —... 
    Performance
    Permanent employment
    Full time
    Temporary work
    Casual work
    Live in
    Work at office
    Flexible hours

    CoreWeave

    Bellevue, WA
    3 days ago
  • $114.6k - $234.6k

     ...same rate. As a Principal Member of Technical...  ...Lightweight Infrastructure (LWI), the next-generation...  ...of software engineering experience building...  ...production and performance issues. Experience...  ...designs complex validation (fault injection,...  ...saving care. And with AI embedded across... 
    Performance
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle Corporation

    Seattle, WA
    7 hours ago
  • $206k - $303k

     ...Essential Cloud for AI™. Built for pioneers...  ...CoreWeave combines superior infrastructure performance with deep technical...  ...of the largest GPU clusters in the world. The AI...  ...pressure. As a Principal Engineer in AI Infrastructure...  ...workload admission, validation, and rollout,... 
    Performance
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Seattle, WA
    19 days ago
  • $193.6k - $414.4k

     ...Description Lead a Network Engineering group responsible for...  ..., security, performance, and cost outcomes. Own...  ...Automation, CI/CD, and Infrastructure as Code: Governs...  ...policy-as-code, and validation platforms. Sets target...  ...saving care. And with AI embedded across our products... 
    Performance
    Temporary work
    Flexible hours

    Oracle

    Seattle, WA
    2 days ago
  • $135.2k - $306.4k

     ...Oracle Cloud Infrastructure (OCI) delivers mission...  ...to enhance engineering efficiency by concentrating...  ...with high performance that can be adopted...  ...deliver a database validation framework that...  ...Responsibilities As a Senior Principal Engineer, you...  ...care. And with AI embedded across... 
    Performance
    Temporary work
    Worldwide
    Flexible hours

    Ll Oefentherapie

    Seattle, WA
    4 days ago
  • $272k - $431.25k

     ...the unlimited potential of AI to define the next era of...  ...device into a high-performance gaming rig. GeForce NOW automatically...  ...are currently seeking a Principal Systems Software Engineer to join our team,...  ...related CI/CD, patching infrastructure, and improving it to take... 
    Performance
    Full time
    Remote work

    Nvidia

    Seattle, WA
    2 days ago
  • $180.5k - $225.6k

     ...running the world's best AI and data infrastructure platform so our customers...  ...their business. Founded by engineers — and customer obsessed —...  ...Converge disparate CPU and GPU cluster management paths into a...  ...include eligibility for annual performance bonus, equity, and the... 
    Performance
    Local area
    Remote work
    Worldwide

    DataBricks

    Bellevue, WA
    2 days ago
  • Veeda AI is building the next generation of multimodal foundation world models for Physical AI. The Member of Technical Staff - AI Infrastructure will design, deploy, and operate high-performance GPU clusters running Slurm and Kubernetes. You will own interconnect fabric... 
    Performance

    Veeda AI

    Seattle, WA
    3 days ago
  • $222k - $300k

     ...world's best data and AI infrastructure platform so our...  ...business. Founded by engineers — and customer obsessed...  ...stronger guardrails, validation, and graduated rollout...  ...team into a Staff and Principal attracting org whose...  ..., and retaining high performing teamsComfortable partnering... 
    Performance
    For contractors
    Local area
    Worldwide

    DataBricks

    Bellevue, WA
    2 days ago
  •  ...DescriptionPrincipal Consultant - AI Platform & Governance Engineer - Infosys...  ...Advisory Practice is seeking a Principal Consultant specializing...  ...protocol.DevOps: CI/CD, infrastructure-as-code (Terraform/Helm)...  ...and monitor key performance indicators (KPIs) to track... 
    Performance
    Full time
    Temporary work
    Local area

    Infosys Technologies

    Seattle, WA
    4 days ago
  • $228.7k - $306.7k

    Job Posting Title:Senior Principal Machine Learning Engineer, Ad PlatformsReq ID:10150390Job Description...  ...- driving advertising performance, innovation, and value in Disney...  ...Machine Learning and AI patterns, platforms and infrastructure, and leadership skills to unblock... 
    Performance
    Full time

    Hulu

    Seattle, WA
    2 days ago
  • Nscale seeks a Principal Infrastructure Engineer to own AI cluster performance and validation across thousands of GPUs. You’ll define healthy-at-scale criteria, set the architecture, and partner with senior leaders to production-readiness goals. You’ll run real AI workloads... 
    Performance

    Nscale

    Seattle, WA
    2 days ago
  • $172.6k - $259k

     ...in!We are looking for a Principal Network Engineer to act as the technical lead...  ...of Network & Cloud Infrastructure and work alongside IAM, Cloud...  ...implementation.Build and validate the full ZTNA architecture...  ...for satisfactory performance in this position. It is not... 
    Performance
    Local area

    Hasbro

    Renton, WA
    7 hours ago
  •  ...discovery to powering AI and the...  ...Customer Solutions Engineer on AMD's Applied...  ...AMD Instinct GPU clusters from delivery to...  ...workload deployment and performance, production...  ...onboarding, performance validation, and sustained...  ...production software or infrastructure engineering,... 
    Performance
    Permanent employment
    Flexible hours

    AMD

    Bellevue, WA
    7 hours ago
  • $120k - $170k

     ...the vertically integrated AI cloud engineered for AI. We own and...  ...services—delivering high-performance infrastructure to AI-native companies, enterprises...  ...switch‑port faults, and validate topology across...  ...AI training and inference clusters. Confident with nvidia-smi... 
    Performance
    Remote work
    Flexible hours

    Nscale

    Seattle, WA
    2 days ago
  • $191k - $297k

    Job DescriptionSr Engineering Manager - Network InfrastructureSNAAP...  ...OCI, Juniper Mist wireless infrastructure, and Infoblox DNS/IPAM. You’...  ...and site typesAnalyze network performance and cost data to optimize routing...  ...EX/QFX switching, Mist AI wireless, or comparable campus... 
    Performance
    Full time

    Nordstrom

    Seattle, WA
    4 days ago
  • $184k - $287.5k

     ...rack, multi-tenant AI/ML datacenters...  ...Senior Software Engineer for our CSP (Cloud...  ...you’ll be doing:Perform deep-dive debugging...  ...rack, multi-tenant clusters: scheduler...  ...environments; automate validation and benchmark...  ...OpenTelemetry), and infrastructure-as-code.Excellent... 
    Performance
    Full time
    Remote work

    Nvidia

    Seattle, WA
    7 hours ago
  •  ...alerts; and designs complex validation (fault injection, brownouts...  ..., and processing. Design performance and load testing plans. Build...  ...gaps. Ensure cloud infrastructure complies with industry standards...  ...of experience in security engineering, application development, or... 
    Performance
    Temporary work
    Flexible hours

    Oracle

    Seattle, WA
    2 days ago
  •  ...Docker is seeking a Senior Principal Engineer to serve as the technical visionary...  ..., Billing, Data, Platform Infrastructure, Developer Tools and...  ...pricing, our expansion into AI and security products, and...  ...infrastructure while maintaining performance and reliability standards... 
    Performance
    Contract work
    Immediate start
    Remote work
    Home office

    Docker, Inc.

    Seattle, WA
    1 day ago
  • $207k - $275k

     ...Essential Cloud for AI™. Built for pioneers...  ...CoreWeave combines superior infrastructure performance with deep technical...  ...ll Do The Endpoint Engineering team at CoreWeave...  ...'s Teleport cluster: IAM node joining, RBAC...  .../OOB node setup and validation. Configure Okta SAML... 
    Performance
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours

    CoreWeave

    Bellevue, WA
    2 days ago
  • We Are:The Global AI Infrastructure team is at the center of...  ...computing, and high-performance workloads. We bring together...  ...and manage XPU-based clusters (GPU, DPU, LPU, CPU)...  ...with LLM inference engines (TensorRT-LLM),...  ...GPU benchmarking and validation tools (MLPerf, NCCL tests... 
    Performance
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    Seattle, WA
    2 days ago
  • $224k - $356.5k

     ...unlimited potential of AI to define the next...  ...own the platform — performance, CI/CD pipelines, validated recipes, and model bring-up infrastructure — that lets developers...  ...efficiency on edge cluster...  ...Computer Science, Computer Engineering, Electrical Engineering... 
    Performance
    Full time
    Local area

    Nvidia

    Seattle, WA
    7 hours ago
  •  ...your potential in a high-performance culture. It's knowing that...  ...THE TEAM The mission of the AI Enablement Department is...  ...strategic guidance and AI infrastructure. THE OPPORTUNITY Aritzia...  ...Aritzia’s AI evolution. As the Principal AI Platform Engineer, you will use your... 
    Performance
    Work at office

    Aritzia

    Seattle, WA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Infrastructure Engineer, AI Cluster Performance & Validation. Be the first to apply!