Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior System Architect, Infrastructure Reliability

$184k - $287.5k

NVIDIA

NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects.

What you'll be doing:
  • Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure. Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs.
  • Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters.
  • Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams.
  • Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs.
What we need to see:
  • Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming. Experience building automated RCA (Root Cause Analysis) pipelines for HPC or cloud-scale environments.
  • CPU Architecture Deep-Dive: Expert knowledge of x86/ARM node-level metrics: IPC (Instructions Per Cycle), cache contention, NUMA imbalance, and hardware interrupts.
  • Programming Proficiency: Strong C++ and Python skills, with the ability to build high-performance daemons that monitor system health without impacting workload performance.
  • Scale Experience: Familiarity with cluster resource managers (Slurm, LSF, or Kubernetes) and how they manage job lifecycle and signal propagation.
Ways To Stand Out From The Crowd:
  • Low-Level Diagnostics: Expert knowledge of the Linux kernel and its error-reporting interfaces (/dev/mcelog, dmesg, journald). Understand how the kernel handles hardware exceptions and memory faults.
  • GPU Infrastructure Proficiency: Deep experience with the NVIDIA DCGM (Data Center GPU Manager) and NVIDIA Management Library (NVML) for monitoring device health and capturing state-dumps. Experience with tools doing non-intrusive monitoring of application health and syscall-level failure patterns. Experience with checkpoint/restore technologies (like CRIU) and their application in long-running EDA flows. #LI-Hybrid

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 2, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

NVIDIA pioneered accelerated computing. Today, our AI infrastructure powers global intelligence, transforming every industry. Learn more about NVIDIA.

#J-18808-Ljbffr
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior System Architect, Infrastructure Reliability in Redmond, WA vacancy
  • $184k - $287.5k

    NVIDIA is looking for an experienced infrastructure Solutions Architect. Do you want to be part of a team that brings Artificial Intelligence (AI) hardware...  ..., NICs/HCAs, switches, and GPUsGood understanding of system hardware architecture impact on network performance,... 
    Senior
    Full time

    Nvidia

    Redmond, WA
    4 days ago
  •  ...A leading data and AI infrastructure provider is seeking a Senior Staff Technical Program Manager for Reliability to enhance the reliability and performance of their multi-cloud...  ...in cloud infrastructure and distributed systems management. You will shape strategies and... 
    Senior

    Menlo Ventures

    Bellevue, WA
    19 hours ago
  •  ...law.  Summary of Position: The Senior Infrastructure Engineer is Denali Advanced Integration...  ...Dell EMC storage, and Cohesity backup systems across the primary datacenter and secondary...  ...to ensure backups operate reliably, efficiently, and at capacity  Monitor... 
    Senior
    Hourly pay
    Contract work
    Temporary work
    Work experience placement
    Work at office
    Local area

    3MD Inc.

    Redmond, WA
    29 days ago
  • $119.8k - $234.7k

     ...Overview The Infrastructure team in Finetuning, Inference and Training...  ...FIT) group is looking for a Senior Software Engineer who loves...  ...the speed, security, and reliability of our deployments, secure our...  ...throughput services and the systems that deploy, observe, and operate... 
    Senior
    Ongoing contract
    Local area

    Microsoft Corporation

    Redmond, WA
    3 days ago
  • $137.2k - $246.79k

     ...drones, and security of borders, critical infrastructure, and smart cities. The company...  ...Security, Smart Cities, Uncrewed Aircraft Systems (UAS), and Airspace Management...  ...Mobility (UTM).Echodyne is seeking a Senior Systems Architect, Radar Development to join our fast-growing... 
    Senior
    Full time
    Temporary work

    Echodyne

    Kirkland, WA
    19 hours ago
  • $232k - $319k

     ...potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era....  ...help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and... 
    Senior
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    3 days ago
  • $165k - $230k

     ...ultimate goal of enabling human life on Mars.SR. HARDWARE / INFRASTRUCTURE SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our...  ...deploy Starlink, the world’s most advanced broadband internet system. Starlink is the world’s largest satellite constellation... 
    Senior
    Permanent employment
    Temporary work
    Work at office
    Worldwide
    Monday to Friday
    Weekend work

    SpaceX

    Redmond, WA
    19 hours ago
  • $165k - $270k

     ...human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD) At SpaceX...  ...operate all parts of the system - receivers that allow users...  ...Starshield's software and GPU infrastructure, you will design, operate...  ...train junior engineersAs a senior engineer you must lead the team... 
    Senior
    Permanent employment
    Temporary work
    Work at office
    Immediate start
    Monday to Friday
    Weekend work

    SpaceX

    Redmond, WA
    2 days ago
  • $116.9k - $203.6k

     ...responsible for enabling the hardware infrastructure underlying this growth...  ...time to support scalable, reliable cloud growth. Through data-...  ...beyond. We are seeking a Senior Technical Solution Manager....  ...Science, Management Information Systems, Engineering, or related... 
    Senior
    Ongoing contract
    Work at office
    Local area
    Worldwide
    3 days per week

    Microsoft

    Redmond, WA
    16 hours ago
  • $142.8k - $274.8k

     ...culture every day. Join the Systems Pathfinding and Architecture...  ...Azure Hardware Systems and Infrastructure (AHSI) organization, the team...  ...passionate Principal AI System Architect to join the team. In this...  ..., network topology design), reliability and serviceability, datacenter... 
    Ongoing contract
    Work at office
    Local area
    Worldwide
    3 days per week

    Microsoft

    Redmond, WA
    2 days ago
  • Senior Machine Learning Engineer, Data InfrastructureUnity Vector...  ...across the company.Our systems operate at scale across batch...  ...ensure our ML pipelines remain reliable, scalable, and...  ...focuses on building reliable infrastructure for generating data infrastructure... 
    Senior
    Full time
    Work at office
    Worldwide
    Relocation package

    Unity Technologies

    Bellevue, WA
    1 day ago
  • $160k - $220k

     ...AI by building the trusted, neutral infrastructure that enables organizations to safely...  ...mission. If you are too, let's talk. Senior Database Reliability Engineer (DBRE) Experience Level:...  ...our large-scale, mission-critical systems. You will work closely with SRE,... 
    Senior
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta, Inc.

    Bellevue, WA
    1 day ago
  •  ...Job Full Description We are seeking an Embedded Systems Architect who is equally comfortable thinking at the architectural level and...  ...is about more than technology-it is about creating reliable products that solve real customer problems. • Entrepreneurial... 
    Day shift

    Express Employment Professionals Defunct

    Redmond, WA
    2 days ago
  • $119.8k - $234.7k

     ...EngineeringDiscipline: Site Reliability EngineeringCompany: MicrosoftOverviewMicrosoft...  ...5 Copilot, providing shared infrastructure, identity, messaging,...  ...demanding workloads. As a Senior Site Reliability Engineer,...  .... Build Scalable Systems: Develop automation for monitoring... 
    Senior
    Ongoing contract
    Local area
    3 days per week

    Microsoft

    Redmond, WA
    1 day ago
  • $184k - $287.5k

     ...as well as developing scalable AI infrastructure services globally. We are seeking...  ...Agentic AI in production.As a senior DGX Cloud AI Infrastructure software...  ...Define meaningful and actionable reliability metrics to track and improve system and service reliability.Skilled... 
    Senior
    Full time
    Remote work

    Nvidia

    Redmond, WA
    11 hours ago
  • $184k - $287.5k

     ...building the software and systems that power the world’...  ...We are looking for a Senior Software Engineer to...  ...run efficiently and reliably at scale. You will...  ...large-scale AI clusters, infrastructure, and end-to-end...  ...Proven track record of architecting, debugging, and scaling... 
    Senior
    Full time
    Remote work

    Nvidia

    Redmond, WA
    19 hours ago
  • $153k - $204k

     ...CoreWeave combines superior infrastructure performance with deep technical...  ...ship software quickly, reliably, and safely. We own the deployment...  ...and monitoring, and release systems — built on a Go, Kubernetes,...  ...Role We are seeking a Senior Software Engineer to design,... 
    Senior
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Bellevue, WA
    15 days ago
  • $188k - $275k

     ...CoreWeave combines superior infrastructure performance with deep technical...  ...at scale has a seamless, reliable, and high-performance experience...  ...data centers, hardware systems, and customer workloads to maintain...  ...experience as a Solutions Architect, Field Engineer,... 
    Senior
    Permanent employment
    Full time
    Contract work
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Bellevue, WA
    9 days ago
  • $185k - $235k

     ...model outputs meet the freshness, latency, reliability, and throughput requirements of real-...  ...data quality, labeling, and validation systems and identify gaps, bias, and freshness...  .... Collaborate with Ads, Data, and Infrastructure teams on signal and feature design.... 
    Senior
    Full time

    NewsBreak

    Bellevue, WA
    2 days ago
  • $160k - $210k

     ...Now, we're growing! We are looking for a Senior Site Reliability Engineer to strengthen our AWS infrastructure and improve service management across Cognitiv....  ...Qualifications AWS certifications (e.g., Solutions Architect, SysOps Administrator)  Experience with... 
    Senior
    Work at office
    Immediate start
    Remote work
    Work from home

    Cognitiv

    Bellevue, WA
    15 days ago
  • $182k - $242k

     ...enterprises, CoreWeave combines superior infrastructure performance with deep technical...  ...This Is a Critical Step To Make Agents Reliable Enough To Perform Long Tasks Autonomously...  ...acceleration, and large scale model training systems Experience leading technically... 
    Senior
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Bellevue, WA
    19 hours ago
  • $185k - $235k

     ...advanced AI, recommendation systems, and adtech. Recognized by...  ...our mission: building the infrastructure layer for content intelligence...  ...looking for an experienced Senior Machine Learning Engineer...  ...the freshness, latency, and reliability that real-time bidding demands... 
    Senior
    Full time
    Local area
    Work from home

    NewsBreak

    Bellevue, WA
    2 days ago
  • $153.6k - $207.8k

     ...experienced and motivated Senior Solutions Architect to partner with our Telecommunications...  ...needs, and design reliable, cost-effective, and...  ...broadly competent across infrastructure, security, DevOps, databases...  ...large-scale, distributed systems spanning on-premises, hybrid... 
    Senior
    Flexible hours

    AmazonWebServices

    Bellevue, WA
    2 days ago
  •  ...Role: Senior GO Software Engineer Location: Redmond WA (Onsite) Job Description...  ...Maintain disciplined handoffs and reliable progress reporting against assigned priorities...  .... Platform engineering, developer infrastructure, build engineering, or release... 
    Senior
    Local area

    United IT Solutions

    Kirkland, WA
    3 days ago
  • Company DescriptionUSM Business Systems Inc. is a quickly developing...  ...Consulting and IT Infrastructure. Our other offerings include...  ...a project-driven firm that reliably meets the IT needs of our State...  ...EngineeringExperience level: Mid-Senior LevelIndustry: Information... 
    Senior
    Worldwide

    USM Systems

    Bellevue, WA
    4 days ago
  • $166k - $244k

    # Senior Software Engineer, Infrastructure, Google Cloud Compute InfrastructureGoogle • onsite • 451 7th Ave S,...  ...distributed computing, large-scale system design, networking and data storage...  ...at unparalleled scale, efficiency, reliability and velocity. Our customers include... 
    Senior
    Worldwide

    Epic Games

    Kirkland, WA
    1 day ago
  • $119.8k - $234.7k

     ...delivering core datacenter infrastructure for Microsoft’s cloud business...  ...a passionate, high-energy Architect to help build the cloud datacenters...  ...Engineering (DCE), Senior Architects provide technical...  ...financial impact, quality, reliability, and time duration.Review design... 
    Senior
    Ongoing contract
    For contractors
    Work at office
    Local area
    Remote work
    Worldwide

    Microsoft

    Redmond, WA
    11 hours ago
  • $152k - $241.5k

     ...the world’s largest customers? NVIDIA is looking for an Infrastructure Solutions Architect to lead deployment and bring‑up of our next‑generation Data...  ...and performance data, identifying product health trends, system bottlenecks, and operational risks.Solve challenging... 
    Full time
    Remote work
    Worldwide

    Nvidia

    Redmond, WA
    19 hours ago
  •  ...DescriptionHi,Hope you are doing great!!We have an urgent opening for senior reliability engineer and the job description is as follows :Location:...  ...assimilate/ review / define engineering specifications from system data and carryout reliability qualitative / quantitative... 
    Senior

    EROS Technologies

    Redmond, WA
    19 hours ago
  • $119.8k - $234.7k

     ...components of the Windows operating system. We own critical platform...  ...networking, boot, and system infrastructure. Teams across Microsoft and...  ...secure, performant, and reliable Windows experiences to millions...  ...of customers worldwide.As a Senior and Principal Product Manager... 
    Senior
    Ongoing contract
    Local area
    Worldwide
    3 days per week

    Microsoft

    Redmond, WA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior System Architect, Infrastructure Reliability. Be the first to apply!