Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior System Architect, Infrastructure Reliability

$184k - $287.5k

NVIDIA

NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects.What you'll be doing:Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure.Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs.Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters.Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams.Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs.What we need to see:Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming.Experience building automated RCA (Root Cause Analysis) pipelines for HPC or cloud-scale environments.CPU Architecture Deep-Dive: Expert knowledge of x86/ARM node-level metrics: IPC (Instructions Per Cycle), cache contention, NUMA imbalance, and hardware interrupts.Programming Proficiency: Strong C++ and Python skills, with the ability to build high-performance daemons that monitor system health without impacting workload performance.Scale Experience: Familiarity with cluster resource managers (Slurm, LSF, or Kubernetes) and how they manage job lifecycle and signal propagation.Ways To Stand Out From The Crowd:Low-Level Diagnostics: Expert knowledge of the Linux kernel and its error-reporting interfaces (/dev/mcelog, dmesg, journald). Understand how the kernel handles hardware exceptions and memory faults.GPU Infrastructure Proficiency: Deep experience with the NVIDIA DCGM (Data Center GPU Manager) and NVIDIA Management Library (NVML) for monitoring device health and capturing state-dumps.Experience with tools doing non-intrusive monitoring of application health and syscall-level failure patterns.Experience with checkpoint/restore technologies (like CRIU) and their application in long-running EDA flows.#LI-Hybrid Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until June 19, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, MA, Westford; US, TX, Austin; US, NC, Durham; US, WA, RedmondType: Full time

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Senior System Architect, Infrastructure Reliability in Redmond, WA vacancy
  • $184k - $287.5k

    NVIDIA is looking for an experienced infrastructure Solutions Architect. Do you want to be part of a team that brings Artificial Intelligence (AI) hardware...  ..., NICs/HCAs, switches, and GPUsGood understanding of system hardware architecture impact on network performance,... 
    Senior
    Full time

    Nvidia

    Redmond, WA
    3 days ago
  • $184k - $287.5k

     ...and maintain our leadership. NVIDIA is seeking a motivated system architect to define future aspects of our GPU through employing pioneering...  ...center workloads.Develop and enhance architecture analysis infrastructure, including performance simulators, testbench components and... 
    Senior
    Full time
    Work experience placement
    Remote work
    Night shift

    Nvidia

    Redmond, WA
    4 days ago
  • $137.2k - $246.79k

     ...drones, and security of borders, critical infrastructure, and smart cities. The company...  ...Security, Smart Cities, Uncrewed Aircraft Systems (UAS), and Airspace Management...  ...Mobility (UTM).Echodyne is seeking a Senior Systems Architect, Radar Development to join our fast-growing... 
    Senior
    Full time
    Temporary work

    Echodyne

    Kirkland, WA
    3 days ago
  • $165k - $230k

     ...ultimate goal of enabling human life on Mars.SR. HARDWARE / INFRASTRUCTURE SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our...  ...deploy Starlink, the world’s most advanced broadband internet system. Starlink is the world’s largest satellite constellation... 
    Senior
    Permanent employment
    Temporary work
    Work at office
    Worldwide
    Monday to Friday
    Weekend work

    SpaceX

    Redmond, WA
    4 days ago
  • $184k - $287.5k

     ...Team means contributing to the infrastructure that powers our innovative...  ...implementing software and systems engineering practices to ensure...  ...of AI systems.As a senior DGX Cloud AI Infrastructure...  ...Define meaningful and actionable reliability metrics to track and improve... 
    Senior
    Full time
    Remote work

    Nvidia

    Redmond, WA
    4 days ago
  • $119.8k - $234.7k

     ...EngineeringDiscipline: Site Reliability EngineeringCompany: MicrosoftOverviewMicrosoft...  ...5 Copilot, providing shared infrastructure, identity, messaging,...  ...demanding workloads. As a Senior Site Reliability Engineer,...  .... Build Scalable Systems: Develop automation for monitoring... 
    Senior
    Ongoing contract
    Local area
    3 days per week

    Microsoft

    Redmond, WA
    15 hours ago
  •  ...or local law. Summary of Position:The Senior Infrastructure Engineer is Denali Advanced Integration...  ...Dell EMC storage, and Cohesity backup systems across the primary datacenter and...  ...infrastructure to ensure backups operate reliably, efficiently, and at capacity Monitor... 
    Senior
    Hourly pay
    Contract work
    Temporary work
    Work at office
    Local area

    Denali Advanced Integration

    Redmond, WA
    2 days ago
  • $184k - $287.5k

     ...as well as developing scalable AI infrastructure services globally. We are seeking...  ...Agentic AI in production.As a senior DGX Cloud AI Infrastructure software...  ...Define meaningful and actionable reliability metrics to track and improve system and service reliability.Skilled... 
    Senior
    Full time
    Remote work

    Nvidia

    Redmond, WA
    15 hours ago
  • $184k - $287.5k

     ...building the software and systems that power the world’...  ...We are looking for a Senior Software Engineer to...  ...run efficiently and reliably at scale. You will...  ...large-scale AI clusters, infrastructure, and end-to-end...  ...Proven track record of architecting, debugging, and scaling... 
    Senior
    Full time
    Remote work

    Nvidia

    Redmond, WA
    4 days ago
  • $184k - $287.5k

     ...clusters even more performant. As a networking Sr. Solutions Architect at NVIDIA you will have agency and palpable effects on the business...  ...such as Perl, python, and shell scripts)Knowledge in Cloud infrastructure and AI workflowsLinux Environment and Linux... 
    Senior
    Full time
    Remote work

    Nvidia

    Redmond, WA
    4 days ago
  • Company DescriptionUSM Business Systems Inc. is a quickly developing...  ...Consulting and IT Infrastructure. Our other offerings include...  ...a project-driven firm that reliably meets the IT needs of our State...  ...EngineeringExperience level: Mid-Senior LevelIndustry: Information... 
    Senior
    Worldwide

    USM Systems

    Bellevue, WA
    3 days ago
  • $152k - $241.5k

     ...the world’s largest customers? NVIDIA is looking for an Infrastructure Solutions Architect to lead deployment and bring‑up of our next‑generation Data...  ...and performance data, identifying product health trends, system bottlenecks, and operational risks.Solve challenging... 
    Full time
    Remote work
    Worldwide

    Nvidia

    Redmond, WA
    4 days ago
  •  ...DescriptionHi,Hope you are doing great!!We have an urgent opening for senior reliability engineer and the job description is as follows :Location:...  ...assimilate/ review / define engineering specifications from system data and carryout reliability qualitative / quantitative... 
    Senior

    EROS Technologies

    Redmond, WA
    4 days ago
  • $153.6k - $207.8k

     ...standard for how builders architect AI workloads that are secure, reliable, and efficient on AWS —...  ...? We are looking for a Senior Security Solutions...  ...security decisions in AI systems.- Raise the Bar Across...  ..., systems engineering, infrastructure, security, networking,... 
    Senior
    Local area
    Flexible hours

    AmazonWebServices

    Bellevue, WA
    2 days ago
  • $119.8k - $234.7k

     ...components of the Windows operating system. We own critical platform...  ...networking, boot, and system infrastructure. Teams across Microsoft and...  ...secure, performant, and reliable Windows experiences to millions...  ...of customers worldwide.As a Senior and Principal Product Manager... 
    Senior
    Ongoing contract
    Local area
    Worldwide
    3 days per week

    Microsoft

    Redmond, WA
    15 hours ago
  • $272k - $431.25k

     ...strength is to innovate how we architect and develop our GPU for the...  ...the scalability of our design, infrastructure and methodology. We are looking for a Principal System Architect with a wealth of experience...  ...communication skills. This senior technical position entails... 
    Full time
    Remote work

    Nvidia

    Redmond, WA
    4 days ago
  • $142.8k - $274.8k

     ...EngineeringDiscipline: Site Reliability EngineeringCompany: MicrosoftOverviewMicrosoft...  ...5 Copilot, providing shared infrastructure, identity, messaging,...  ...practices that make systems reliable by design.This role...  ...through others—developing senior and principal engineers,... 
    Ongoing contract
    Temporary work
    Fixed term contract
    Local area
    Immediate start
    3 days per week

    Microsoft

    Redmond, WA
    15 hours ago
  • $174k - $253k

     ...troubleshooting large-scale distributed systems. 2 years of experience leading...  ...Science or Engineering. About The Job Site Reliability Engineering (SRE) is what you get when...  ...the architecture built by the Technical Infrastructure team to keep it running. From developing... 
    Senior
    Temporary work

    Google

    Kirkland, WA
    4 days ago
  • $192k - $278k

    Senior Product Manager, AI Infrastructure Advanced Experience owning outcomes and decision making, solving ambiguous...  ...TPU supercomputers as effortless, reliable, and accessible to operate. We own...  ..., Data Center operations, systems research, and much more. Individual... 
    Senior
    Temporary work
    Worldwide

    Google Inc.

    Kirkland, WA
    1 day ago
  • $165.6k - $296.4k

     ...in the industry. Cloud Infrastructure Automation and...  ...datacenter capacity, and safe, reliable, and efficient...  ....We are looking for a Senior Principal Technical Program...  ...highly complex set of systems and processes, and partner...  ..., and technical architects to design scalable, future... 
    Senior
    Ongoing contract
    Local area
    3 days per week

    Microsoft

    Redmond, WA
    15 hours ago
  • $184k - $287.5k

     ...looking for someone who wants to build AI speed infrastructure for Tegra: a ridiculously fast build, test, and validation system for the future, aimed at supporting...  ..., and scarce device-lab resources. You will architect how risk-based gating operates at agent-speed... 
    Senior
    Full time
    Local area
    Remote work

    Nvidia

    Redmond, WA
    4 days ago
  • $119.8k - $234.7k

     ...evaluation backbone for safe, reliable, and efficient agentic...  ...Security. This team will create systems that determine when an AI agent...  ...burden. Role mission As a Senior Security Researcher and...  ...platform implementation, test infrastructure, telemetry, measurement, security... 
    Senior
    Ongoing contract
    Local area
    Shift work
    3 days per week

    Microsoft

    Redmond, WA
    4 days ago
  • $140k - $200k

    Water / Wastewater Infrastructure Senior Project EngineerWater and Environment Jobs with David Evans...  ...team excels in managing complex whole-system resources. We offer comprehensive...  ...our clients by blending innovation with reliability and sustainable practices. This approach... 
    Senior
    Work at office
    Local area
    Flexible hours

    David Evans and Associates

    Bellevue, WA
    2 days ago
  • $119.8k - $234.7k

     ...of the highest-scale experimentation platforms - infrastructure that enables rapid iteration in AI systems and product features. You will design and build services...  ...your expertise in distributed systems, service reliability, and experimentation methodologies. You will... 
    Senior
    Ongoing contract
    Local area
    3 days per week

    Microsoft

    Redmond, WA
    4 days ago
  • $137.3k - $185.7k

     ...organizations operating in places without reliable connectivity.As an Wireless Test Architect you will engage with an...  ...to develop new, advanced test systems for our Digital Phased Array systems...  ...them with our existing test infrastructure. We have near field spherical scanners... 
    Senior
    Permanent employment
    Flexible hours

    Amazon

    Redmond, WA
    4 days ago
  • $162.1k - $187k

     ...ultrasound and GlideScope video laryngoscopy & bronchoscopy systems effectively address unmet needs for healthcare providers...  ...information, please visit . Overview Verathon is seeking a Senior System Architect to join the Information Technology organization. This plays... 
    Senior
    Full time

    Verathon

    Bothell, WA
    4 days ago
  • $137.3k - $185.7k

     ...consumers, businesses, government agencies, and other organizations operating in places without reliable connectivity.We are building the next generation of optical interconnect systems, and reliability is at the heart of what we do. We are seeking a highly motivated and... 
    Senior
    Permanent employment
    Flexible hours

    Amazon

    Redmond, WA
    15 hours ago
  • $119.8k - $234.7k

     ...Profession: Software EngineeringDiscipline: Site Reliability EngineeringCompany: MicrosoftOverviewAre...  ...engineering team. We are looking for a Senior Site Reliability Engineer who will be...  ...teams to ensure services and systems are highly stable, meet performance SLAs... 
    Senior
    Ongoing contract
    Local area
    3 days per week

    Microsoft

    Redmond, WA
    15 hours ago
  • $159.2k - $215.3k

     ...network. Our mission is to deliver fast, reliable internet connectivity to customers beyond...  ...are seeking a highly skilled and motivated Senior Reliability Engineer to join the Hardware...  ...and performance of satellite hardware systems in the harsh environment of space. You will... 
    Senior
    Permanent employment
    Flexible hours

    Amazon

    Redmond, WA
    3 days ago
  • $97.6k - $188.4k

     ...drive their business more successfully. The Microsoft Cloud Infrastructure Finance team is an exciting and fast-evolving finance team at...  ...the future of Microsoft cloud products. This is a role for a systems-oriented finance professional who can operate across Finance,... 
    Senior
    Ongoing contract
    Work experience placement
    Local area
    3 days per week

    Microsoft

    Redmond, WA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior System Architect, Infrastructure Reliability. Be the first to apply!