Senior System Architect, Infrastructure Reliability
$184k - $287.5kNVIDIA
Senior System Architect: Heterogeneous EDA Systems
NVIDIA is seeking a Senior System Architect to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects.
What you'll be doing:
- Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure.
- Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs.
- Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters.
- Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams.
- Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs.
What we need to see:
- Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming.
- Experience building automated RCA (Root Cause Analysis) pipelines for HPC or cloud-scale environments.
- CPU Architecture Deep-Dive: Expert knowledge of x86/ARM node-level metrics: IPC (Instructions Per Cycle), cache contention, NUMA imbalance, and hardware interrupts.
- Programming Proficiency: Strong C++ and Python skills, with the ability to build high-performance daemons that monitor system health without impacting workload performance.
- Scale Experience: Familiarity with cluster resource managers (Slurm, LSF, or Kubernetes) and how they manage job lifecycle and signal propagation.
Ways To Stand Out From The Crowd:
- Low-Level Diagnostics: Expert knowledge of the Linux kernel and its error-reporting interfaces (/dev/mcelog, dmesg, journald). Understand how the kernel handles hardware exceptions and memory faults.
- GPU Infrastructure Proficiency: Deep experience with the NVIDIA DCGM (Data Center GPU Manager) and NVIDIA Management Library (NVML) for monitoring device health and capturing state-dumps.
- Experience with tools doing non-intrusive monitoring of application health and syscall-level failure patterns.
- Experience with checkpoint/restore technologies (like CRIU) and their application in long-running EDA flows.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until June 19, 2026. This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
$184k - $287.5k
NVIDIA is looking for an experienced infrastructure Solutions Architect. Do you want to be part of a team that brings Artificial Intelligence (AI) hardware... ..., NICs/HCAs, switches, and GPUsGood understanding of system hardware architecture impact on network performance,...SeniorFull time- ...Senior Infrastructure EngineerThe Senior Infrastructure Engineer is Denali Advanced Integration's... ...Dell EMC storage, and Cohesity backup systems across the primary datacenter and secondary... ...to ensure backups operate reliably, efficiently, and at capacityMonitor ServiceNow...SeniorContract workWork at office
- ...Position: Senior Site Reliability Engineer (SRE) Location: Redmond WA... ...scalable Azure platforms. Architect production-grade cloud... ...). Create end-to-end system architecture, deployment... ...excellence. Develop Infrastructure as Code (IaC) using Bicep...SeniorFull time
$240k - $333k
...Senior Leadership Technical Program Manager, Customer Experience, Global Infrastructure In accordance with Washington state law, we are highlighting our comprehensive benefits... ...governance across GGI to deliver scalable, reliable infrastructure solutions. Google Cloud...SeniorTemporary workWork at office- DXC Technology seeks an experienced AWS Infrastructure Administrator to manage, secure, and optimize enterprise AWS infrastructure, ensuring high availability and reliability across production environments. The role involves hands-on administration of EC2, S3, RDS, IAM...Senior
$137.2k - $246.79k
...drones, and security of borders, critical infrastructure, and smart cities. The company... ...Security, Smart Cities, Uncrewed Aircraft Systems (UAS), and Airspace Management... ...Mobility (UTM).Echodyne is seeking a Senior Systems Architect, Radar Development to join our fast-growing...SeniorFull timeTemporary work$232k - $319k
...potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era.... ...help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and...SeniorPermanent employmentLocal areaWorldwideFlexible hours- Google in Kirkland, WA seeks software engineers to build infrastructure at scale, across distributed systems, networking, and data storage. You’ll tackle projects critical to Google’s needs with opportunities to switch teams as you grow. We offer comprehensive Washington...Senior
$165k - $230k
...ultimate goal of enabling human life on Mars.SR. HARDWARE / INFRASTRUCTURE SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our... ...deploy Starlink, the world’s most advanced broadband internet system. Starlink is the world’s largest satellite constellation...SeniorPermanent employmentTemporary workWork at officeWorldwideMonday to FridayWeekend work- ...Enterprise Development LP is seeking a Principal Presales Systems Engineer to join our remote networking team. You... ...with customers, account teams, and engineering to architect data center, WAN, cloud and AI infrastructure solutions that align with business goals. In...Remote job
$92k - $138k
...decision-making across the company. Our systems operate at scale across batch and... ...Learning Engineer to join our Offline Infrastructure team. This is an ideal role for a recent... ...systems that ensure our ML pipelines are reliable, scalable, and efficient. This role...Work at officeWorldwideRelocation package- ...available, performant, scalable, and secure infrastructure for security SaaS offerings. Build and operate tooling for Snowflake data systems. Develop automation for... ...a strong interest in applying AI/ML to reliability, security, or operational efficiency is...SeniorFull timeWork at officeFlexible hours
$147k - $202.4k
...AI by building the trusted, neutral infrastructure that enables organizations to safely... ...company is seeking a highly skilled Senior Site Reliability Engineer to join our team. We are a SaaS... ...specializing in securing large-scale systems. This role is a blend of software engineering...SeniorWork at officeLocal areaWorldwideFlexible hoursShift work$160k - $220k
...AI by building the trusted, neutral infrastructure that enables organizations to safely... ...mission. If you are too, let's talk. Senior Database Reliability Engineer (DBRE) Experience Level:... ...our large-scale, mission-critical systems. You will work closely with SRE,...SeniorPermanent employmentWork at officeLocal areaWorldwideFlexible hours$153k - $192k
...Senior AI/ML Solutions Architect Hybrid (Office 3 days/wk – Onsite-Flex) within Oregon, Washington,... ...economically sustainable health care system. Who We Are Looking For:... ...domains: can explain AI system needs to infrastructure teams and enterprise constraints to...SeniorWork at officeFlexible hours$112k - $218.4k
...: Our team is looking for a Senior Active Directory Site Reliability Engineer. Our mission is to improve the availability, latency, performance and security of the Identity systems behind Microsoft's cloud. Like traditional operations, we keep important revenue-critical...SeniorFull timeLocal area$200.4k - $260.5k
Senior Machine Learning Engineer, Data InfrastructureUnity Vector... ...across the company.Our systems operate at scale across batch... ...ensure our ML pipelines remain reliable, scalable, and... ...focuses on building reliable infrastructure for generating data infrastructure...SeniorFull timeWork at officeWorldwideRelocation package- Salesforce is seeking a Senior Director of Product Management to lead the vision, roadmap, and delivery of our core platform infrastructure, including the HyperForce runtime platform for AI-first, multi-substrate foundation. You will guide a 3-5 PM team, bridge internal...Senior
$184k - $287.5k
...Team means contributing to the infrastructure that powers our innovative... ...implementing software and systems engineering practices to ensure... ...of AI systems.As a senior DGX Cloud AI Infrastructure... ...Define meaningful and actionable reliability metrics to track and improve...SeniorFull timeRemote work$119.8k - $234.7k
...EngineeringDiscipline: Site Reliability EngineeringCompany: MicrosoftOverviewMicrosoft... ...5 Copilot, providing shared infrastructure, identity, messaging,... ...demanding workloads. As a Senior Site Reliability Engineer,... .... Build Scalable Systems: Develop automation for monitoring...SeniorOngoing contractLocal area3 days per week$160k - $210k
...Now, we're growing! We are looking for a Senior Site Reliability Engineer to strengthen our AWS infrastructure and improve service management across Cognitiv.... ...Qualifications AWS certifications (e.g., Solutions Architect, SysOps Administrator) Experience with...SeniorWork at officeImmediate startRemote workWork from home$232.3k - $406.6k
...'ll get to work on complex distributed systems and solve massive scale problems centered... ...roadmap, ensuring scalability, reliability, and alignment with industry best practices... ...distributed development, product, operations, infrastructure, cybersecurity, and support teams to...SeniorWork experience placementWork at officeLocal areaImmediate start$184k - $287.5k
...building the software and systems that power the world’... ...We are looking for a Senior Software Engineer to... ...run efficiently and reliably at scale. You will... ...large-scale AI clusters, infrastructure, and end-to-end... ...Proven track record of architecting, debugging, and scaling...SeniorFull timeRemote work$182k - $242k
...enterprises, CoreWeave combines superior infrastructure performance with deep technical... ...This is a critical step to make agents reliable enough to perform long tasks autonomously... ...acceleration, and large scale model training systems Experience leading technically...SeniorPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours- Salesforce is seeking an experienced backend engineer to design, implement, and operate large-scale cloud infrastructure. You will build and maintain distributed systems in public cloud environments, emphasizing availability, security, and performance. Expect to work with...Senior
- ...Senior Go Software EngineerWe are seeking a highly skilled Senior Go Software Engineer... ....Maintain disciplined handoffs and reliable progress reporting against assigned priorities... ....Platform engineering, developer infrastructure, build engineering, or release engineering...SeniorLocal area
$153.6k - $207.8k
...experienced and motivated Senior Solutions Architect to partner with our Telecommunications... ...needs, and design reliable, cost-effective, and... ...broadly competent across infrastructure, security, DevOps, databases... ...large-scale, distributed systems spanning on-premises, hybrid...SeniorFlexible hours- Company DescriptionUSM Business Systems Inc. is a quickly developing... ...Consulting and IT Infrastructure. Our other offerings include... ...a project-driven firm that reliably meets the IT needs of our State... ...EngineeringExperience level: Mid-Senior LevelIndustry: Information...SeniorWorldwide
$122.3k - $158.5k
...driven tools, live services, and infrastructure that ensure global scale,... ...risk across partners and systems, and ensure compliance with... ...operate securely at scale. The Senior Machine Learning Engineer... ...consistency, data quality, and reliable offline and online feature...Full timeWork at officeLocal areaRemote work$152k - $241.5k
...the world’s largest customers? NVIDIA is looking for an Infrastructure Solutions Architect to lead deployment and bring‑up of our next‑generation Data... ...and performance data, identifying product health trends, system bottlenecks, and operational risks.Solve challenging...Full timeRemote workWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior System Architect, Infrastructure Reliability. Be the first to apply!
- system architect Redmond, WA
- pega system architect Redmond, WA
- technical architect Redmond, WA
- senior developer Redmond, WA
- senior aws cloud engineer Redmond, WA
- remote senior salesforce administrator Redmond, WA
- senior marketing operations manager Redmond, WA
- senior manager tax Redmond, WA
- senior tax Redmond, WA
- senior helpdesk technician Redmond, WA


