Senior System Architect, Infrastructure Reliability
$184k - $356.5kNVIDIA
NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects.
What you'll be doing:
Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure.
Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs.
Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters.
Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams.
Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs.
What we need to see:
Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming.
Experience building automated RCA (Root Cause Analysis) pipelines for HPC or cloud-scale environments.
CPU Architecture Deep-Dive: Expert knowledge of x86/ARM node-level metrics: IPC (Instructions Per Cycle), cache contention, NUMA imbalance, and hardware interrupts.
Programming Proficiency: Strong C++ and Python skills, with the ability to build high-performance daemons that monitor system health without impacting workload performance.
Scale Experience: Familiarity with cluster resource managers (Slurm, LSF, or Kubernetes) and how they manage job lifecycle and signal propagation.
Ways To Stand Out From The Crowd:
Low-Level Diagnostics: Expert knowledge of the Linux kernel and its error-reporting interfaces (/dev/mcelog, dmesg, journald). Understand how the kernel handles hardware exceptions and memory faults.
GPU Infrastructure Proficiency: Deep experience with the NVIDIA DCGM (Data Center GPU Manager) and NVIDIA Management Library (NVML) for monitoring device health and capturing state-dumps.
Experience with tools doing non-intrusive monitoring of application health and syscall-level failure patterns.
Experience with checkpoint/restore technologies (like CRIU) and their application in long-running EDA flows.
#LI-Hybrid
You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until September 29, 2026.This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.$164.2k - $295.6k
...highly skilled and motivated Systems / Product Infrastructure Engineer to join our Data... ...and dependable.Job Title: Senior Principal Systems... ...energy efficiency and system reliability.Join forces with multiple... ...related incidents.Design and architect end-to-end data center infrastructure...SeniorFull timeTemporary workWork at officeLocal areaRemote workWorldwide$308.6k - $417.5k
Position We are looking for an experienced system performance architect to drive the definition and exploration of performance to craft the future of compute infrastructure performance across data center and networking domains. Responsibilities Define SoC system and subsystem...Senior- NVIDIA is hiring a Senior Software Architect for the Agent Harness & Runtime Engineering team to build foundational AI systems. You’ll work across agent runtimes, inference, evaluation... ...essential, with experience in GPU infrastructure and cloud environments. #J-18808-...Senior
- ...are rebuilding logistics infrastructure for the national security... ...allied partners. We build AI systems that determine how... ...We are looking for a Senior Solutions Architect to lead the design, deployment... ...automate workflows, and improve reliability and scalability of deployed...SeniorWork at officeLocal area
$184k - $287.5k
NVIDIA is looking for an experienced infrastructure Solutions Architect. Do you want to be part of a team that brings Artificial Intelligence (AI) hardware... ..., NICs/HCAs, switches, and GPUsGood understanding of system hardware architecture impact on network performance,...SeniorFull time- ...has to match. The role We\'re looking for a Senior SRE to own the reliability, scalability, and operational posture of Satsuma\'s multi-cloud infrastructure. You\'ll be the person who keeps things running, builds the systems that prevent fires, and makes on-call not terrible...Senior
- CVS Health is seeking a Senior Mainframe DB2 Database Administrator (DBA) to support, optimize... ...support across large-scale mainframe systems. You will collaborate with development and infrastructure teams to improve reliability and availability. The ideal candidate brings...Senior
- ...America’s next-generation power company building the backend data infrastructure for telemetry, market signals, and operations. As a Data... ...software, hardware, markets, and operations teams to deliver reliable, scalable data solutions for dashboards, models, and APIs in...Senior
$128.7k - $261.3k
...breakthrough hardware and battery systems to intuitive design,... ...where we build the critical infrastructure that powers every machine learning... ...to use, and exceptionally reliable. Your success will be... ...advanced driverless vehicles.As a Senior ML Infra Engineer, you will...SeniorFull timeWork at officeLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours- ...Job title: Senior Cloud Solutions Architect Location: Austin, TX Duration: Long Term Job Description... ...\'\'\'\'s cloud computing infrastructure and applications. Implements and... ...ensuring security, compliance, and system reliability. Required Skill: 8 Years of Multi...Senior
- ...served immigrants from 150+ countries! As a Senior Cloud Engineer, you will: - Lead infrastructure engineering projects to improve banking... ...infrastructure issues faster and increase the reliability of Stilt's systems. - Build close partnerships with other tech...Senior
- Visa is seeking a Sr. ML Engineer to design, build, and maintain scalable ML platform infrastructure powering AI/ML applications. You will guide the team on deployment patterns, create tooling to productionize models, and establish best practices across the ML platform...Senior
$199.75k - $258.5k
Senior Principal Electrical Systems ArchitectFrom applied research to advanced engineering, the Engineering... ...Principal Electrical Systems Architect on our Engineering Technologist team... ...productization, and deliver high-performance, reliable, and cost-effective solutions...Senior$199.75k - $258.5k
Senior Principal Mechanical Systems ArchitectFrom applied research to advanced engineering, the Engineering... ...Principal Mechanical Systems Architect on our Engineering Technologist team... ...productization, and deliver high-performance, reliable, and cost-effective solutions...Senior- ...in execution and impactThe RoleAs a Senior Site Reliability Engineer at Proofpoint you will develop... ...team player who cares about the infrastructure, remains calm in crisis,collaborates... ...experience in troubleshooting, and tuning in systems, networking, and cloudservices.• A...SeniorFull timeFlexible hours
$155.4k - $210.2k
...growing business unit within Amazon.com which provides cloud services. Since early 2006, AWS has provided a highly reliable, scalable, low-cost infrastructure platform that powers hundreds of thousands of businesses in 190 countries around the world. We are looking for a...SeniorLocal areaFlexible hours- ...join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a... ...available, secure, and scalable Azure cloud infrastructure supporting TeamViewer’s global SaaS... ...monitoring, alerting, and automation systems to guarantee 24/7 service reliability...SeniorTemporary workCasual workWorldwide
- ...Job Description Job Description Senior Solutions Architect – Cloud & AI Optimization Location... ..., analytics, and decision-support systems Build insights that drive cost optimization... ...workflows Platform & Infrastructure (Preferred) Kubernetes Cloud-native...SeniorRemote work
- NVIDIA in Austin, TX, is hiring a Senior Systems Engineer for Agentic AI (Finance). You will architect scalable systems across agent runtimes, inference, evaluation, and orchestration, delivering production-ready software. The position requires 12+ years, a BS/MS in CS...Senior
$126.2k - $264.1k
...solve complex problems in distributed systems, networking, multi-tenant Infrastructure-as-a-Service (IaaS), and Software... .... We are looking for an Architect who will contribute to and direct... ...leadership skills and can influence senior leadership in a positive way to make...SeniorTemporary workFlexible hours$136.2k - $214.01k
...execution and impact The Role As a Senior Site Reliability Engineer at Proofpoint you will... ...enthusiastic team player who cares about the infrastructure, remains calm in crisis,... ...experience in troubleshooting, and tuning in systems, networking, and cloud services....SeniorFull timeFlexible hours- ...Job Description Bicycle Health - Senior Infrastructure Software Engineer SUMMARY Company... ...Acceptable Tech Background: Distributed Systems, Docker, Kubernetes, Node, TypeScript,... ...for empowering teams to build highly reliable and performant services that support patients...SeniorWork at officeRemote workFlexible hours
- ...AI and Bitcoin mining infrastructure. Bitdeer is committed... ...We are seeking a Senior AI Storage Infrastructure... ...be responsible for architecting high-performance storage... ...to parallel file system integration—that enable... ...maintain high standards of reliability and performance...SeniorLocal area
$186.07k - $218.9k
...at Coinbase. Coinbase’s Developer Infrastructure group builds the systems every Coinbase engineer relies on to... ...trustworthy feedback. We’re hiring Senior Software Engineers across several... ...build speed, deploy safety, and test reliability scale across thousands of engineers...SeniorLocal area$152k - $241.5k
Memory System Architecture Engineer NVIDIA is searching for experienced candidates with a track record of architecture development to join our memory system architecture team! This team drives memory system architecture in NVIDIA's world changing SOCs for deep-learning...Senior- ...Senior Solutions Architect New York, Austin, Miami, Dallas Appnovation is a global, full-service... ...remains in place within its source systems. This position requires a senior,... ...OAuth/OIDC and identity providers; infrastructure-as-code (Terraform). Leadership...SeniorContract work
$100k - $150k
...Linux Canopy Connect is building the infrastructure that powers best-in-class insurance... ...pipeline and tooling to maintain world-class reliability and security. We are a product-... ...our customers. What you'll do Architect, implement and enforce best practices...Senior$208k - $327.75k
NVIDIA Enterprise Platforms Group is seeking a Senior System Architect to define, design, and validate enterprise AI factory reference architectures... ...system architecture, customer requirements, and hands-on infrastructure validation, helping turn NVIDIA accelerated computing,...SeniorFull time- ...experience in complex mechanical design and system analysis for enterprise-scale hardware.... ...technical intent, manage product reliability, and collaborate effectively with manufacturing... ...Skills Mechanical design Data center infrastructure Feasibility studies Product testing...SeniorRelocation package
- ...together as a team, and how we develop as individuals. The Senior Solutions Architect is responsible for the design of complex end to end... ...Have architectural skills dealing with data modeling and infrastructure solutions, such as classical data warehousing and big data...SeniorWork at officeRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior System Architect, Infrastructure Reliability. Be the first to apply!
- pega system architect Austin, TX
- embedded systems architect Austin, TX
- servicenow technical architect Austin, TX
- technical architect Austin, TX
- system architect Austin, TX
- senior operations technician Austin, TX
- senior operations associate Austin, TX
- senior cloud service delivery manager Austin, TX
- senior it service manager Austin, TX
- senior project engineer Austin, TX




