Senior System Architect, Infrastructure Reliability
$184k - $287.5kNVIDIA
NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects.What you'll be doing:Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure.Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs.Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters.Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams.Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs.What we need to see:Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming.Experience building automated RCA (Root Cause Analysis) pipelines for HPC or cloud-scale environments.CPU Architecture Deep-Dive: Expert knowledge of x86/ARM node-level metrics: IPC (Instructions Per Cycle), cache contention, NUMA imbalance, and hardware interrupts.Programming Proficiency: Strong C++ and Python skills, with the ability to build high-performance daemons that monitor system health without impacting workload performance.Scale Experience: Familiarity with cluster resource managers (Slurm, LSF, or Kubernetes) and how they manage job lifecycle and signal propagation.Ways To Stand Out From The Crowd:Low-Level Diagnostics: Expert knowledge of the Linux kernel and its error-reporting interfaces (/dev/mcelog, dmesg, journald). Understand how the kernel handles hardware exceptions and memory faults.GPU Infrastructure Proficiency: Deep experience with the NVIDIA DCGM (Data Center GPU Manager) and NVIDIA Management Library (NVML) for monitoring device health and capturing state-dumps.Experience with tools doing non-intrusive monitoring of application health and syscall-level failure patterns.Experience with checkpoint/restore technologies (like CRIU) and their application in long-running EDA flows.#LI-Hybrid Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until June 19, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, MA, Westford; US, TX, Austin; US, NC, Durham; US, WA, RedmondType: Full time
$164.2k - $295.6k
...highly skilled and motivated Systems / Product Infrastructure Engineer to join our Data... ...and dependable.Job Title: Senior Principal Systems... ...energy efficiency and system reliability.Join forces with multiple... ...related incidents.Design and architect end-to-end data center infrastructure...SeniorFull timeTemporary workWork at officeLocal areaRemote workWorldwide$184k - $287.5k
NVIDIA is looking for an experienced infrastructure Solutions Architect. Do you want to be part of a team that brings Artificial Intelligence (AI) hardware... ..., NICs/HCAs, switches, and GPUsGood understanding of system hardware architecture impact on network performance,...SeniorFull time$151.2k - $204.6k
...touch operations. You'll own the roadmap for an AI-powered infrastructure reliability platform that prevents, detects, and resolves incidents... ...scientists and engineers. You will shape how LLMs, multi-agent systems, and machine learning are applied to one of the most...SeniorFull timeTemporary workSeasonal workFlexible hours$154.6k - $209.1k
...foundation that keeps Amazon's fulfillment network running 24/7. Infrastructure Reliability is building an AI-powered platform that detects, diagnoses... ...OpenSearch, CloudWatch, and internal incident management systems, publishing curated datasets to our data lake and Andes...SeniorFull timeTemporary workSeasonal workFlexible hours$308.6k - $417.5k
...In the Arm Central Technology Systems group we define SoC and system architectures and... ...for an experienced system performance architect to drive the definition and exploration... ...performance to craft the future of compute infrastructure performance across data center and...SeniorWork at officeLocal area$208k - $327.75k
NVIDIA Enterprise Platforms Group is seeking a Senior System Architect to define, design, and validate enterprise AI factory reference architectures... ...system architecture, customer requirements, and hands-on infrastructure validation, helping turn NVIDIA accelerated computing,...SeniorFull time$127k - $249k
We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs.The InfraSec team collaborates...SeniorLocal areaRemote workWorldwideFlexible hours$382.5k
...manufactured, and delivered. Our company is seeking a Fellow - AI Systems Architect for Product Operations to define the next generation of... ..., verification, Product & Test Engineering (PETE), Quality, Reliability, Product Operations, and Supply Chain Management. The...SeniorWork at officeLocal area$153.6k - $207.8k
...trajectory to success?As a Senior Solutions Architect within the AWS Worldwide... ...use these inputs to design reliable and scalable cloud architectures... ...with development, infrastructure, security and IT operations... ...development, cloud computing, systems engineering,...SeniorLocal areaWorldwideFlexible hours$153.6k - $207.8k
...generative AI and agentic systems. We are looking for a Solutions Architect who wants to be at the... ...less. AWS is hiring a Senior Solutions Architect to... ...platforms for performance, reliability, and cost as they scale... ...in applications and infrastructures experience- 5+ years of...SeniorLocal areaWorldwideFlexible hours$153.6k - $207.8k
...for an experienced Solution Architect to help space companies achieve... ..., robotic/autonomous systems, mission control center operations... ..., cost, performance, reliability and operational efficiency.... ...computing, systems engineering, infrastructure, security, networking, data...SeniorLocal areaFlexible hours$126.2k - $264.1k
...About Oracle Health Applications & Infrastructure Oracle Health Applications & Infrastructure... .... About the Role As an Senior Principal Product Manager, you will own... ...architecture, APIs, data flows, integrations, reliability, security, and scalability considerations...SeniorTemporary workFlexible hours$152k - $241.5k
...the world.We seek a Solutions Architect to join our focused and hardworking AI Factory infrastructure deployment team. NVIDIA is at... ...edge computing. If you enjoy system building and have a... ...experience in NCP, CSP, site reliability, and virtualization technologies...Full timeWork experience placementRemote workWorldwide- ...Senior Solutions Architect At Gallatin, we are rebuilding logistics infrastructure for the national security missions of the United States... ...partners. We build AI systems that determine how logistics... ...automate workflows, and improve reliability and scalability of deployed...SeniorWork at officeLocal area
$224k - $356.5k
We are seeking Systems Engineers and Software Engineers interested in building and running reliable large scale infrastructure platform services. In this organization, you will ensure that our internal and external facing EDA services atop of NVIDIA hardware are running...SeniorFull timeRemote work- Cloudflare is hiring a Senior Systems Engineer to help build AI Gateway, Cloudflare’s edge-focused AI infrastructure layer. You will work across the stack—from distributed systems... ...building durable, observable systems for reliable AI inference at scale. You will...Senior
$150k - $218k
...lead efforts to improve existing test infrastructure or create new test infrastructure to increase... ...and tools.Experience in embedded systems.Experience in test automation.Preferred... ...Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include...SeniorWorldwide- ...responsible for the core systems that drive our fleet’... ...We are looking for a Senior Software Engineer to... ...this role, you will architect systems that handle... ...availability services where reliability is paramount. Your... ...(AWS) and modern infrastructure-as-code practices....SeniorFull timeRemote workRelocationShift work
- ...serve our clients. As a Senior Engineer on AI.x, you... ....As a Senior AI Site Reliability Engineer you will... ...will work closely with architects and engineers to... ...reliability of critical AI systems. Above all, you will... ...improvementsChampion Infrastructure-as-Code (IaC)...SeniorFull time
- ...centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and... ...Basis of Design (BoD) documents.Architect end‑to‑end data center solutions including... ...inference, mixed workloads) into scalable, reliable infrastructure designs optimized for AMD platforms....
- ...2K SRE team owns the infrastructure behind every player connection... ...-service events push systems to their limits, and... ...!The RoleThe Senior SRE at 2K is a hands-... ...network engineers, systems architects, and game studio... ..., influencing reliability from architecture review...Senior
$184k - $287.5k
...computing (HPC), gaming, virtual reality, and autonomous vehicles? Come join the CPU performance architecture team as a Senior System Simulation Architect and help us push performance boundaries for NVIDIA’s line of CPU products!What you’ll be doing:Develop full-system...SeniorFull time$152k - $241.5k
NVIDIA is searching for experienced candidates with a track record of architecture development to join our memory system architecture team! This team drives memory system architecture in NVIDIA’s world changing SOCs for deep-learning, autonomous vehicles and robotics,...SeniorFull time- ...Role: We are looking for a Senior SRE to join our Platform Engineering... ...’ll be responsible for the reliability, scalability, and continued... ...faster insight into their systems. What You’ll Work On:... ...for on-premises observability infrastructure - including Elasticsearch cluster...SeniorFull time
$184k - $287.5k
...Team means contributing to the infrastructure that powers our innovative... ...implementing software and systems engineering practices to ensure... ...of AI systems.As a senior DGX Cloud AI Infrastructure... ...Define meaningful and actionable reliability metrics to track and improve...SeniorFull timeRemote work$184k - $287.5k
...building the software and systems that power the world’... ...We are looking for a Senior Software Engineer to... ...run efficiently and reliably at scale. You will... ...large-scale AI clusters, infrastructure, and end-to-end... ...Proven track record of architecting, debugging, and scaling...SeniorFull timeRemote work$126.2k - $264.1k
...solve complex problems in distributed systems, networking, multi-tenant Infrastructure-as-a-Service (IaaS), and Software... ...as designed.We are looking for an Architect who will contribute to and direct... ...skills and can influence senior leadership in a positive way to make...SeniorTemporary workFlexible hours- ...of out-of-the-box communication systems for satellites, UAVs, launch vehicles... ...join our team.We are seeking a Senior Software Engineer - Test Automation & Infrastructure to develop and scale automated... ...and infrastructure that ensure reliable, repeatable system validation....SeniorPermanent employmentFull timeContract workWork experience placementLocal area
- ...engineering teams with tools, processes, and infrastructure to develop software for our business. About the Role:We are looking for a Senior SRE to serve as the operations owner for... ...tooling ecosystemsOwn the operational reliability of developer tooling ecosystems,...SeniorFull timeLocal area
- ...global engineering force has the most reliable and efficient build environments possible... ...an experienced and highly skilled Senior Infrastructure Engineer tosupport and operate a modern... ...identity platforms, endpoint management systems,cloud infrastructure, and security...SeniorFull timeWork at officeRemote workWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior System Architect, Infrastructure Reliability. Be the first to apply!
- pega system architect Austin, TX
- system architect Austin, TX
- embedded systems architect Austin, TX
- technical architect Austin, TX
- senior manufacturing manager Austin, TX
- senior business analyst Austin, TX
- senior cost estimator Austin, TX
- senior manager tax Austin, TX
- senior devops Austin, TX
- senior recruiter Austin, TX


