Senior System Architect, Infrastructure Reliability
$184k - $287.5kNVIDIA
NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects.What you'll be doing:Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure.Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs.Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters.Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams.Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs.What we need to see:Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming.Experience building automated RCA (Root Cause Analysis) pipelines for HPC or cloud-scale environments.CPU Architecture Deep-Dive: Expert knowledge of x86/ARM node-level metrics: IPC (Instructions Per Cycle), cache contention, NUMA imbalance, and hardware interrupts.Programming Proficiency: Strong C++ and Python skills, with the ability to build high-performance daemons that monitor system health without impacting workload performance.Scale Experience: Familiarity with cluster resource managers (Slurm, LSF, or Kubernetes) and how they manage job lifecycle and signal propagation.Ways To Stand Out From The Crowd:Low-Level Diagnostics: Expert knowledge of the Linux kernel and its error-reporting interfaces (/dev/mcelog, dmesg, journald). Understand how the kernel handles hardware exceptions and memory faults.GPU Infrastructure Proficiency: Deep experience with the NVIDIA DCGM (Data Center GPU Manager) and NVIDIA Management Library (NVML) for monitoring device health and capturing state-dumps.Experience with tools doing non-intrusive monitoring of application health and syscall-level failure patterns.Experience with checkpoint/restore technologies (like CRIU) and their application in long-running EDA flows.#LI-Hybrid Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until June 19, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, MA, Westford; US, TX, Austin; US, NC, Durham; US, WA, RedmondType: Full time
$164.2k - $295.6k
...highly skilled and motivated Systems / Product Infrastructure Engineer to join our Data... ...and dependable.Job Title: Senior Principal Systems... ...energy efficiency and system reliability.Join forces with multiple... ...related incidents.Design and architect end-to-end data center infrastructure...SeniorFull timeTemporary workWork at officeLocal areaRemote workWorldwide$184k - $287.5k
NVIDIA is looking for an experienced infrastructure Solutions Architect. Do you want to be part of a team that brings Artificial Intelligence (AI) hardware... ..., NICs/HCAs, switches, and GPUsGood understanding of system hardware architecture impact on network performance,...SeniorFull time$308.6k - $417.5k
...In the Arm Central Technology Systems group we define SoC and system architectures and... ...for an experienced system performance architect to drive the definition and exploration... ...performance to craft the future of compute infrastructure performance across data center and...SeniorWork at officeLocal area$208k - $327.75k
NVIDIA Enterprise Platforms Group is seeking a Senior System Architect to define, design, and validate enterprise AI factory reference architectures... ...system architecture, customer requirements, and hands-on infrastructure validation, helping turn NVIDIA accelerated computing,...SeniorFull time$127k - $249k
We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs.The InfraSec team collaborates...SeniorLocal areaRemote workWorldwideFlexible hours$153.6k - $207.8k
...generative AI and agentic systems. We are looking for a Solutions Architect who wants to be at the... ...less. AWS is hiring a Senior Solutions Architect to... ...platforms for performance, reliability, and cost as they scale... ...in applications and infrastructures experience- 5+ years of...SeniorLocal areaWorldwideFlexible hours$153.6k - $207.8k
...trajectory to success?As a Senior Solutions Architect within the AWS Worldwide... ...use these inputs to design reliable and scalable cloud architectures... ...with development, infrastructure, security and IT operations... ...development, cloud computing, systems engineering,...SeniorLocal areaWorldwideFlexible hours$153.6k - $207.8k
...for an experienced Solution Architect to help space companies achieve... ..., robotic/autonomous systems, mission control center operations... ..., cost, performance, reliability and operational efficiency.... ...computing, systems engineering, infrastructure, security, networking, data...SeniorLocal areaFlexible hours$152k - $241.5k
...the world.We seek a Solutions Architect to join our focused and hardworking AI Factory infrastructure deployment team. NVIDIA is at... ...edge computing. If you enjoy system building and have a... ...experience in NCP, CSP, site reliability, and virtualization technologies...Full timeWork experience placementRemote workWorldwide- ...Cloudflare is hiring a Senior Systems Engineer to help build AI Gateway, Cloudflare’s edge-focused AI infrastructure layer. You will work across the stack—from distributed systems... ...building durable, observable systems for reliable AI inference at scale. You will...Senior
$150k - $218k
...lead efforts to improve existing test infrastructure or create new test infrastructure to increase... ...and tools.Experience in embedded systems.Experience in test automation.Preferred... ...Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include...SeniorWorldwide- ...responsible for the core systems that drive our fleet’... ...We are looking for a Senior Software Engineer to... ...this role, you will architect systems that handle... ...availability services where reliability is paramount. Your... ...(AWS) and modern infrastructure-as-code practices....SeniorFull timeRemote workRelocationShift work
- ...centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and... ...Basis of Design (BoD) documents.Architect end‑to‑end data center solutions including... ...inference, mixed workloads) into scalable, reliable infrastructure designs optimized for AMD platforms....
- ...serve our clients. As a Senior Engineer on AI.x, you... ....As a Senior AI Site Reliability Engineer you will... ...will work closely with architects and engineers to... ...reliability of critical AI systems. Above all, you will... ...improvementsChampion Infrastructure-as-Code (IaC)...SeniorFull time
$152k - $241.5k
NVIDIA is searching for experienced candidates with a track record of architecture development to join our memory system architecture team! This team drives memory system architecture in NVIDIA’s world changing SOCs for deep-learning, autonomous vehicles and robotics,...SeniorFull time$184k - $287.5k
...computing (HPC), gaming, virtual reality, and autonomous vehicles? Come join the CPU performance architecture team as a Senior System Simulation Architect and help us push performance boundaries for NVIDIA’s line of CPU products!What you’ll be doing:Develop full-system...SeniorFull time- ...2K SRE team owns the infrastructure behind every player connection... ...-service events push systems to their limits, and... ...!The RoleThe Senior SRE at 2K is a hands-... ...network engineers, systems architects, and game studio... ..., influencing reliability from architecture review...Senior
- ...join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a... ...available, secure, and scalable Azure cloud infrastructure supporting TeamViewer’s global SaaS... ...monitoring, alerting, and automation systems to guarantee 24/7 service reliability...SeniorTemporary workCasual workWorldwide
$184k - $287.5k
...Team means contributing to the infrastructure that powers our innovative... ...implementing software and systems engineering practices to ensure... ...of AI systems.As a senior DGX Cloud AI Infrastructure... ...Define meaningful and actionable reliability metrics to track and improve...SeniorFull timeRemote work$188k - $274k
...leader in industry forums to shape digital infrastructure while proactively addressing market... ...Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include... ...Networking, Data Center operations, systems research, and much more.Individual pay...SeniorWorldwide- ...Role: We are looking for a Senior SRE to join our Platform Engineering... ...’ll be responsible for the reliability, scalability, and continued... ...faster insight into their systems. What You’ll Work On:... ...for on-premises observability infrastructure - including Elasticsearch cluster...SeniorFull time
$184k - $287.5k
...building the software and systems that power the world’... ...We are looking for a Senior Software Engineer to... ...run efficiently and reliably at scale. You will... ...large-scale AI clusters, infrastructure, and end-to-end... ...Proven track record of architecting, debugging, and scaling...SeniorFull timeRemote work$126.2k - $264.1k
...solve complex problems in distributed systems, networking, multi-tenant Infrastructure-as-a-Service (IaaS), and Software... ...as designed.We are looking for an Architect who will contribute to and direct... ...skills and can influence senior leadership in a positive way to make...SeniorTemporary workFlexible hours- ...engineering teams with tools, processes, and infrastructure to develop software for our business. About the Role:We are looking for a Senior SRE to serve as the operations owner for... ...tooling ecosystemsOwn the operational reliability of developer tooling ecosystems,...SeniorFull timeLocal area
- ...of out-of-the-box communication systems for satellites, UAVs, launch vehicles... ...join our team.We are seeking a Senior Software Engineer - Test Automation & Infrastructure to develop and scale automated... ...and infrastructure that ensure reliable, repeatable system validation....SeniorPermanent employmentFull timeContract workWork experience placementLocal area
- ...global engineering force has the most reliable and efficient build environments possible... ...an experienced and highly skilled Senior Infrastructure Engineer tosupport and operate a modern... ...identity platforms, endpoint management systems,cloud infrastructure, and security...SeniorFull timeWork at officeRemote workWorldwide
$127k - $249k
...We are looking for an experienced Senior Engineer for our SRE, Atlas team... ...be able to design & build complex systems, operate with autonomy and act as... ...are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This role requires engineers...SeniorLocal areaRemote workWorldwideFlexible hours$185k - $335.3k
...strong, impact-driven expert in ML Training Infrastructure with a demonstrated ability to lead... ...design and development of scalable, reliable, and high-performance AI/ML platform infrastructure... ...hardware environments.Raise the bar on system observability, debuggability,...Full timeLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours$127k - $249k
...is responsible for a range of critical infrastructure and operational functions that support... ...mesh), and observability and alerting systems.The Fleet Management team provides the... ...critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager...SeniorWork at officeLocal areaRemote workWorldwideFlexible hours$133.1k - $306.4k
As a Senior Manager, you will lead a team responsible for the development... ...fabrics and supporting systems. This role requires deep... ...make these fabrics more reliable, observable, and efficient at... ...for large-scale distributed infrastructure.Only Oracle brings together...SeniorTemporary workFlexible hoursNight shift
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior System Architect, Infrastructure Reliability. Be the first to apply!
- pega system architect Austin, TX
- system architect Austin, TX
- servicenow technical architect Austin, TX
- embedded systems architect Austin, TX
- technical architect Austin, TX
- senior manufacturing manager Austin, TX
- senior business analyst Austin, TX
- senior risk manager Austin, TX
- senior cost estimator Austin, TX
- senior manager tax Austin, TX


