Senior System Architect, Infrastructure Reliability
$184k - $287.5kNVIDIA
NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects.What you'll be doing:Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure.Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs.Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters.Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams.Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs.What we need to see:Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming.Experience building automated RCA (Root Cause Analysis) pipelines for HPC or cloud-scale environments.CPU Architecture Deep-Dive: Expert knowledge of x86/ARM node-level metrics: IPC (Instructions Per Cycle), cache contention, NUMA imbalance, and hardware interrupts.Programming Proficiency: Strong C++ and Python skills, with the ability to build high-performance daemons that monitor system health without impacting workload performance.Scale Experience: Familiarity with cluster resource managers (Slurm, LSF, or Kubernetes) and how they manage job lifecycle and signal propagation.Ways To Stand Out From The Crowd:Low-Level Diagnostics: Expert knowledge of the Linux kernel and its error-reporting interfaces (/dev/mcelog, dmesg, journald). Understand how the kernel handles hardware exceptions and memory faults.GPU Infrastructure Proficiency: Deep experience with the NVIDIA DCGM (Data Center GPU Manager) and NVIDIA Management Library (NVML) for monitoring device health and capturing state-dumps.Experience with tools doing non-intrusive monitoring of application health and syscall-level failure patterns.Experience with checkpoint/restore technologies (like CRIU) and their application in long-running EDA flows.#LI-Hybrid Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until June 19, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, MA, Westford; US, TX, Austin; US, NC, Durham; US, WA, RedmondType: Full time
$184k - $287.5k
...seeking outstanding AI Solutions Architects to assist and support... ...design, deploy and optimize AI infrastructure.This role will focus on... ...technical advisor for accelerated systems architecture, GPU and... ...improve cluster utilization, reliability, performance and workload insightBuild...SeniorFull timeRemote work$255k - $340k
...is a leader in AI cloud infrastructure serving tens of thousands... ...from rack and pod level system arrangement that maximizes... ...infrastructure.We're looking for a Senior HPC Systems Architect with extensive experience... ..., scalability, and reliability.Evaluate emerging...SeniorWork at officeLocal areaWork from homeFlexible hours$184k - $287.5k
NVIDIA is seeking a Hardware Systems Architect to lead rack-level and platform pathfinding for... ...center teams to deliver high-performance, reliable AI platforms. Other responsibilities... ....Solid understanding of data center infrastructure: rack power distribution, network fabrics...SeniorFull time$184k - $287.5k
NVIDIA is looking for an experienced infrastructure Solutions Architect. Do you want to be part of a team that brings Artificial Intelligence (AI) hardware... ..., NICs/HCAs, switches, and GPUsGood understanding of system hardware architecture impact on network performance,...SeniorFull time$184k - $287.5k
...supercomputers.We are seeking a highly motivated Senior Solutions Architect to join the NVIDIA Cloud Partners team with a focus on GPU, NVLink, and infrastructure design. In this role, you will be at... ...in designing large-scale distributed systems, AI clusters, or HPC infrastructure....SeniorFull timeRemote work$262k - $364k
...distributed computing, large-scale system design, networking and data... ...Google Cloud customers.This senior software engineering... ...strong focus on quality and reliability throughout the manufacturing... ...deployment life-cycle.The AI and Infrastructure team is redefining what’s...SeniorWorldwide$155.42k - $205.9k
...team owns the cloud-agnostic, reliable, and cost-efficient platform... ...We’re proud to serve as the infrastructure platform for teams... ...the Role: We are seeking a Senior ML Infrastructure engineer to... ...running scalable distributed systems. They will rapidly test and...SeniorFull timeLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours$160k - $200k
...to join its fast-growing teams.As a Senior ML Infrastructure Engineer at Plus, you will design scalable... ...for managing model versioning systems and experiment tracking frameworks, which... ....Ensure high availability and reliability of the ML platform by implementing robust...SeniorFull time$153.2k - $234.1k
...breakthrough hardware and battery systems to intuitive design,... ...solutions that support safe and reliable autonomous vehicle behavior... ...real-world scenarios. As a Senior ML Infra Engineer, you will... ...systems, applications, or ML infrastructure. Experience designing robust...SeniorFull timeLocal areaRemote workWork from homeRelocation packageFlexible hours$153.2k - $234.1k
...breakthrough hardware and battery systems to intuitive design,... ...where we build the critical infrastructure that powers every machine learning... ...to use, and exceptionally reliable. Your success will be... ...advanced driverless vehicles.As a Senior ML Infra Engineer, you will...SeniorFull timeWork at officeLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours$170k - $240k
...driven expert in ML Training Infrastructure with a strong ability to execute... ...and building scalable, reliable, and high-performance AI/ML platform... ...initiatives. As a Senior ML Engineer, you will collaborate... ...and save cost.Raise the bar on system observability, debuggability,...SeniorFull timeLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours$152k - $241.5k
...seeking a hands-on, action-oriented Senior Solutions Architect to join our team, focused on the technical... ...role requires a strong passion for system design and a successful history of... ...high-performance, distributed AI infrastructure on-prem or in the cloud built with the...Senior- ...outstanding and visionary Lead Architect to drive the definition and... ...of next-generation system architectures and technologies... ...craft the future of compute infrastructure across data center and networking... ...the architectural vision to senior collaborators and technical...SeniorWork at officeLocal areaRemote work
$184k - $287.5k
...creativity and intelligence.NVIDIA is looking for a Datacenter System Architect to help define & design products for AI, high performance... ...our Datacenter System Architecture team and help build the infrastructure for the next industrial revolution.Your base salary will be...SeniorFull time$184k - $287.5k
...the world.NVIDIA is looking for a highly motivated Datacenter Systems Architect to join our multifaceted and innovative System Engineering... ...optimal balance of performance, quality, scalability, re-use, reliability, cost and time-to-marketDocument designs and plans to...SeniorFull time$184k - $287.5k
...and maintain our leadership. NVIDIA is seeking a motivated system architect to define future aspects of our GPU through employing pioneering... ...center workloads.Develop and enhance architecture analysis infrastructure, including performance simulators, testbench components and...SeniorFull timeWork experience placementRemote workNight shift$208k - $327.75k
NVIDIA Enterprise Platforms Group is seeking a Senior System Architect to define, design, and validate enterprise AI factory reference architectures... ...system architecture, customer requirements, and hands-on infrastructure validation, helping turn NVIDIA accelerated computing,...SeniorFull time$183k - $247.6k
...collaborate with a cross-functional team to drive system architecture across Amazon devices.Key job responsibilitiesAs a Senior System Architect, you will be responsible for defining... ..., Product Design, Industrial Design, Reliability, and Operations.You are a hands-on...SeniorLocal areaFlexible hours$184k - $287.5k
We are looking for a Senior Solutions Architect to help leading Enterprise ISVs... ...production-grade Agentic AI systems on NVIDIA’s accelerated... ...frontier AI capabilities into reliable enterprise products,... ...technologies and GPU-accelerated infrastructure to bring next-generation...SeniorFull time$152k - $241.5k
...and amazing people. NVIDIA is looking for an experienced Senior Software and System Architect to join our Networking Software Architecture group.... ...solutions to complex problemsWriting effective, clear and reliable architecture specificationsEvaluating new technologies,...SeniorFull timeRemote work$148k - $235.75k
...world.Join NVIDIA as a Solution Architect on the First Time Deployment... ...the world's leading AI infrastructure from blueprint to large-scale... ...This involves architectural systems, power distribution, cooling... ...the operational efficiency, reliability, and readiness of data center...SeniorFull timeWork at officeRemote workWorldwide$184k - $287.5k
...experienced Network Solutions Architect Engineer to help bring our... ...server, network, and cluster infrastructure in customer data centers.... ...on advanced GPU and network systems (Spectrum-X, BlueField DPU,... ...software to deliver performant, reliable AI clusters.Identify and...SeniorFull timeRemote work- ...A leading cybersecurity firm based in Santa Clara is seeking a Sr Site Reliability Engineer. The candidate will be responsible for maintaining highly reliable cloud infrastructure and will lead cross-functional initiatives. Ideal candidates should have at least 5 years...Senior
- ...A leading technology firm is in search of a Senior Wireless Network Site Reliability Engineer to manage and enhance their wireless network infrastructure. The ideal candidate has over 8 years of experience in wireless network operations and a strong background in wireless...Senior
$152k - $241.5k
...Cluster Administration and Site Reliability Engineering. You will... ...should be familiar with Linux system administration, Python, and... ...Collaborating with solution architects, engineering or product teams... ...previous work with data center infrastructure experience, from hardware up...Full timeWork at office$145k - $165k
...looking for a highly experienced Site Reliability Engineer (SRE). This role involves maintaining uptime and performance across systems. Exceptional Linux expertise and automation... ...include designing resilient infrastructure, monitoring environments, and responding...Senior$209k - $235k
The Role We're looking for a Senior Software Engineer to join our ML Infrastructure team and support the foundational infrastructure... ...dispatch commands must execute reliably in real time, and ML models need... ...well) Strong distributed systems and infrastructure skills:...SeniorFull timeHome officeFlexible hours- Cerebras Systems builds the world's largest AI chip, 56 times larger... ...ensure Cerebras systems are reliably deployed, operated, and... ...Systems EngineeringAI Cloud Infrastructure & OperationsNetwork & Storage... ...metrics, and operational risks to senior leadershipRequired...Senior
$183.7k - $248.6k
...opportunity Unity is looking for a Senior Machine Learning Infrastructure Engineer to join our Vector Ads team, where we build the real-time systems that power Unity's global... ...bidding, and targeting systems run reliably at scale. This is a great opportunity...SeniorWork at officeRemote workWorldwideRelocation package$224k - $356.5k
...completing one? Do you look at a build system, a release pipeline, or a... ...? We're looking for a senior engineer to invent, set, and construct the infrastructure and AI direction that enables us... ...business partnersOwn platform reliability: define SLIs and SLOs, instrument...SeniorFull timeImmediate start
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior System Architect, Infrastructure Reliability. Be the first to apply!
- pega system architect Santa Clara, CA
- system architect Santa Clara, CA
- technical architect Santa Clara, CA
- senior manager tax Santa Clara, CA
- senior devops Santa Clara, CA
- senior director digital marketing Santa Clara, CA
- senior international accountant Santa Clara, CA
- senior vmware engineer Santa Clara, CA
- sr marketing manager Santa Clara, CA
- sr technical product manager Santa Clara, CA

