Senior System Architect, Infrastructure Reliability
$184k - $287.5kNVIDIA
NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects.What you'll be doing:Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure.Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs.Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters.Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams.Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs.What we need to see:Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming.Experience building automated RCA (Root Cause Analysis) pipelines for HPC or cloud-scale environments.CPU Architecture Deep-Dive: Expert knowledge of x86/ARM node-level metrics: IPC (Instructions Per Cycle), cache contention, NUMA imbalance, and hardware interrupts.Programming Proficiency: Strong C++ and Python skills, with the ability to build high-performance daemons that monitor system health without impacting workload performance.Scale Experience: Familiarity with cluster resource managers (Slurm, LSF, or Kubernetes) and how they manage job lifecycle and signal propagation.Ways To Stand Out From The Crowd:Low-Level Diagnostics: Expert knowledge of the Linux kernel and its error-reporting interfaces (/dev/mcelog, dmesg, journald). Understand how the kernel handles hardware exceptions and memory faults.GPU Infrastructure Proficiency: Deep experience with the NVIDIA DCGM (Data Center GPU Manager) and NVIDIA Management Library (NVML) for monitoring device health and capturing state-dumps.Experience with tools doing non-intrusive monitoring of application health and syscall-level failure patterns.Experience with checkpoint/restore technologies (like CRIU) and their application in long-running EDA flows.#LI-Hybrid Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until June 19, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, MA, Westford; US, TX, Austin; US, NC, Durham; US, WA, RedmondType: Full time
$152k - $241.5k
NVIDIA is searching for experienced candidates with a track record of architecture development to join our memory system architecture team! This team drives memory system architecture in NVIDIA’s world changing SOCs for deep-learning, autonomous vehicles and robotics,...SeniorFull time$184k - $287.5k
...computing (HPC), gaming, virtual reality, and autonomous vehicles? Come join the CPU performance architecture team as a Senior System Simulation Architect and help us push performance boundaries for NVIDIA’s line of CPU products!What you’ll be doing:Develop full-system...SeniorFull time- ...immigration sponsorship for this positionThe RoleOur Site Reliability Engineering group within Enterprise Infrastructure combines Operations Excellence with the... ...have a background in either software engineering or systems engineering with a desire to learn the other or...SeniorFull time
$130 - $150 per hour
...DevOps / Cloud Security / Data Infrastructure Engineer who will build,... ...infrastructure-as-code, and platform reliability so that our engineering and... ...Ideal CandidateYou think in systems: pipelines, infrastructure... .... There may also be a more senior or junior position available...SeniorWork at office$184k - $287.5k
We are seeking an ambitious Senior Solutions Architect - AI Factory Deployment to join our NVIDIA Infrastructure Specialists team in Santa Clara! This role is uniquely positioned... ...) to understand workload behavior and system health.Develop automation (Python, Shell) for...SeniorFull timeRemote work$184k - $287.5k
NVIDIA's Infrastructure Specialists team is hiring a Senior Solutions Architect - AI Factory Observability & Visualization! This remote role develops full-spectrum visibility... ...that supports the smooth functioning of HPC systems and AI factories, transforming intricate...SeniorFull timeRemote work$152k - $241.5k
We are in search of a curious and motivated Senior Solutions Architect to join our NVIDIA Infrastructure Specialists team. In this capacity, you'll support the creation... ..., dashboards) to understand workload behavior and system health.Build automation using Python and Shell for...SeniorFull timeRemote work$152k - $241.5k
...intelligence.We’re looking for a Senior SRE to join our Compute Farm... ...keep critically important systems running while working on the... ...and network fabrics.Use IaC(Infrastructure‑as‑Code) and config... ...lifecycle management, fleet reliability/auto-healing, E2E observability...SeniorFull time- ...Senior Site Reliability Engineer (Enterprise Platform) Location: Remote - US - Open to Europe if... ...reliability of mission-critical, multi-region infrastructure. This is not a traditional support... ...an engineer who has operated real systems at scale and is eager to take end-to-...SeniorContract workCurrently hiringRemote work
- ...resilience of software and infrastructure through incident root... ...distributed multitiered systems at scale. Builds and operates... ...leads enterprise-level reliability strategies.Architects resilient systems and infrastructure... ...tools.Advises senior leadership on reliability...Full time
$174k - $253k
...troubleshooting large-scale distributed systems. 2 years of experience leading... ...Science or Engineering. About The Job Site Reliability Engineering (SRE) is what you get when... ...the architecture built by the Technical Infrastructure team to keep it running. From developing...Senior$160k - $240k
...way IT organizations work. We are currently looking for a Senior Site Reliability Engineer to join our SRE team in the Platform Engineering... ...ll be Doing Diagnose and resolve complex application and infrastructure issues Participate in our 24x7 on-call rotation, SCRUM,...SeniorPermanent employmentFull timeRemote workWork from homeRelocationFlexible hours$184k - $287.5k
NVIDIA is seeking elite ASIC Infrastructure engineers to deliver the tooling and environment that enables DV and RTL Design for the world's... ...GPU front end build flow Keep the GPU Continuous Integration system at the cutting edge of source management methodologies Guide compute...SeniorFull timeWork experience placementWork at officeWorldwide- ...Principal Site Reliability Engineer at Fidelity Investments in Durham, NC to facilitate & orchestrate data recovery events including organizational cloud & on-premise routing, failovers, & evidence captures. Req. Bachelor’s degree and 5 yrs. exp. or Master’s and 3 yrs...
- ...company) is seeking a Senior Electrical Engineer... ...), and other building systems for K–12, higher education... ...operating safely, reliably, and cost-effectively... ...clients navigate aging infrastructure and evolving operational... ...Interface with architects, civil and MEP engineers...SeniorWork from homeFlexible hours2 days per week
- ...IBM and other global technology leaders. We are seeking a Senior Solution Architect to join our Global Architecture team . This role is... ...work closely with business stakeholders, delivery teams, infrastructure teams, security, and application owners. This position is...SeniorRemote workWork from homeWorldwideHome officeFlexible hours
$152k - $241.5k
...work on building and maintaining the core infrastructure for deploying and running these agents... ...in architecture, performance, and reliability, enabling teams to bring to bear LLMs and... ...experience building production-grade software systems, and demonstrated experience shipping...SeniorFull time$152k - $241.5k
...responsible for development and support of infrastructure tools used by design engineers for... ...You'll be Doing:Work as a team to build reliable, scalable and high performance software... ...expertise in modern C++, compiler, build systems, and database.Experienced with static...SeniorFull timeWorldwide$152k - $241.5k
As a member of the Hardware Infrastructure EDA Compute team, you will optimize, scale, and support workload scheduling systems that directly impact design velocity and infrastructure... ...improvements in observability, service reliability, and automation, ensuring the EDA...SeniorFull time- ...modern workplaces? As a Senior Low Voltage Designer,... ...supporting network infrastructure in technology spaces,... ...and other low voltage systems in support of large-scale... ...Ability to collaborate with architects, engineers, project... ...to deliver secure, reliable, future-ready infrastructure...SeniorFull timeWork at officeOverseas
- ...Cybersecurity organization is seeking a Senior Data Engineer to join our... ....Hands-on knowledge of AWS infrastructure, including S3, EC2, RDS, EKS... ...cycle to ensure production reliability.Develop end-to-end solutions... ...and enhance production systems, focusing on high...SeniorFull timeWork experience placement
- Senior Systems Software Engineer This role has been designed as 'Hybrid' with a requirement... ...edge cases, performance limits, and reliability opportunities Debug and resolve challenging... ...characterization, or large-scale infrastructure Experience building tools or platforms...SeniorWork experience placementWork at officeLocal areaImmediate start2 days per week
$84.7k - $144.43k
...implementation of mechanical systems for high-performance,... ..., clean rooms, and critical infrastructure environments. This role involves... ...solutions that meet strict reliability and redundancy requirements.... ...candidate willCollaborate with architects, other engineers, and clients...SeniorWork at office$152k - $241.5k
...impact on the world.We are looking for a Senior Software Engineer to join our mission to continue improving our HPC infrastructure. Our team builds and operates... ...software development, crafting and building reliable distributed systems, and has the ability to implement...SeniorFull time$120k - $141.5k
...Wealthfront is seeking a Senior Staff Accountant to join the finance... ...team in building scalable infrastructure to support a dynamic and high... ...inaccuracies in and maintain reliability of financial documents by setting up internal control systems and embracing proper policies...SeniorRemote work- ...Senior Java Engineer Location: Durham, NC (100 New Millennium... ...architectures and AWS cloud infrastructure to support a strategic... ...developing distributed enterprise systems, optimizing heavy database... ...scalability, volume handling, and reliability. Qualifications &...Senior
$90k - $135k
...learning company, is hiring a Senior Software Engineer to play a... ..., plan, route work, and run reliably in production—not just... ...and distil complex agentic systems into clear, reliable solutions... ...building enterprise agentic AI infrastructure that teams across Pearson rely...SeniorFull time$107.9k - $172.64k
Job DescriptionThe Senior Data Modeler designs, implements, and documents more complex data architecture and data modeling solutions, which include the use of relational, dimensional, and NoSQL databases. These solutions support enterprise data management, business intelligence...SeniorFull timeWork experience placementWork at officeLocal areaRemote workFlexible hours2 days per week$25 - $30 per hour
...is looking for a Lab Support Senior Engineer to handle the hands... ...storage arrays and SAN infrastructure, and is comfortable physically... ...maintenance to keep equipment reliable and available Plan and... ...upgrades, and archival of legacy systems Coordinate with server,...SeniorFull timeContract workLocal areaShift work- ...SQS, DynamoDB, S3). ~ Understanding of SRE (Site Reliability Engineering) principles with supporting skills... ...experience in creating scalable and reliable software systems. ~ Extensive experience working with Infrastructure as Code -- Terraform, AWS Cloud Formation...SeniorContract work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior System Architect, Infrastructure Reliability. Be the first to apply!
- pega system architect Durham, NC
- system architect Durham, NC
- technical architect Durham, NC
- senior business analyst Durham, NC
- senior manager tax Durham, NC
- senior devops Durham, NC
- senior international accountant Durham, NC
- senior vmware engineer Durham, NC
- sr technical product manager Durham, NC
- senior resident engineer Durham, NC


