Senior System Architect, Infrastructure Reliability
$184k - $287.5kNVIDIA
NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects.What you'll be doing:Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure.Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs.Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters.Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams.Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs.What we need to see:Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming.Experience building automated RCA (Root Cause Analysis) pipelines for HPC or cloud-scale environments.CPU Architecture Deep-Dive: Expert knowledge of x86/ARM node-level metrics: IPC (Instructions Per Cycle), cache contention, NUMA imbalance, and hardware interrupts.Programming Proficiency: Strong C++ and Python skills, with the ability to build high-performance daemons that monitor system health without impacting workload performance.Scale Experience: Familiarity with cluster resource managers (Slurm, LSF, or Kubernetes) and how they manage job lifecycle and signal propagation.Ways To Stand Out From The Crowd:Low-Level Diagnostics: Expert knowledge of the Linux kernel and its error-reporting interfaces (/dev/mcelog, dmesg, journald). Understand how the kernel handles hardware exceptions and memory faults.GPU Infrastructure Proficiency: Deep experience with the NVIDIA DCGM (Data Center GPU Manager) and NVIDIA Management Library (NVML) for monitoring device health and capturing state-dumps.Experience with tools doing non-intrusive monitoring of application health and syscall-level failure patterns.Experience with checkpoint/restore technologies (like CRIU) and their application in long-running EDA flows.#LI-Hybrid Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until June 19, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, MA, Westford; US, TX, Austin; US, NC, Durham; US, WA, RedmondType: Full time
$136k - $218.5k
...lasting impact on the world.NVIDIA is seeking a Senior ASIC Compose Verification Infrastructure and Tools Engineer to advance the systems that qualify, integrate, and deliver... ...sophisticated engineering systems faster, more reliable, and easier to reason about. You’ll also...SeniorFull time- ...marketplace, getting the platform right — reliable, observable, standardized, and cost-... ...for execution at marketplace scale.As a Senior Infrastructure Engineer, you'll design and operate the large-scale distributed systems, data infrastructure, and deployment platforms...SeniorFull timeLocal area
$152k - $241.5k
NVIDIA is searching for experienced candidates with a track record of architecture development to join our memory system architecture team! This team drives memory system architecture in NVIDIA’s world changing SOCs for deep-learning, autonomous vehicles and robotics,...SeniorFull time$184k - $287.5k
...computing (HPC), gaming, virtual reality, and autonomous vehicles? Come join the CPU performance architecture team as a Senior System Simulation Architect and help us push performance boundaries for NVIDIA’s line of CPU products!What you’ll be doing:Develop full-system...SeniorFull time$224k - $356.5k
We are seeking Systems Engineers and Software Engineers interested in building and running reliable large scale infrastructure platform services. In this organization, you will ensure that our internal and external facing EDA services atop of NVIDIA hardware are running...SeniorFull timeRemote work$184k - $287.5k
NVIDIA's Infrastructure Specialists team is hiring a Senior Solutions Architect - AI Factory Observability & Visualization! This remote role develops full-spectrum visibility... ...that supports the smooth functioning of HPC systems and AI factories, transforming intricate...SeniorFull timeRemote work$152k - $241.5k
We are in search of a curious and motivated Senior Solutions Architect to join our NVIDIA Infrastructure Specialists team. In this capacity, you'll support the creation... ..., dashboards) to understand workload behavior and system health.Build automation using Python and Shell for...SeniorFull timeRemote work- ...sponsorship for this position The Role Our Site Reliability Engineering group within Enterprise Infrastructure combines Operations Excellence with the... ...have a background in either software engineering or systems engineering with a desire to learn the other or...SeniorFull time
- ...NetApp is seeking a Senior Software Engineer – Cloud Infrastructure to build, operate, and scale mission-critical cloud services supporting SaaS and IaaS... ..., infrastructure, and operations. You will improve reliability, scalability, and operational efficiency through automation...Senior
$184k - $287.5k
NVIDIA is seeking elite ASIC Infrastructure engineers to deliver the tooling and environment that enables DV and RTL Design for the world's... ...GPU front end build flow Keep the GPU Continuous Integration system at the cutting edge of source management methodologies Guide compute...SeniorFull timeWork experience placementWork at officeWorldwide- ...Principal Site Reliability Engineer at Fidelity Investments in Durham, NC to facilitate & orchestrate data recovery events including organizational cloud & on-premise routing, failovers, & evidence captures. Req. Bachelor’s degree and 5 yrs. exp. or Master’s and 3 yrs...
- Kitware is seeking a Research/Development Engineer to design and implement large-scale software frameworks that enable cutting-edge computer vision and machine-learning algorithms to run across embedded, desktop, and cloud environments. You will contribute to open source...Senior
- ...IBM and other global technology leaders. We are seeking a Senior Solution Architect to join our Global Architecture team . This role is... ...work closely with business stakeholders, delivery teams, infrastructure teams, security, and application owners. This position is...SeniorRemote workWork from homeWorldwideHome officeFlexible hours
- .../Responsibilities • 10–14 years experience in cloud infrastructure engineering • Strong expertise in AWS, Kubernetes,... ...capabilities • Proven ability to implement secure, reliable, and observable cloud systems • Experience with vulnerability management and cloud...Senior
$152k - $241.5k
...work on building and maintaining the core infrastructure for deploying and running these agents... ...in architecture, performance, and reliability, enabling teams to bring to bear LLMs and... ...experience building production-grade software systems, and demonstrated experience shipping...SeniorFull time$135.8k - $213.4k
...opportunities to work on revolutionary systems that impact people's lives around... ..., they're making history.As an Infrastructure Product Owner (DevOps) - Level 4,... ...Leaders from entry-level to the most senior chief engineers and architects to Product Owners and Scrum...SeniorFull timeRemote workRelocation packageShift work- ...pharmaceutical companies and health systems make confident decisions and... ...Manager - Application Architect to join our team in Durham,... ..., ensuring scalability, reliability, performance, and security.Provide... ...CI/CD pipelines and cloud infrastructure performance.Lead performance...Full timeTemporary workCasual workInternshipWork at officeMonday to FridayFlexible hoursDay shift3 days per week
$152k - $241.5k
As a member of the Hardware Infrastructure EDA Compute team, you will optimize, scale, and support workload scheduling systems that directly impact design velocity and infrastructure... ...improvements in observability, service reliability, and automation, ensuring the EDA...SeniorFull time$152k - $241.5k
...responsible for development and support of infrastructure tools used by design engineers for... ...You'll be Doing:Work as a team to build reliable, scalable and high performance software... ...expertise in modern C++, compiler, build systems, and database.Experienced with static...SeniorFull timeWorldwide$135k - $150k
...Senior Electrical EngineerAs an Electrical Engineer... ...project feasibility, system performance, and... ...and mission-critical infrastructure.Ability to resolve advanced... ...quality, system reliability, and integration of emerging... ...HVAC Plumbing Architect Architecture Building...SeniorRelocation package- ...Summary IONNA is seeking a visionary Senior Vice President of Product to lead the... ...delivering the industry's most seamless, reliable, and rewarding EV charging experience.... ...effortless through world-class charging infrastructure, exceptional customer experiences, and...SeniorPermanent employmentFull time
$152k - $241.5k
...world.We are now looking for a Senior Full-stack web applications software architect to join our Hardware Infrastructure team! Our team is building... ...and AI agents that are reliable, scalable, and maintainable... ...Detailed knowledge of distributed systems principles, concurrency,...SeniorFull time- ...modern workplaces? As a Senior Low Voltage Designer,... ...supporting network infrastructure in technology spaces,... ...and other low voltage systems in support of large-scale... ...Ability to collaborate with architects, engineers, project... ...to deliver secure, reliable, future-ready infrastructure...SeniorFull timeWork at officeOverseas
$90k - $135k
...Senior Software Engineer Drive the future of Agentic AI at... ..., plan, route work, and run reliably in production—not just prototypes... ...and distil complex agentic systems into clear, reliable... ...building enterprise agentic AI infrastructure that teams across Pearson rely...SeniorFull time$84.7k - $144.43k
...implementation of mechanical systems for high-performance,... ..., clean rooms, and critical infrastructure environments. This role involves... ...solutions that meet strict reliability and redundancy requirements.... ...candidate willCollaborate with architects, other engineers, and clients...SeniorWork at office- ...Join FLBG's Global Infrastructure & Operations team as... ...CloudWatch) to ensure reliability, performance, and QoS... ...integration with existing systems and full compliance... ...Collaborate with senior stakeholders globally... ...Engineer or Solutions Architect advantageous ~ Microsoft...Full timeWork from home
- ...Senior DevOps EngineerLocation: Durham, NC /NJ Remote till covidDuration... ...in Software, Cloud, or Site Reliability Engineering fieldsBS or MS... ...-as-code!)Experience with Infrastructure as Code (Terraform, Python,... ...a plusUnderstanding of well architected framework implementation in...SeniorRemote work
- ...Engineer, you will design and build the data infrastructure that makes Vulcan’s operational and... ...models and pipelines, and building the systems that collect, contextualize, and... ...ready for analytics and AI workloads Build reliable ingest paths for structured data, time-...Permanent employmentContract workFor contractorsFor subcontractorWork at officeVisa sponsorship
- ...routine processes, optimizes infrastructure, and monitors data workflows to ensure availability and reliability. Provides business solutions... ...Technology, Information Systems, or a closely related field... ...3) years of experience as a Senior Software Engineer/Developer...SeniorFull time
- ...availability, performance, reliability, and user adoption Serve as... ...and compliance requirements Architect and deliver solutions that address... ...and support on-premises infrastructure at the Durham location, including... ...and related infrastructure systems Participate in after-hours...SeniorWork at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior System Architect, Infrastructure Reliability. Be the first to apply!
- pega system architect Durham, NC
- technical architect Durham, NC
- system architect Durham, NC
- senior associate architect Durham, NC
- senior dynamics crm developer Durham, NC
- senior application security Durham, NC
- senior cloud data engineer Durham, NC
- senior Durham, NC
- senior customer success engineer Durham, NC
- senior property accountant Durham, NC



