Senior System Architect, Infrastructure Reliability (Durham)
$184k - $287.5kNvidia
NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects.What you'll be doing:Architect Failure Attribution Frameworks: Build a scalable flight recorder for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure.Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs.Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters.Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as Hardware Fault, Software Bug, or Environment Issue. This reduces the Mean Time to Identify (MTTI) for R&D teams.Resiliency Engineering: Work closely with hardware and infrastructure teams to define signals of impending failure, enabling proactive job migration or check-pointing before a crash occurs.What we need to see:Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming.Experience building automated RCA (Root Cause Analysis) pipelines for HPC or cloud-scale environments.CPU Architecture Deep-Dive: Expert knowledge of x86/ARM node-level metrics: IPC (Instructions Per Cycle), cache contention, NUMA imbalance, and hardware interrupts.Programming Proficiency: Strong C++ and Python skills, with the ability to build high-performance daemons that monitor system health without impacting workload performance.Scale Experience: Familiarity with cluster resource managers (Slurm, LSF, or Kubernetes) and how they manage job lifecycle and signal propagation.Ways To Stand Out From The Crowd:Low-Level Diagnostics: Expert knowledge of the Linux kernel and its error-reporting interfaces (/dev/mcelog, dmesg, journald). Understand how the kernel handles hardware exceptions and memory faults.GPU Infrastructure Proficiency: Deep experience with the NVIDIA DCGM (Data Center GPU Manager) and NVIDIA Management Library (NVML) for monitoring device health and capturing state-dumps.Experience with tools doing non-intrusive monitoring of application health and syscall-level failure patterns.Experience with checkpoint/restore technologies (like CRIU) and their application in long-running EDA flows.#LI-Hybrid Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until June 19, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, MA, Westford; US, TX, Austin; US, NC, Durham; US, WA, RedmondType: Full time
$184k - $287.5k
...We are seeking an ambitious Senior Solutions Architect - AI Factory Deployment to join our NVIDIA Infrastructure Specialists team in Santa Clara... ...workload behavior and system health.Develop automation (Python... ...; US, TX, Austin; US, NC, Durham; US, CA, Santa ClaraType: Full...SeniorFull timePart timeRemote work$184k - $287.5k
...NVIDIA's Infrastructure Specialists team is hiring a Senior Solutions Architect - AI Factory Observability & Visualization! This remote... ...the smooth functioning of HPC systems and AI factories, transforming... ...; US, TX, Austin; US, NC, Durham; US, CA, Santa ClaraType: Full...SeniorFull timePart timeRemote work$184k - $287.5k
...computing (HPC), gaming, virtual reality, and autonomous vehicles? Come join the CPU performance architecture team as a Senior System Simulation Architect and help us push performance boundaries for NVIDIA’s line of CPU products!What you’ll be doing:Develop full-system...SeniorFull timePart time$80 per hour
...Onsite - Research Triangle Park, Durham, NC 27709 Max vendor rate is $80/hr... ...days per week/ 8 hours per day Senior Cloud Infrastructure Engineer Top Skills Required... ...troubleshoot, and optimize platform reliability, performance, and security • Comfortable...Senior$184k - $287.5k
...NVIDIA is seeking elite ASIC Infrastructure engineers to deliver the tooling and environment that enables DV and RTL Design for the world'... ...GPU front end build flow Keep the GPU Continuous Integration system at the cutting edge of source management methodologies Guide compute...SeniorFull timePart timeWork experience placementWork at officeWorldwide$174k - $253k
...following: Raleigh, NC, USA; Durham, NC, USA. Minimum... ...troubleshooting large-scale distributed systems. 2 years of experience... ...Engineering. About The Job Site Reliability Engineering (SRE) is what... ...architecture built by the Technical Infrastructure team to keep it running....Senior$272k - $431.25k
...help scale up its AI Infrastructure. We expect you to have... ...for production systems that enable large scalable... ...enable industry leading reliability, availability, and... ...teams, principles, and architects and coordinate... ...SummaryLocation: US, NC, Durham; US, RemoteType: Full...Full timePart time$224k - $356.5k
..., including AI/DL researchers, hardware architects, and software engineers.Participate in technology... ...development.Expertise with programming systems such as Python, C+, CUDA, and deep... ....SummaryLocation: US, CA, Santa Clara; US, NC, Durham; US, NC, RemoteType: Full time...SeniorFull timePart time$184k - $287.5k
...seeking a world-class computer architect to contribute to the... ...high-performance computing systems, with a focus on enhancing the... ...outside partners on system infrastructure. What we want to see:10+ yrs... ...Clara; US, TX, Austin; US, NC, Durham; US, CA, Remote; US, WA, SeattleType...SeniorFull timePart timeRemote work$224k - $356.5k
..., validated recipes, and model bring-up infrastructure — that lets developers run groundbreaking... ...with depth in GPU computing, ML systems, or high-performance inferenceStrong Python... ...Clara; US, MA, Westford; US, TX, Austin; US, NC, Durham; US, WA, SeattleType: Full time...SeniorFull timePart timeLocal area$272k - $431.25k
...: giving robotic arms and dexterous systems the perception, grasping, motion, and... ...accelerated computing and modern ML infrastructure, and how to architect robotics software to take advantage... ...Austin; US, VA, Charlottesville; US, NC, Durham; US, CO, Boulder; US, VA,...SeniorFull timePart time- ...Senior DevOps EngineerThis role has been designed as ‘Hybrid’... ...with deep expertise in Linux systems, Kubernetes platforms, virtualization... ...automate enterprise-scale infrastructure and platform services that... ...environments while ensuring reliability, scalability, security, and...SeniorPart timeWork experience placementWork at office2 days per week
$228.4k - $289.2k
...foundation models that enhance reliability, strengthen security,... ...environment. Your Impact As a Senior Engineering Manager, AI, you... ...fluency in machine learning systems, large language models, and... ...revolutionizing how data and infrastructure connect and protect organizations...SeniorFull timeTemporary workPart timeLocal areaFlexible hours- ...Senior Director of Product Management, Private Cloud EdgeThis role... ...deliver solutions that span infrastructure and software capabilities. The... ....Broad Technical Depth: Systems level understanding of cloud... ...America; San Jose, California; Durham, North Carolina; Ft. Collins,...SeniorFull timePart timeWork experience placementWork at officeLocal areaImmediate startRemote work2 days per week
$176.1k - $250.4k
...number of applications are received.Senior Staff Product Manager – Platform... ...enterprises to secure and ensure the reliability of their digital systems. Join us as we pursue our next chapter... ...to deliver crypto-resilient infrastructure that customers rely on every second...SeniorFull timeTemporary workPart timeLocal areaFlexible hours$165k - $241.4k
...and throughout Cisco to help build the infrastructure and data pipelines to power their data-... ...Linux environment. Focus on increasing reliability, performance, scalability, testability,... ...requests. Minimum QualificationsDistributed Systems: Experience analyzing, scaling,...SeniorPermanent employmentFull timeTemporary workPart timeLocal areaFlexible hours- ...Job Title: Senior Ruby and Node Developer Job Description: We are seeking a Senior... ...designing, developing, and optimizing back-end systems and ensuring that our applications are... ....js applications, ensuring scalability, reliability, and performance. Collaborate with...Senior
- ...HOSPITALITY GROUP: HOUSEKEEPING SUPERVISOR JOB DESCRIPTION SUMMARY: The Housekeeping Supervisor for Summit Hospitality Group/Hyatt Place Durham Southpoint is responsible for ensuring property guestrooms/suites, all public space, and associate areas are clean and well...SeniorFlexible hoursShift work
$184k - $287.5k
...has scaled exponentially. We are seeking a Senior Systems Performance Engineer to build our next generation of profiling infrastructure. You will be responsible for measuring,... ...future of computing!What you'll be doing:Architecting and maintaining custom profiling frameworks...SeniorFull timePart time$168k - $270.25k
...expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, MA, Westford; US, TX, Austin; US, NC, Durham; US, WA, Redmond; US, NY, New YorkType: Full time...SeniorFull timePart timeWeekend work- ...to join our team in Durham, NC. Our professionals... ...), and other building systems for K–12, higher... ...facilities operating safely, reliably, and cost-effectively... ...navigate aging infrastructure and evolving operational... ...solutions Interface with architects, civil and MEP...SeniorWork from homeFlexible hours2 days per week
- ...business goals.• Influence how platforms and systems enable business and customer outcomes... ...at all levels of the organizationYou have senior level experience leading the development... ...certain Criminal Histories.SummaryLocation: Durham, NC; Merrimack, NH; Smithfield, RIType:...SeniorFull timePart timeLocal area
- ...Senior Data Platform Engineer This role requires... ...time onsite presence in Durham, NC. Please note: We... ...own and evolve the data infrastructure that powers our EV... ...responsible for keeping it reliable, performant, and... ...issues across multiple systems. Excellent communication...SeniorPermanent employmentFull time
$224k - $356.5k
...We are seeking a senior system software engineer to work on next-generation Data Center GPU... ...C++ diagnostic workloads and software infrastructure required for new chip development,... .... Assessing new hardware features and architecting manufacturing and field diagnostic tests...SeniorFull timePart time$184k - $287.5k
...searching for a highly motivated, technical engineer to join the Tegra system-on-chip (SoC) software organization. You will work on key... ...SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, OR, Hillsboro; US, NC, Durham; US, CA, Remote; US, WA, RedmondType: Full time...SeniorFull timePart timeRemote work$184k - $287.5k
...We are now looking for a Senior GPU Architect! The NVIDIA GPU Architecture group is looking for... ...programming models, new architectures and new infrastructure that is required to make this... ..., test infrastructures or metrics systems including databases.Work in a team to...SeniorFull timePart time$184k - $287.5k
...We’re currently seeking a Senior Developer Technology Engineer, Artificial Intelligence!... ...rewarding to investigate, find, and eliminate system bottlenecks to achieve the best possible... ...; US, TX, Austin; US, OR, Hillsboro; US, NC, Durham; US, CA, RemoteType: Full time...SeniorFull timePart timeWork experience placement- ...Cloud Services portfolio across multiple systems, platforms, and applications.Management Level... ...and leadership to design and develop reliable, cost-effective, and high-quality solutions... ..., Texas, United States of America; Durham, North Carolina, United States of America...Full timePart timeWork experience placementWork at officeLocal areaImmediate start2 days per week
- ...West Colony Place, Durham, NC Reports to... ...SUMMARY The Senior Network Engineer... ...enterprise network infrastructure. This role... ...technology teams to architect resilient network... ...segmentation. Ensure reliable connectivity for... ...for critical systems. Participate in...SeniorFull timeWork at officeRemote work
$184k - $287.5k
...reality. We are seeking visionary computer architects to design and develop models for the... ...MS preferred) 8+ years of experience in systems architecture and modeling, especially performance... ...Austin; US, OR, Hillsboro; US, Remote; US, NC, Durham; US, CO, BoulderType: Full time...SeniorFull timePart timeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior System Architect, Infrastructure Reliability (Durham). Be the first to apply!
- senior hvac project manager Durham, NC
- senior medical science liaison Durham, NC
- senior accountant remote Durham, NC
- senior robotics software engineer Durham, NC
- sr project manager Durham, NC
- senior dynamics crm developer Durham, NC
- senior compensation manager Durham, NC
- senior storage engineer Durham, NC
- senior vice president of operations Durham, NC
- senior cloud engineer Durham, NC






