Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior System Architect, Infrastructure Reliability

$184k - $287.5k

NVIDIA

NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects.What you'll be doing:Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure.Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs.Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters.Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams.Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs.What we need to see:Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming.Experience building automated RCA (Root Cause Analysis) pipelines for HPC or cloud-scale environments.CPU Architecture Deep-Dive: Expert knowledge of x86/ARM node-level metrics: IPC (Instructions Per Cycle), cache contention, NUMA imbalance, and hardware interrupts.Programming Proficiency: Strong C++ and Python skills, with the ability to build high-performance daemons that monitor system health without impacting workload performance.Scale Experience: Familiarity with cluster resource managers (Slurm, LSF, or Kubernetes) and how they manage job lifecycle and signal propagation.Ways To Stand Out From The Crowd:Low-Level Diagnostics: Expert knowledge of the Linux kernel and its error-reporting interfaces (/dev/mcelog, dmesg, journald). Understand how the kernel handles hardware exceptions and memory faults.GPU Infrastructure Proficiency: Deep experience with the NVIDIA DCGM (Data Center GPU Manager) and NVIDIA Management Library (NVML) for monitoring device health and capturing state-dumps.Experience with tools doing non-intrusive monitoring of application health and syscall-level failure patterns.Experience with checkpoint/restore technologies (like CRIU) and their application in long-running EDA flows.#LI-Hybrid Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until June 19, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, MA, Westford; US, TX, Austin; US, NC, Durham; US, WA, RedmondType: Full time

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior System Architect, Infrastructure Reliability in Durham, NC vacancy
  • $136k - $218.5k

     ...lasting impact on the world.NVIDIA is seeking a Senior ASIC Compose Verification Infrastructure and Tools Engineer to advance the systems that qualify, integrate, and deliver...  ...sophisticated engineering systems faster, more reliable, and easier to reason about. You’ll also... 
    Senior
    Full time

    Nvidia

    Durham, NC
    1 day ago
  •  ...marketplace, getting the platform right — reliable, observable, standardized, and cost-...  ...for execution at marketplace scale.As a Senior Infrastructure Engineer, you'll design and operate the large-scale distributed systems, data infrastructure, and deployment platforms... 
    Senior
    Full time
    Local area

    Fyber

    Durham, NC
    1 day ago
  • $152k - $241.5k

    NVIDIA is searching for experienced candidates with a track record of architecture development to join our memory system architecture team! This team drives memory system architecture in NVIDIA’s world changing SOCs for deep-learning, autonomous vehicles and robotics,... 
    Senior
    Full time

    Nvidia

    Durham, NC
    1 day ago
  • $184k - $287.5k

     ...computing (HPC), gaming, virtual reality, and autonomous vehicles? Come join the CPU performance architecture team as a Senior System Simulation Architect and help us push performance boundaries for NVIDIA’s line of CPU products!What you’ll be doing:Develop full-system... 
    Senior
    Full time

    Nvidia

    Durham, NC
    4 days ago
  • $224k - $356.5k

    We are seeking Systems Engineers and Software Engineers interested in building and running reliable large scale infrastructure platform services. In this organization, you will ensure that our internal and external facing EDA services atop of NVIDIA hardware are running... 
    Senior
    Full time
    Remote work

    Nvidia

    Durham, NC
    3 days ago
  • $184k - $287.5k

    NVIDIA's Infrastructure Specialists team is hiring a Senior Solutions Architect - AI Factory Observability & Visualization! This remote role develops full-spectrum visibility...  ...that supports the smooth functioning of HPC systems and AI factories, transforming intricate... 
    Senior
    Full time
    Remote work

    Nvidia

    Durham, NC
    1 day ago
  • $152k - $241.5k

    We are in search of a curious and motivated Senior Solutions Architect to join our NVIDIA Infrastructure Specialists team. In this capacity, you'll support the creation...  ..., dashboards) to understand workload behavior and system health.Build automation using Python and Shell for... 
    Senior
    Full time
    Remote work

    Nvidia

    Durham, NC
    4 days ago
  •  ...sponsorship for this position The Role Our Site Reliability Engineering group within Enterprise Infrastructure combines Operations Excellence with the...  ...have a background in either software engineering or systems engineering with a desire to learn the other or... 
    Senior
    Full time

    Fidelity Investments

    Durham, NC
    1 day ago
  •  ...NetApp is seeking a Senior Software Engineer – Cloud Infrastructure to build, operate, and scale mission-critical cloud services supporting SaaS and IaaS...  ..., infrastructure, and operations. You will improve reliability, scalability, and operational efficiency through automation... 
    Senior

    NetApp

    Morrisville, NC
    1 day ago
  •  ...workstations, smartphones, tablets), infrastructure (server, storage, edge, high...  ...workloads worldwide. We are seeking a Senior Data Center System Architect to drive end-to-end architecture...  ...performance, scalability, power efficiency, reliability, thermal design, serviceability,... 
    Senior
    Full time
    Local area
    Worldwide

    Lenovo

    Morrisville, NC
    4 days ago
  •  ...Morrisville, NC is seeking an Advisory Researcher in AI Compute and Data Infrastructure to provide senior technical leadership for research, architecture, and development of intelligent Hybrid AI systems. You will shape architectures across GPUs, CPUs, memory, storage, and... 
    Senior

    Lenovo

    Morrisville, NC
    1 day ago
  • A leading technology solutions provider in North Carolina is seeking a Senior Software Engineer focused on platform performance and resilience. The role involves ensuring system reliability and performance across multi-tier systems using AI-enabled automation. Candidates... 
    Senior

    Toshiba Global Commerce Solutions - External

    Durham, NC
    1 day ago
  • NetApp seeks a Senior Software Engineer - Cloud Infrastructure to design, operate, and scale cloud services across SaaS and IaaS...  ...NetApp teams and cloud providers to enhance reliability, performance, and security for distributed systems at scale. #J-18808-Ljbffr NetApp
    Senior

    NetApp

    Morrisville, NC
    2 days ago
  • Google is seeking an experienced Site Reliability Engineer (SRE) to help design, deploy, and maintain highly available services. You will...  ...planetary scale. Ideal candidates bring strong software and systems background, with a track record of leading technical initiatives... 
    Senior

    Google

    Durham, NC
    3 days ago
  •  ...responsibilities for Synapse Enterprise Information System (EIS) product and the delivery of high-...  ...The role focuses on enhancing software reliability through strong manual testing practices...  ...solutions to keep transportation infrastructure, aerospace, and oil and gas assets safe... 
    Senior
    Local area
    Flexible hours
    Night shift

    FUJIFILM Corporation

    Durham, NC
    1 day ago
  • $216k - $414k

     ...work. Come join the team and see how you can make a lasting impact on the world. NVIDIA is looking for a highly motivated, creative architect with experience in high-performance networking technologies, DPUs and GPUs to join the NVIDIA Network Software Architecture team.... 
    Senior

    NVIDIA

    Durham, NC
    1 day ago
  • $184k - $287.5k

    NVIDIA is seeking elite ASIC Infrastructure engineers to deliver the tooling and environment that enables DV and RTL Design for the world's...  ...GPU front end build flow Keep the GPU Continuous Integration system at the cutting edge of source management methodologies Guide compute... 
    Senior
    Full time
    Work experience placement
    Work at office
    Worldwide

    Nvidia

    Durham, NC
    11 hours ago
  •  ...Principal Site Reliability Engineer at Fidelity Investments in Durham, NC to facilitate & orchestrate data recovery events including organizational cloud & on-premise routing, failovers, & evidence captures. Req. Bachelor’s degree and 5 yrs. exp. or Master’s and 3 yrs... 

    Fidelity Investments

    Durham, NC
    1 day ago
  • Direct Supply, Inc. is looking for a Senior Staff Software Engineer in Durham, NC to lead the design of AI-driven systems impacting senior living technology. This role requires collaboration with engineering leaders and a strong skill set in strategic thinking, technology... 
    Senior

    Direct Supply, Inc.

    Durham, NC
    1 hour ago
  • Kitware is seeking a Research/Development Engineer to design and implement large-scale software frameworks that enable cutting-edge computer vision and machine-learning algorithms to run across embedded, desktop, and cloud environments. You will contribute to open source...
    Senior

    Jobleads-US

    Carrboro, NC
    2 days ago
  • Lenovo is seeking a Senior Data Center System Architect in Morrisville, NC, to drive end-to-end data center architecture for cloud-scale platforms. You will define platform topology, coordinate CPU/GPU/memory, PCIe/CXL, and power/cooling decisions, and shape roadmaps with... 
    Senior

    Lenovo

    Morrisville, NC
    4 days ago
  •  ...IBM and other global technology leaders. We are seeking a Senior Solution Architect to join our Global Architecture team . This role is...  ...work closely with business stakeholders, delivery teams, infrastructure teams, security, and application owners. This position is... 
    Senior
    Remote work
    Work from home
    Worldwide
    Home office
    Flexible hours

    Syntax

    Morrisville, NC
    2 days ago
  •  .../Responsibilities • 10–14 years experience in cloud infrastructure engineering • Strong expertise in AWS, Kubernetes,...  ...capabilities • Proven ability to implement secure, reliable, and observable cloud systems • Experience with vulnerability management and cloud... 
    Senior

    Omni Inclusive

    Durham, NC
    3 days ago
  • $152k - $241.5k

     ...work on building and maintaining the core infrastructure for deploying and running these agents...  ...in architecture, performance, and reliability, enabling teams to bring to bear LLMs and...  ...experience building production-grade software systems, and demonstrated experience shipping... 
    Senior
    Full time

    Nvidia

    Durham, NC
    3 days ago
  •  ...pharmaceutical companies and health systems make confident decisions and...  ...Manager - Application Architect to join our team in Durham,...  ..., ensuring scalability, reliability, performance, and security.Provide...  ...CI/CD pipelines and cloud infrastructure performance.Lead performance... 
    Full time
    Temporary work
    Casual work
    Internship
    Work at office
    Monday to Friday
    Flexible hours
    Day shift
    3 days per week

    Laboratory Corporation of America

    Durham, NC
    1 day ago
  • $135.8k - $213.4k

     ...opportunities to work on revolutionary systems that impact people's lives around...  ..., they're making history.As an Infrastructure Product Owner (DevOps) - Level 4,...  ...Leaders from entry-level to the most senior chief engineers and architects to Product Owners and Scrum... 
    Senior
    Full time
    Remote work
    Relocation package
    Shift work

    Northrop Grumman

    Morrisville, NC
    2 days ago
  • FUJIFILM Biotechnologies is seeking a Principle Cloud Infrastructure Engineer to lead design, build, and operation of secure, scalable cloud platforms across IT and OT environments. You will drive cloud strategy, implement IaC, and enable compliant, observable services... 
    Senior

    FUJIFILM Biotechnologies

    Morrisville, NC
    2 days ago
  • $152k - $241.5k

    As a member of the Hardware Infrastructure EDA Compute team, you will optimize, scale, and support workload scheduling systems that directly impact design velocity and infrastructure...  ...improvements in observability, service reliability, and automation, ensuring the EDA... 
    Senior
    Full time

    Nvidia

    Durham, NC
    1 day ago
  • $152k - $241.5k

     ...responsible for development and support of infrastructure tools used by design engineers for...  ...You'll be Doing:Work as a team to build reliable, scalable and high performance software...  ...expertise in modern C++, compiler, build systems, and database.Experienced with static... 
    Senior
    Full time
    Worldwide

    Nvidia

    Durham, NC
    1 day ago
  • $146k - $220k

    Downtown Boulder Partnership is seeking a Director of IT in Morrisville, NC. The successful candidate will lead Infrastructure Engineering initiatives, focusing on IaaS and Data Intelligence. They should possess extensive experience in infrastructure management, cloud... 
    Senior

    Downtown Boulder Partnership

    Morrisville, NC
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior System Architect, Infrastructure Reliability. Be the first to apply!