Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior System Architect, Infrastructure Reliability

$184k - $287.5k

NVIDIA

US, CA, Santa Clara

US, MA, Westford

US, TX, Austin

US, NC, Durham

US, WA, Redmond

Full time

JR2013698

NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects.

What you'll be doing:

  • Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure.

  • Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs.

  • Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters.

  • Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams.

  • Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs.

What we need to see:

  • Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming.

  • Experience building automated RCA (Root Cause Analysis) pipelines for HPC or cloud-scale environments.

  • CPU Architecture Deep-Dive: Expert knowledge of x86/ARM node-level metrics: IPC (Instructions Per Cycle), cache contention, NUMA imbalance, and hardware interrupts.

  • Programming Proficiency: Strong C++ and Python skills, with the ability to build high-performance daemons that monitor system health without impacting workload performance.

  • Scale Experience: Familiarity with cluster resource managers (Slurm, LSF, or Kubernetes) and how they manage job lifecycle and signal propagation.

Ways To Stand Out From The Crowd:

  • Low-Level Diagnostics: Expert knowledge of the Linux kernel and its error-reporting interfaces (/dev/mcelog, dmesg, journald). Understand how the kernel handles hardware exceptions and memory faults.

  • GPU Infrastructure Proficiency: Deep experience with the NVIDIA DCGM (Data Center GPU Manager) and NVIDIA Management Library (NVML) for monitoring device health and capturing state-dumps.

  • Experience with tools doing non-intrusive monitoring of application health and syscall-level failure patterns.

  • Experience with checkpoint/restore technologies (like CRIU) and their application in long-running EDA flows.

#LI-Hybrid

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

You will also be eligible for equity and benefits ( .

Applications for this job will be accepted at least until September 2, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

NVIDIA pioneered accelerated computing. Today, our AI infrastructure powers global intelligence, transforming every industry.

Learn more about NVIDIA .

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior System Architect, Infrastructure Reliability in Santa Clara, CA vacancy
  •  ...Oracle Cloud Infrastructure (OCI) seeks a Senior Principal Engineer to lead the design and implementation of reliability validation for OCI control plane services, focusing on a high-performance, low-level systems approach. You will mentor engineers, define validation... 
    Senior

    Jobleads-US

    Santa Clara, CA
    5 hours ago
  • $323k

     ...outstanding and visionary Lead Architect to drive the definition and...  ...of next-generation system architectures and technologies...  ...craft the future of compute infrastructure across data center and networking...  ...the architectural vision to senior collaborators and technical... 
    Senior
    Work at office
    Local area
    Remote work

    Arm Limited

    San Jose, CA
    3 days ago
  • $184k - $287.5k

     ...seeking outstanding AI Solutions Architects to assist and support...  ...design, deploy and optimize AI infrastructure.This role will focus on...  ...technical advisor for accelerated systems architecture, GPU and...  ...improve cluster utilization, reliability, performance and workload insightBuild... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  • $160k - $200k

     ...join its fast-growing teams. As a Senior ML Infrastructure Engineer at Plus, you will design scalable...  ...for managing model versioning systems and experiment tracking frameworks, which...  .... Ensure high availability and reliability of the ML platform by implementing robust... 
    Senior

    PlusAI

    Santa Clara, CA
    16 days ago
  • $152k - $241.5k

     ...and amazing people. NVIDIA is looking for an experienced Senior Software and System Architect to join our Networking Software Architecture group....  ...solutions to complex problems Writing effective, clear and reliable architecture specifications Evaluating new... 
    Senior
    Remote work

    NVIDIA Gruppe

    Santa Clara, CA
    2 days ago
  • $255k - $340k

     ...is a leader in AI cloud infrastructure serving tens of thousands...  ...from rack and pod level system arrangement that maximizes...  ...infrastructure.We're looking for a Senior HPC Systems Architect with extensive experience...  ..., scalability, and reliability.Evaluate emerging... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    7 days ago
  •  ...architecture, CI/CD, and containerization. The role focuses on reliability, scalability, and modernization with cross-team...  ...pipelines, containerization strategy, and modernization of hybrid infrastructures across AWS, OCI, and more. Strong SRE background and mentorship... 
    Senior

    Jobleads-US

    San Jose, CA
    1 day ago
  • $184k - $287.5k

    NVIDIA is seeking a Hardware Systems Architect to lead rack-level and platform pathfinding for...  ...center teams to deliver high-performance, reliable AI platforms. Other responsibilities...  ....Solid understanding of data center infrastructure: rack power distribution, network fabrics... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...Senior Infrastructure Engineer We are seeking a Senior Infrastructure Engineer with a strong focus...  ...and automation to build scalable, reliable, and efficient infrastructure solutions...  ...monitoring, and scaling processes to enhance system efficiency and reliability.... 
    Senior

    Omni Inclusive

    San Jose, CA
    2 days ago
  • $184k - $287.5k

    NVIDIA is looking for an experienced infrastructure Solutions Architect. Do you want to be part of a team that brings Artificial Intelligence (AI) hardware...  ..., NICs/HCAs, switches, and GPUsGood understanding of system hardware architecture impact on network performance,... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...alert: Principal Hardware Systems Architect-Compute Platforms Req ID...  ...Platforms to provide senior technical leadership for the...  ...understanding that next-generation AI infrastructure can no longer be optimized...  ..., power, thermal density, reliability, cost, scalability,... 
    Local area

    Celestica Inc.

    San Jose, CA
    4 days ago
  • $220k - $280k

     ...Credo is looking for a Principal AI System Architect to join our team in San Jose, CA, reporting...  .../GPU/custom silicon) into working AI infrastructure. This role needs to follow and define...  ...next. Our technology powers the most reliable and energy‑efficient connections around... 
    Work experience placement

    Socket.dev

    San Jose, CA
    15 hours ago
  •  ...performance, exceptional reliability, and outstanding...  ...software, operating-system memory management, accelerator...  ...responsibilities ~ Architect the model-aware...  ...designs; mentor the senior engineer; and translate...  ...systems, ML infrastructure, embedded systems, or... 
    Local area
    Remote work

    Lexarenterprise

    San Jose, CA
    3 days ago
  • $184k - $287.5k

     ...supercomputers.We are seeking a highly motivated Senior Solutions Architect to join the NVIDIA Cloud Partners team with a focus on GPU, NVLink, and infrastructure design. In this role, you will be at...  ...in designing large-scale distributed systems, AI clusters, or HPC infrastructure.... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    5 days ago
  • $262k - $364k

     ...distributed computing, large-scale system design, networking and data...  ...Google Cloud customers.This senior software engineering...  ...strong focus on quality and reliability throughout the manufacturing...  ...deployment life-cycle.The AI and Infrastructure team is redefining what’s... 
    Senior
    Worldwide

    Google

    Sunnyvale, CA
    5 days ago
  •  ...enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service...  ...Your Impact You will be the most senior technical individual contributor on...  ...the team's automation direction and infrastructure architecture decisions with regional... 
    Senior

    Outshift by Cisco

    San Jose, CA
    5 days ago
  •  ...development, and deployment of advanced AI agents and agentic systems. Architect and implement complex multi-agent systems, including...  ...autonomy and intelligence. Build robust, scalable, and reliable infrastructure to support the deployment and operation of AI agents at... 
    Senior
    Full time
    Work experience placement

    Eightfold

    Santa Clara, CA
    15 hours ago
  • $155.42k - $205.9k

     ...team owns the cloud-agnostic, reliable, and cost-efficient platform...  ...We’re proud to serve as the infrastructure platform for teams...  ...the Role: We are seeking a Senior ML Infrastructure engineer to...  ...running scalable distributed systems. They will rapidly test and... 
    Senior
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    4 days ago
  • $170k - $240.8k

     ...driven expert in ML Training Infrastructure with a strong ability to execute...  ...and building scalable, reliable, and high-performance AI/ML platform...  ...initiatives. As a Senior ML Engineer, you will collaborate...  ...training at scale.Raise the bar on system observability, debuggability,... 
    Senior
    Full time
    Local area
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    5 days ago
  • $153.2k - $234.1k

     ...breakthrough hardware and battery systems to intuitive design,...  ...solutions that support safe and reliable autonomous vehicle behavior...  ...real-world scenarios. As a Senior ML Infra Engineer, you will...  ...systems, applications, or ML infrastructure. Experience designing robust... 
    Senior
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    3 days ago
  • $160k - $240k

     ...Senior Site Reliability Engineer Calling all innovators - find your future at Fiserv. We're Fiserv...  ...monitoring, logging and alerting systems to ensure strong observability across...  ...Kubernetes). Working knowledge of Infrastructure as Code and configuration management... 
    Senior

    BentoBox

    Sunnyvale, CA
    4 days ago
  • $184k - $287.5k

     ...supercomputers. We are seeing a highly motivated Senior Solutions Architect to join the Cluster Design and...  ...AI supercomputers and enterprise AI infrastructure in the field. As a Solutions...  ...designing large-scale distributed systems, AI clusters, or HPC infrastructure... 
    Senior

    NVIDIA Gruppe

    Santa Clara, CA
    2 days ago
  •  ...via getnomad.app. Summary: As a Senior Solution Architect, you will be responsible for all...  ...understand LotusFlare's DNO platform, system architecture and integration components...  ...solutions around the LF DNO platform and infrastructure. Delivery: Drive project... 
    Senior
    Work at office
    Worldwide
    Flexible hours

    LotusFlare

    Santa Clara, CA
    15 hours ago
  •  ...We are seeking a Senior Database Reliability Engineer (DBRE) to design, operate, and improve reliable, scalable, secure, and...  ...engineering, site reliability engineering, Linux systems administration, and infrastructure automation. The ideal candidate will have strong... 
    Senior
    Full time

    Neshent Technologies

    San Jose, CA
    25 days ago
  •  ...with previous work with cloud infrastructure and managed data services, including...  ...production database or distributed data systems across application, database, operating...  ...Roles & Responsibilities As a Senior Database Reliability Engineer, you will help make Meraki'... 
    Senior
    Full time

    Northern Base

    San Jose, CA
    2 days ago
  • $240k - $333k

     ...Senior Leadership Technical Program Manager, Customer Experience, Global Infrastructure In accordance with Washington state law, we are highlighting our comprehensive benefits...  ...governance across GGI to deliver scalable, reliable infrastructure solutions. Google Cloud... 
    Senior
    Temporary work
    Work at office

    Google

    Sunnyvale, CA
    3 days ago
  •  ...NVIDIA is seeking a Senior Solutions Architect to join the Cluster Design and Architecture team, focusing on networking technologies. You...  ...thousands of GPUs, enabling AI supercomputers and enterprise AI infrastructures in the field. Responsibilities include collaborating... 
    Senior

    Jobleads-US

    Santa Clara, CA
    2 days ago
  •  ...Palo Alto Networks is seeking a Senior Principal Engineer/Architect to serve as the technical authority...  ...complex requirements into automated infrastructure. You will own projects end-to-...  ...innovation with AI tooling, and ensure reliability and cost efficiency in cloud-... 
    Senior

    Jobleads-US

    Santa Clara, CA
    4 days ago
  •  ...Dell Technologies in Santa Clara, CA is seeking a Senior AI Storage Solutions Lead, Solutions Architecture to drive end-to-end PoCs...  ...and leadership across multiple customer engagements, ensuring data infrastructure aligns with AI pipelines and #J-18808-Ljbffr Jobleads-US
    Senior

    Jobleads-US

    Santa Clara, CA
    4 days ago
  •  ...Role     We're looking for a distributed ML infrastructure engineer to help extend and scale our training systems. You’ll work side-by-side with world-class...  ...external visibility  • Improve training system reliability, maintainability, and performance  • While much... 
    Flexible hours

    Institute of Foundation Models

    Sunnyvale, CA
    16 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior System Architect, Infrastructure Reliability. Be the first to apply!