Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior System Architect, Infrastructure Reliability

$184k - $287.5k

NVIDIA

Senior System Architect: Heterogeneous EDA Systems

NVIDIA is seeking a Senior System Architect to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects.

What You'll Be Doing:

  • Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure.
  • Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs.
  • Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters.
  • Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams.
  • Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs.

What We Need To See:

  • Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming.
  • Experience building automated RCA (Root Cause Analysis) pipelines for HPC or cloud-scale environments.
  • CPU Architecture Deep-Dive: Expert knowledge of x86/ARM node-level metrics: IPC (Instructions Per Cycle), cache contention, NUMA imbalance, and hardware interrupts.
  • Programming Proficiency: Strong C++ and Python skills, with the ability to build high-performance daemons that monitor system health without impacting workload performance.
  • Scale Experience: Familiarity with cluster resource managers (Slurm, LSF, or Kubernetes) and how they manage job lifecycle and signal propagation.

Ways To Stand Out From The Crowd:

  • Low-Level Diagnostics: Expert knowledge of the Linux kernel and its error-reporting interfaces (/dev/mcelog, dmesg, journald). Understand how the kernel handles hardware exceptions and memory faults.
  • GPU Infrastructure Proficiency: Deep experience with the NVIDIA DCGM (Data Center GPU Manager) and NVIDIA Management Library (NVML) for monitoring device health and capturing state-dumps.
  • Experience with tools doing non-intrusive monitoring of application health and syscall-level failure patterns.
  • Experience with checkpoint/restore technologies (like CRIU) and their application in long-running EDA flows.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until June 19, 2026.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Senior System Architect, Infrastructure Reliability in Austin, TX vacancy
  • $126.2k - $264.1k

     ...About Oracle Health Applications & Infrastructure Oracle Health Applications & Infrastructure...  .... About the Role As an Senior Principal Product Manager, you will own...  ...architecture, APIs, data flows, integrations, reliability, security, and scalability considerations... 
    Senior
    Temporary work
    Flexible hours

    Oracle

    Austin, TX
    4 days ago
  •  ...Senior Solutions Architect At Gallatin, we are rebuilding logistics infrastructure for the national security missions of the United States...  ...partners. We build AI systems that determine how logistics...  ...automate workflows, and improve reliability and scalability of deployed... 
    Senior
    Work at office
    Local area

    Gallatin AI, Inc.

    Austin, TX
    1 day ago
  •  ...Descriptions: Key Responsibilities: • Infrastructure & GitOpsK8s & Containerization: Design|...  ...via Flyway.SRE & Observability Reliability: Own end-to-end production availability...  ...incident response| deep-dive distributed system debugging| RCAs| and on-call rotations.... 
    Suggested

    Purple Drive

    Austin, TX
    7 days ago
  • $124k - $195.5k

    NVIDIA is seeking a Solutions Architect in Data Center Infrastructure to join our Infrastructure Specialists team...  ...centers including power/cooling systems, cabling and network provisioning and...  ...to validate the functionality and reliability of infrastructure components.... 
    Suggested
    Work experience placement
    Work at office
    Worldwide
    Flexible hours

    Nvidia Corporation

    Austin, TX
    2 days ago
  • $152k - $195k

     ...Senior Site Reliability Engineer Austin, TX (Hybrid) SecurityScorecard is the global leader in cybersecurity ratings, with...  ...the design and optimization of our Kubernetes-based infrastructure and CI/CD systems. You will also own the infrastructure behind our AI tooling... 
    Senior

    SecurityScorecard

    Austin, TX
    3 days ago
  •  ...contractors, every day, with reliable parts and coupled with...  ...our seamless gutter systems and offer the most...  ...is looking for a senior engineer responsible for...  ...and automating hybrid infrastructure using Scale Computing...  ...AWS Certified Solutions Architect). Experience supporting... 
    Senior
    Full time
    For contractors
    Monday to Friday
    Shift work
    Night shift
    Weekend work

    Senox

    Austin, TX
    1 day ago
  •  ...SUMMARY We are seeking an experienced Site Reliability Engineer to own and maintain the deployment of our cloud-based infrastructure to customer sites. In this role, you will...  ...and deploys models to real-time robotic systems. Providing a reliable framework will accelerate... 
    Senior
    Full time
    Local area

    Synthesia

    Austin, TX
    1 day ago
  •  ...Senior Site Reliability Engineer Austin, Texas, United States Who We...  ...The 2K SRE team owns the infrastructure behind every player connection...  ...live-service events push systems to their limits, and this...  ...network engineers, systems architects, and game studio developers... 
    Senior

    2K

    Austin, TX
    3 days ago
  •  ...the vision while completing key deliverables. The Senior Principal Solutions Architect provides hands-on advisors, using strong interpersonal...  ...understanding of modern technical architecture, including cloud infrastructure and applications Proficiency in data integration... 
    Senior
    Worldwide

    Dun & Bradstreet

    Austin, TX
    14 hours ago
  •  ...Senior Site Reliability Engineer Zello is a voice-first communication platform, powered by our...  ...who likes operating real production systems, doesn't get stage fright in incidents...  ...~7+ years in SRE, DevOps, platform, infrastructure, or database reliability roles, with... 
    Senior
    Permanent employment
    Local area
    Flexible hours

    Zello

    Austin, TX
    3 days ago
  • $158k - $210k

     ...with food, mining, and transport. Our systems are designed to understand, predict, and...  ...physical operations into something more reliable, more scalable, and more productive....  ...you'll do Work on a data intelligence infrastructure team, which is focused on gaining intelligence... 
    Senior
    Full time
    Temporary work
    Work at office
    Flexible hours

    ATOMS Careers page

    Austin, TX
    5 days ago
  • $200k - $210k

     ...sharp, self‑motivated, problem‑solving Senior Infrastructure Engineer to join our team. As a Senior...  ...infrastructure to make it more reliable, resilient, and secure. You'll bring depth...  .... You'll help us scale and harden the systems that underpin CompanyCam's products as... 
    Senior
    Hourly pay
    For contractors
    Work experience placement
    Remote work
    Night shift
    Weekend work

    CompanyCam

    Austin, TX
    1 day ago
  •  ...fabrication, and integration of electrical systems for data center power and cooling...  ...MV/HV power distribution, phased power infrastructure, and controls architecture. The ideal candidate...  ...and a passion for delivering scalable, reliable, and compliant infrastructure solutions... 
    Senior
    Contract work
    Remote work

    Jabil Circuit, Inc.

    Round Rock, TX
    6 days ago
  • $158k - $210k

     ...Senior Software Engineer - Data Infrastructure Austin, TX About the Company Atoms is building the machines that power...  ...food, mining, and transport. Our systems are designed to understand,...  ...physical operations into something more reliable, more scalable, and more productive... 
    Senior
    Full time
    Temporary work
    Work at office
    Worldwide
    Flexible hours

    ATOMS Careers page

    Austin, TX
    1 day ago
  • $121.4k - $218.6k

     ...generation dedicated AI hardware infrastructure. You will be responsible for...  ...best-in-class uptime and reliability of our AI hardware...  ...scalability, and performance of our systems. You'll define key performance...  ...they are breached. As a Senior Site Reliability Engineer,... 
    Senior
    Work experience placement
    Work at office

    Akamai

    Austin, TX
    5 days ago
  • $110.7k - $171.8k

     ...platform components, including: Cloud infrastructure primitives Kubernetes clusters and...  ...in on-call rotation as a platform reliability escalation point Incident response,...  ...troubleshooting skills for distributed systems, including root-cause analysis and reliability... 
    Senior
    Work experience placement
    Work at office
    Local area

    Visa

    Austin, TX
    1 day ago
  • $127k - $249k

     ...is responsible for a range of critical infrastructure and operational functions that support...  ...mesh), and observability and alerting systems. The Fleet Management team provides the...  ...components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager... 
    Senior
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Austin, TX
    2 days ago
  •  ...you. Key Responsibilities Cloud Infrastructure Engineering Design, build, and support...  ...and self-service capabilities. Reliability, Monitoring & Incident Response...  ...activities. Continuously improving system reliability and operational processes.... 
    Senior
    Permanent employment
    Temporary work
    Work at office
    Flexible hours

    Corient Capital Partners

    Austin, TX
    2 days ago
  •  ...global engineering force has the most reliable and efficient build environments possible...  ...an experienced and highly skilled Senior Infrastructure Engineer to support and operate a modern...  ...platforms, endpoint management systems, cloud infrastructure, and security technologies... 
    Senior
    Work at office
    Remote work
    Worldwide

    Actian

    Round Rock, TX
    2 days ago
  •  ...Senior Machine Learning Engineer We are seeking a Senior Machine...  ...production ready AI systems for secure and distributed environments...  ...scalable, efficient, and reliable production systems that...  ...hardware from government cloud infrastructure to edge devices in... 
    Senior
    Live out
    Work at office
    Flexible hours

    Webai

    Austin, TX
    5 days ago
  • Seekr is building the infrastructure that powers the next generation...  ...enterprise AI. As a Senior AI Infrastructure...  ...across distributed systems, Kubernetes, GPU infrastructure...  ...scalable, and highly reliable systems capable of...  ...edge environments. Architect and optimize high‑... 
    Senior
    Work experience placement
    Flexible hours

    Seekr

    Austin, TX
    7 days ago
  • $120k - $150k

     ...Capital adopts a unique approach to digital infrastructure investment. Leveraging experience and...  ...Responsibilities Own availability and reliability analysis for BTM power solutions across...  ...for multi-technology BTM power systems (e.g., gas engines, turbines, fuel cells... 
    Senior
    Work at office
    Flexible hours

    Tract Capital Management

    Austin, TX
    1 day ago
  • $120k - $150k

    Senior Technology Architect Hiring Department: Applied Research Laboratories Position Open To: All...  ...Purpose We are seeking a talented systems specialist, with deep familiarity and...  ...implementation and maintenance of central IT infrastructure and applications. Responsibilities... 
    Senior
    Work at office
    Immediate start
    Weekend work
    Afternoon shift

    The University of Texas at Austin

    Austin, TX
    2 days ago
  •  ...just deploying AI—they’re building systems that remain reliable, adaptable, and ready to scale in an...  ...dynamic environments. The Role As a Senior Machine Learning Engineer at Striveworks...  ...(e.g., Docker, Kubernetes [k8s], infrastructure as code, major cloud architectures)... 
    Senior
    Work at office
    Remote work

    Strive Works

    Austin, TX
    1 day ago
  • $138k - $208k

     ...Visits, March 2025) Day to Day As a Senior Machine Learning Engineer on our...  ...new agentic experiences, and LLMOps reliability and infrastructure. Responsibilities Autonomously...  ...and building recommendation / ranking systems Develop LLM and machine learning model... 
    Senior
    Work experience placement
    Local area

    Indeed

    Austin, TX
    3 days ago
  •  ...Senior Machine Learning Engineer Hybrid At Cloudflare...  ...large-scale data systems, own the company's data lake, ingestion infrastructure, and platform tooling,...  ...datasets into fast, reliable, business-critical...  ...will be the principal architect behind the next generation... 
    Senior
    Local area

    Cloudflare Inc

    Austin, TX
    4 days ago
  • $111.6k - $186k

     .... Job Description Senior Site Reliability Engineer Department...  ..., scalable, and resilient systems. In this role you will...  ...improvements across our production infrastructure while mentoring engineers...  ...planning ~ Architect and operate container... 
    Senior
    Remote work
    Relocation
    Flexible hours
    Shift work

    Cox Communications

    Austin, TX
    4 days ago
  • $164k - $205k

    Join to apply for the Senior Site Reliability Engineer role at BetterUp Let’s face it, a company whose mission is human transformation...  ...monitor, troubleshoot, and maintain production systems Build and operate cloud infrastructure on AWS, using Terraform to codify and version‑... 
    Senior
    Work experience placement
    Summer holiday
    Work at office
    Local area
    Flexible hours
    Shift work
    2 days per week

    BetterUp

    Austin, TX
    2 days ago
  •  ...If you are a Cloud Platform Infrastructure Engineer professional looking for an opportunity...  ...Build, optimize, and manage cloud-native systems such as Kubernetes clusters and cloud resources...  ...productivity, energy security and reliability. With global operations and a... 
    Senior
    Temporary work
    Flexible hours

    Emerson

    Round Rock, TX
    5 days ago
  • $81.1k - $187k

    Overview We are looking for a Site Reliability Engineer 3 to support...  ...work closely with development, infrastructure, security, and operations...  ...proactive steps to design and architect infrastructure and/or...  ...to capacity needs, ensuring systems can handle current and future... 
    Senior
    Temporary work
    Flexible hours
    Shift work

    Oracle

    Austin, TX
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior System Architect, Infrastructure Reliability. Be the first to apply!