Senior System Architect, Infrastructure Reliability
$184k - $287.5kNVIDIA
Senior System Architect: Heterogeneous EDA Systems
NVIDIA is seeking a Senior System Architect to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects.
What You'll Be Doing:
- Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure.
- Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs.
- Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters.
- Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams.
- Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs.
What We Need To See:
- Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming.
- Experience building automated RCA (Root Cause Analysis) pipelines for HPC or cloud-scale environments.
- CPU Architecture Deep-Dive: Expert knowledge of x86/ARM node-level metrics: IPC (Instructions Per Cycle), cache contention, NUMA imbalance, and hardware interrupts.
- Programming Proficiency: Strong C++ and Python skills, with the ability to build high-performance daemons that monitor system health without impacting workload performance.
- Scale Experience: Familiarity with cluster resource managers (Slurm, LSF, or Kubernetes) and how they manage job lifecycle and signal propagation.
Ways To Stand Out From The Crowd:
- Low-Level Diagnostics: Expert knowledge of the Linux kernel and its error-reporting interfaces (/dev/mcelog, dmesg, journald). Understand how the kernel handles hardware exceptions and memory faults.
- GPU Infrastructure Proficiency: Deep experience with the NVIDIA DCGM (Data Center GPU Manager) and NVIDIA Management Library (NVML) for monitoring device health and capturing state-dumps.
- Experience with tools doing non-intrusive monitoring of application health and syscall-level failure patterns.
- Experience with checkpoint/restore technologies (like CRIU) and their application in long-running EDA flows.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until June 19, 2026.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
$126.2k - $264.1k
...About Oracle Health Applications & Infrastructure Oracle Health Applications & Infrastructure... .... About the Role As an Senior Principal Product Manager, you will own... ...architecture, APIs, data flows, integrations, reliability, security, and scalability considerations...SeniorTemporary workFlexible hours- ...Senior Solutions Architect At Gallatin, we are rebuilding logistics infrastructure for the national security missions of the United States... ...partners. We build AI systems that determine how logistics... ...automate workflows, and improve reliability and scalability of deployed...SeniorWork at officeLocal area
- ...Descriptions: Key Responsibilities: • Infrastructure & GitOpsK8s & Containerization: Design|... ...via Flyway.SRE & Observability Reliability: Own end-to-end production availability... ...incident response| deep-dive distributed system debugging| RCAs| and on-call rotations....Suggested
$124k - $195.5k
NVIDIA is seeking a Solutions Architect in Data Center Infrastructure to join our Infrastructure Specialists team... ...centers including power/cooling systems, cabling and network provisioning and... ...to validate the functionality and reliability of infrastructure components....SuggestedWork experience placementWork at officeWorldwideFlexible hours$152k - $195k
...Senior Site Reliability Engineer Austin, TX (Hybrid) SecurityScorecard is the global leader in cybersecurity ratings, with... ...the design and optimization of our Kubernetes-based infrastructure and CI/CD systems. You will also own the infrastructure behind our AI tooling...Senior- ...contractors, every day, with reliable parts and coupled with... ...our seamless gutter systems and offer the most... ...is looking for a senior engineer responsible for... ...and automating hybrid infrastructure using Scale Computing... ...AWS Certified Solutions Architect). Experience supporting...SeniorFull timeFor contractorsMonday to FridayShift workNight shiftWeekend work
- ...SUMMARY We are seeking an experienced Site Reliability Engineer to own and maintain the deployment of our cloud-based infrastructure to customer sites. In this role, you will... ...and deploys models to real-time robotic systems. Providing a reliable framework will accelerate...SeniorFull timeLocal area
- ...Senior Site Reliability Engineer Austin, Texas, United States Who We... ...The 2K SRE team owns the infrastructure behind every player connection... ...live-service events push systems to their limits, and this... ...network engineers, systems architects, and game studio developers...Senior
- ...the vision while completing key deliverables. The Senior Principal Solutions Architect provides hands-on advisors, using strong interpersonal... ...understanding of modern technical architecture, including cloud infrastructure and applications Proficiency in data integration...SeniorWorldwide
- ...Senior Site Reliability Engineer Zello is a voice-first communication platform, powered by our... ...who likes operating real production systems, doesn't get stage fright in incidents... ...~7+ years in SRE, DevOps, platform, infrastructure, or database reliability roles, with...SeniorPermanent employmentLocal areaFlexible hours
$158k - $210k
...with food, mining, and transport. Our systems are designed to understand, predict, and... ...physical operations into something more reliable, more scalable, and more productive.... ...you'll do Work on a data intelligence infrastructure team, which is focused on gaining intelligence...SeniorFull timeTemporary workWork at officeFlexible hours$200k - $210k
...sharp, self‑motivated, problem‑solving Senior Infrastructure Engineer to join our team. As a Senior... ...infrastructure to make it more reliable, resilient, and secure. You'll bring depth... .... You'll help us scale and harden the systems that underpin CompanyCam's products as...SeniorHourly payFor contractorsWork experience placementRemote workNight shiftWeekend work- ...fabrication, and integration of electrical systems for data center power and cooling... ...MV/HV power distribution, phased power infrastructure, and controls architecture. The ideal candidate... ...and a passion for delivering scalable, reliable, and compliant infrastructure solutions...SeniorContract workRemote work
$158k - $210k
...Senior Software Engineer - Data Infrastructure Austin, TX About the Company Atoms is building the machines that power... ...food, mining, and transport. Our systems are designed to understand,... ...physical operations into something more reliable, more scalable, and more productive...SeniorFull timeTemporary workWork at officeWorldwideFlexible hours$121.4k - $218.6k
...generation dedicated AI hardware infrastructure. You will be responsible for... ...best-in-class uptime and reliability of our AI hardware... ...scalability, and performance of our systems. You'll define key performance... ...they are breached. As a Senior Site Reliability Engineer,...SeniorWork experience placementWork at office$110.7k - $171.8k
...platform components, including: Cloud infrastructure primitives Kubernetes clusters and... ...in on-call rotation as a platform reliability escalation point Incident response,... ...troubleshooting skills for distributed systems, including root-cause analysis and reliability...SeniorWork experience placementWork at officeLocal area$127k - $249k
...is responsible for a range of critical infrastructure and operational functions that support... ...mesh), and observability and alerting systems. The Fleet Management team provides the... ...components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager...SeniorWork at officeLocal areaRemote workWorldwideFlexible hours- ...you. Key Responsibilities Cloud Infrastructure Engineering Design, build, and support... ...and self-service capabilities. Reliability, Monitoring & Incident Response... ...activities. Continuously improving system reliability and operational processes....SeniorPermanent employmentTemporary workWork at officeFlexible hours
- ...global engineering force has the most reliable and efficient build environments possible... ...an experienced and highly skilled Senior Infrastructure Engineer to support and operate a modern... ...platforms, endpoint management systems, cloud infrastructure, and security technologies...SeniorWork at officeRemote workWorldwide
- ...Senior Machine Learning Engineer We are seeking a Senior Machine... ...production ready AI systems for secure and distributed environments... ...scalable, efficient, and reliable production systems that... ...hardware from government cloud infrastructure to edge devices in...SeniorLive outWork at officeFlexible hours
- Seekr is building the infrastructure that powers the next generation... ...enterprise AI. As a Senior AI Infrastructure... ...across distributed systems, Kubernetes, GPU infrastructure... ...scalable, and highly reliable systems capable of... ...edge environments. Architect and optimize high‑...SeniorWork experience placementFlexible hours
$120k - $150k
...Capital adopts a unique approach to digital infrastructure investment. Leveraging experience and... ...Responsibilities Own availability and reliability analysis for BTM power solutions across... ...for multi-technology BTM power systems (e.g., gas engines, turbines, fuel cells...SeniorWork at officeFlexible hours$120k - $150k
Senior Technology Architect Hiring Department: Applied Research Laboratories Position Open To: All... ...Purpose We are seeking a talented systems specialist, with deep familiarity and... ...implementation and maintenance of central IT infrastructure and applications. Responsibilities...SeniorWork at officeImmediate startWeekend workAfternoon shift- ...just deploying AI—they’re building systems that remain reliable, adaptable, and ready to scale in an... ...dynamic environments. The Role As a Senior Machine Learning Engineer at Striveworks... ...(e.g., Docker, Kubernetes [k8s], infrastructure as code, major cloud architectures)...SeniorWork at officeRemote work
$138k - $208k
...Visits, March 2025) Day to Day As a Senior Machine Learning Engineer on our... ...new agentic experiences, and LLMOps reliability and infrastructure. Responsibilities Autonomously... ...and building recommendation / ranking systems Develop LLM and machine learning model...SeniorWork experience placementLocal area- ...Senior Machine Learning Engineer Hybrid At Cloudflare... ...large-scale data systems, own the company's data lake, ingestion infrastructure, and platform tooling,... ...datasets into fast, reliable, business-critical... ...will be the principal architect behind the next generation...SeniorLocal area
$111.6k - $186k
.... Job Description Senior Site Reliability Engineer Department... ..., scalable, and resilient systems. In this role you will... ...improvements across our production infrastructure while mentoring engineers... ...planning ~ Architect and operate container...SeniorRemote workRelocationFlexible hoursShift work$164k - $205k
Join to apply for the Senior Site Reliability Engineer role at BetterUp Let’s face it, a company whose mission is human transformation... ...monitor, troubleshoot, and maintain production systems Build and operate cloud infrastructure on AWS, using Terraform to codify and version‑...SeniorWork experience placementSummer holidayWork at officeLocal areaFlexible hoursShift work2 days per week- ...If you are a Cloud Platform Infrastructure Engineer professional looking for an opportunity... ...Build, optimize, and manage cloud-native systems such as Kubernetes clusters and cloud resources... ...productivity, energy security and reliability. With global operations and a...SeniorTemporary workFlexible hours
$81.1k - $187k
Overview We are looking for a Site Reliability Engineer 3 to support... ...work closely with development, infrastructure, security, and operations... ...proactive steps to design and architect infrastructure and/or... ...to capacity needs, ensuring systems can handle current and future...SeniorTemporary workFlexible hoursShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior System Architect, Infrastructure Reliability. Be the first to apply!
- servicenow technical architect Austin, TX
- technical architect Austin, TX
- pega system architect Austin, TX
- system architect Austin, TX
- senior vice president communications Austin, TX
- senior manager quality engineering Austin, TX
- senior device engineer Austin, TX
- sr operations manager Austin, TX
- senior supervisor Austin, TX
- senior client services manager Austin, TX

