Senior System Architect, Infrastructure Reliability
$184k - $287.5kNVIDIA
NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects.What you'll be doing:Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure.Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs.Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters.Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams.Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs.What we need to see:Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming.Experience building automated RCA (Root Cause Analysis) pipelines for HPC or cloud-scale environments.CPU Architecture Deep-Dive: Expert knowledge of x86/ARM node-level metrics: IPC (Instructions Per Cycle), cache contention, NUMA imbalance, and hardware interrupts.Programming Proficiency: Strong C++ and Python skills, with the ability to build high-performance daemons that monitor system health without impacting workload performance.Scale Experience: Familiarity with cluster resource managers (Slurm, LSF, or Kubernetes) and how they manage job lifecycle and signal propagation.Ways To Stand Out From The Crowd:Low-Level Diagnostics: Expert knowledge of the Linux kernel and its error-reporting interfaces (/dev/mcelog, dmesg, journald). Understand how the kernel handles hardware exceptions and memory faults.GPU Infrastructure Proficiency: Deep experience with the NVIDIA DCGM (Data Center GPU Manager) and NVIDIA Management Library (NVML) for monitoring device health and capturing state-dumps.Experience with tools doing non-intrusive monitoring of application health and syscall-level failure patterns.Experience with checkpoint/restore technologies (like CRIU) and their application in long-running EDA flows.#LI-Hybrid Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until September 2, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, MA, Westford; US, TX, Austin; US, NC, Durham; US, WA, RedmondType: Full time
$208k - $327.75k
Senior System Architect NVIDIA Enterprise Platforms Group is seeking a Senior System Architect to define, design, and validate enterprise... ...system architecture, customer requirements, and hands-on infrastructure validation, helping turn NVIDIA accelerated computing, networking...Senior$136k - $218.5k
...lasting impact on the world.NVIDIA is seeking a Senior ASIC Compose Verification Infrastructure and Tools Engineer to advance the systems that qualify, integrate, and deliver... ...sophisticated engineering systems faster, more reliable, and easier to reason about. You’ll also...SeniorFull time$224k - $356.5k
...next generation architecture into its EDA Infrastructure. We expect you to have a deep... ...product introductions (NPIs), distributed systems, familiarity with software testing and... ...problems and write effective, clear and reliable architecture specificationTranslate requirements...SeniorFull time$152k - $241.5k
...and management of tooling and release infrastructure for chip designers. We are constantly evolving... ...compatibility and keeping tools reliable and scalable.What You'll be Doing:Research... ...metrics of the build and deployment systems.Research and adapt the latest CI/CD practices...SeniorFull time$184k - $287.5k
NVIDIA is looking for a Systems & Software Engineer interested in building and running reliable large scale infrastructure platform services. In this role you will create and support automation that manages infrastructure, as part of our organization running EDA platforms...SeniorFull timeRemote work$152k - $241.5k
Sr Software Engineer - Distributed Systems Engineer, EDA InfrastructureNVIDIA is hiring engineers to build and scale the infrastructure that supports our Electronic Design Automation... ...abilities. You will help design reliable automation and platform services that manage...SeniorFull timeRemote work- Applied Research Solutions seeks a Senior Systems Engineer at Hanscom AFB, Bedford, MA, onsite. The role focuses on translating operational needs into technical solutions, developing architectures, and guiding risk assessment across the DoD acquisition lifecycle. Collaboration...Senior
£80k - £95k per year
Job Description This Senior Dev Ops Engineer role in Chelmsford involves managing and improving infrastructure to support technology systems in the Technology & Telecoms industry. You'll... ...processes to improve efficiency and reliability. \n Monitor system performance...SeniorPermanent employment- Hardware Infrastructure EDA Compute Team Member As a member of the Hardware Infrastructure EDA... ...scale, and support workload scheduling systems that directly impact design velocity... ...improvements in observability, service reliability, and automation, ensuring the EDA compute...Senior
$184k - $287.5k
...responsible for development and support of infrastructure tools used by design engineers for... ...architectural, RTL, and gate level designs. As a senior software engineer, you will craft... ...be Doing: Work as a team to build reliable, scalable and high performance software...SeniorFull timeWorldwide$184k - $287.5k
...the world. We are seeking a Senior Quantum/HPC Systems Engineer to help architect, deploy, and operate a first-of-... ...the intersection of data center infrastructure, high-performance computing, and... ...experimental quantum hardware into a reliable, high-impact research platform....SeniorWork at office$184k - $287.5k
We are seeking a Linux Systems Engineer to join our Infrastructure Engineering team, dedicated to supporting high-performance EDA compute farms. You will be responsible for designing, deploying, and maintaining enterprise-grade Linux environments across RPM-based distributions...SeniorFull timeWork experience placementRemote work$152k - $241.5k
Senior Full-stack Web Applications Software Architect NVIDIA has been transforming computer graphics... ...Architect to join our Hardware Infrastructure team! Our team is... ...and AI agents that are reliable, scalable, and... ...knowledge of distributed systems principles, concurrency...Senior$184k - $287.5k
NVIDIA is seeking elite ASIC Infrastructure engineers to deliver the tooling and environment that enables DV and RTL Design for the world's... ...GPU front end build flow Keep the GPU Continuous Integration system at the cutting edge of source management methodologies Guide compute...SeniorFull timeWork experience placementWork at officeWorldwide- ...Assurance (QA) Engineer Nokia's Network Infrastructure group is at the heart of a revolution... ...increasing demand for higher capacity, greater reliability, faster speeds and lower costs. Your... ...appraisals of programming languages, systems and computation software. Your Skills...SeniorTemporary workWorldwide
- Principal Network / Systems Architect Koniag IT Systems, LLC, a Koniag Government Services company, is seeking a Principal Network / Systems... ...& Duties may include, but are not limited to: Convergent Infrastructure Design: Architect integrated solutions where Software-...Local areaRemote workFlexible hours
- Berkshire Grey, Inc. in Massachusetts is seeking a Principal Solutions Architect to lead the design of intelligent automation systems that leverage our technology to optimize customer operations. You will analyze complex data and processes to ensure performance and scalability...Senior
- ...security frameworks, governance processes, and industry best practices. Support security architecture reviews for applications, infrastructure, and technology initiatives. Drive implementation and optimization of Microsoft Purview security and compliance solutions....Senior
$184k - $287.5k
AI Infrastructure Engineers at NVIDIA build the systems, tooling, and data infrastructure that enable operation of our GPU cloud services. We are enabling engineering teams to innovate while proactively identifying, tracking, and mitigating risks across the entire technical...SeniorFull timeRemote work- The Middlesex Corporation is a national leader in heavy civil construction, delivering infrastructure projects since 1972. The Estimator will price work, guide the bid team, and coordinate with subs and vendors for road, bridge, marine, and heavy/civil projects valued...SeniorPrice work
- ...trailer-unload computer vision problems for real-world robotic systems. Your work will help robots better perceive, reason about, and... ...conditions. These efforts will enhance the performance, reliability, and throughput of our robotic unloading solutions while unlocking...SeniorShift work
- ...seeking a qualified individual in Chelmsford, MA, to serve as a technical expert in wetland and waterways permitting for complex infrastructure projects. You will lead permitting efforts, prepare applications, and collaborate with multidisciplinary teams to ensure...Senior
- Senior Pega Engineer/Developer The Senior Pega Engineer/Developer... ...’s degree in Information Systems, Computer Science, software... ...DevOps principles, including infrastructure as code (IaC), continuous integration... ...high availability and reliability of Pega applications....Senior
- Mantu is seeking a Senior Embedded Linux Software Engineer in Westford, MA. This role focuses... ...for safety-critical fire detection systems, working with Yocto Linux-based... ...collaborating with global teams to ensure reliability and compliance with safety standards. #J...Senior
- ...Mechanical Engineer in Littleton, MA to lead fluid and thermal system architecture from concept to production. You will mentor engineers... ...activities, and cross‑functional execution to deliver reliable, scalable cooling solutions while advancing #J-18808-Ljbffr Flextronics...SeniorFlexible hours
- ...list through its consistent efforts to safely build America’s infrastructure. The Middlesex Corporation specializes in building and reconstructing... ...regional presence and reputation. Position Summary: The Senior Estimator is responsible for the accurate takeoff and pricing...SeniorFull timePart timeFor subcontractor
$115k - $125k
...Overview Senior Logistics Manager WORK LOCATION : Hanscom AFB,... ...the DoD. Proficiency with IT infrastructure Logistics Preferred Qualificaitons... ...link and communications systems Responsibilities The... ...metrics like supply levels, reliability, and maintainability, feeding...SeniorFull timeFor contractors$55 per hour
...Trident Consulting is seeking a "Senior Cloud DevOps Engineer" for one of our client... ...Responsibilities • Design and manage infrastructure using AWS, Kubernetes, and Terraform,... ...CI/CD pipelines to enable efficient and reliable deployments. • Automate...SeniorContract workRemote work- ...implementation, coordinating contributors across product, engineering, infrastructure, data, and user-facing technology. The environment requires... ...clearly with technical teams, business stakeholders, and senior leadership. Apply a product-oriented mindset to new AI and...SeniorRemote work
- ...JD: We are seeking a Senior Cloud DevOps Engineer to assist in the design, implementation... ...scalable, secure, and resilient cloud infrastructure. This role will drive our... ...performing team focused on rapid delivery, system reliability, and continuous improvement. Working...SeniorRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior System Architect, Infrastructure Reliability. Be the first to apply!
- senior consulting engineer Westford, MA
- senior medical science liaison Westford, MA
- senior software engineer remote Westford, MA
- senior mulesoft developer Westford, MA
- senior performance engineer Westford, MA
- senior manager m&a tax Westford, MA
- senior living Westford, MA
- senior performance tester Westford, MA
- senior international accountant Westford, MA
- senior sas developer Westford, MA


