Principal Cluster Reliability Architect
Advanced Micro Devices Inc
WHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. THE ROLEAMD is seeking a Principal Cluster Reliability Architect to define and drive the reliability strategy for next-generation AI and HPC cluster platforms. This role serves as the technical authority for end-to-end cluster reliability, spanning compute, networking, storage, and control plane architectures. The successful candidate will influence architecture, deployment, validation, and operational readiness across AMD reference designs and large-scale customer deployments.The Principal Cluster Reliability Architect will establish reliability frameworks, drive cross-functional alignment, and ensure that reliability considerations are embedded throughout the cluster lifecycle from architecture through production operations.THE PERSONYou are a highly experienced systems architect with deep expertise in distributed systems, large-scale infrastructure, and reliability engineering. You possess a strong understanding of AI/HPC environments and have demonstrated success designing and operating production-scale platforms.You excel at leading through influence, aligning diverse engineering teams around common objectives, and translating operational learnings into architectural improvements. You combine technical depth with strategic thinking and are passionate about delivering resilient, scalable, and operationally efficient infrastructure platforms.KEY RESPONSIBILITIESReliability Architecture & DesignDefine and implement a comprehensive cluster reliability architecture framework spanning the full infrastructure lifecycle.Establish reliability requirements for AMD reference architectures and customer deployment models.Drive architecture decisions that enhance fault tolerance, redundancy, failure isolation, and system recoverability.Influence hardware, firmware, software, and operational design decisions to improve system resiliency.Operational Readiness & Day-2 OperationsPartner with Site Reliability Engineering (SRE) and Platform Operations teams to define sustainable operational models.Establish standards for:Observability and telemetryHealth monitoring frameworksJob-aware failure detection and recovery mechanismsLifecycle management, including patching and upgradesCreate closed-loop feedback mechanisms that convert operational insights into architectural improvements.Validation & Reliability EngineeringDefine reliability KPIs, SLAs, SLOs, and deployment acceptance criteria.Develop and execute reliability validation strategies at production scale.Establish methodologies for:Failure injection testingChaos engineeringRecovery validationResiliency benchmarkingDrive data-driven reliability improvements through empirical testing and analysis.Lifecycle Integration & Process MaturityEnsure reliability consistency across:ArchitectureDeploymentCluster bring-upValidationOperationsIdentify and eliminate gaps in ownership, governance, handoff processes, and operational readiness.Drive repeatable engineering practices that improve scalability and deployment quality.Cross-Functional Technical LeadershipAct as AMD's subject matter expert for cluster reliability architecture.Lead alignment across Architecture, Deployment Engineering, Validation Engineering, Platform Operations, and Product organizations.Mentor senior engineers and architects.Represent reliability strategy in customer engagements and strategic programs.Influence product and platform roadmaps to improve reliability outcomes across AMD cluster initiatives.PREFERRED SKILLSTechnical ExpertiseLarge-scale cluster architecture across compute, networking, storage, and control plane systems.Distributed systems architecture, reliability engineering, and failure modeling.Linux infrastructure and systems software.Production-scale cloud, AI, HPC, or platform infrastructure.AI/HPC TechnologiesKubernetesSlurmROCmGPU cluster architecturesAI training and inference infrastructureNetworking & Platform TechnologiesRDMA networking technologies:InfiniBandRoCEHigh-performance cluster interconnectsDistributed storage architecturesReliability EngineeringSite Reliability Engineering (SRE) methodologiesObservability platforms and telemetry systemsIncident management and root cause analysisAutomation and infrastructure-as-code practicesChaos engineering and resilience testingLeadership & InfluenceDevelopment of reference architectures and platform standardsCross-functional technical leadershipExecutive-level communication and stakeholder managementExperience influencing technical direction across large organizations without direct authorityACADEMIC CREDENTIALSBachelor's degree in Computer Engineering, Computer Science, Electrical Engineering, or a related technical discipline.This role is not eligible for visa sponsorship.#LI-KW1Benefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.
$272k - $431.25k
...platforms. You will ensure the seamless scaling of models to clusters comprising hundreds of thousands of nodes.AI Communication Library... ...Library Co-Design: Partner with application developers to architect and implement specialized communication primitives. You will ensure...PrincipalFull time- ...your career. THE ROLE: AMD is looking for a Server System RAS Architect to join our Datacenter System Architecture and Engineering... ...of next generation system designs to continually improve the reliability, availability and serviceability of Data Center Solutions.Influence...Principal
- Job Overview:Arm is seeking an experienced SoC Reliability, Availability and Serviceability (RAS) Architect to drive the RAS strategy for our next-generation SoCs. In this pivotal role, you will collaborate closely with design, verification, manufacturing, and product...PrincipalWork at officeLocal area
- ...(DCGPU) Validation and Engineering team ensures the quality, reliability, security, and performance of AMD's industry-leading AI, ML,... ...readiness, and deliver exceptional customer quality. Validation Architects play a key role in shaping future products by driving design-...Principal
$192k - $267k
...practical experience. ~10+ years of experience as an enterprise architect or in a customer‑facing role. ~ Experience in cloud market... ...built in the cloud. Our products are developed for security, reliability and scalability, running the full stack from infrastructure to...PrincipalWork at officeRemote work$236k - $275k
...Principal Architect HybridSeattle preferred / Remote OK too Full-Time We're hiring a Principal Architect to reinvent how legal work... ...accuracy, observability, or speed, at the scale required to reliably serve thousands of people every year. We are looking for...PrincipalFull timeTemporary workWork at officeImmediate startRemote workShift work2 days per week- Job DescriptionWe are seeking a Principal Router Architect with 10 to 15 years of experience to define the backbone of next-generation Zonal and... .... This role focuses on architecting high-bandwidth, ultra-reliable Ethernet Routing Engines and PCIe subsystems that serve as...PrincipalLocal areaRemote workFlexible hoursNight shift
- ...of how clients manage their money by providing innovative and reliable technology products and services as a part of our ongoing commitment... ....Collaborate with business leaders, product teams, peer architects, enterprise architects, engineers, and external vendors to translate...PrincipalFull timeWork at office
- ...activity and maintain compliance with regulatory requirements.The Principal Architect, Stock Plan Technology serves as the senior technology... ...the architectural vision, modernization strategy, Site Reliability Engineering and target-state technology roadmap for the platform...PrincipalFull timeWork at office
- ...broken so they can fix it. This has been the case for over 25 years. The Alerting Platform & Workflows team focuses on providing a reliable, fast and intuitive experience for configuring, receiving, and responding to alerts.Our users—software developers—need to...PrincipalFull timeRemote workFlexible hours
- ...of how clients manage their money by providing innovative and reliable technology products and services as a part of our ongoing commitment... ...Investment Management Operations & Distribution Platform Architect to lead architecture strategy and platform modernization...PrincipalFull timeWork at office
$272k - $431.25k
Do you want to help drive the development of CPU technology for architectures used for artificial intelligence (AI), agentic workloads, deep learning (DL), high-performance computing (HPC), cloud service providers (CSP), gaming, virtual reality, and autonomous vehicles?...PrincipalFull time- ...diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. THE ROLE:As a Performance Architect in AMD’s Embedded Business Unit, you will define and shape the next generation of x86 SoCs through deep pre-silicon performance...PrincipalNight shift
- ...silicon environments to enable performance validation, debug, and optimization.THE PERSON: As a passionate and technically strong Principal SoC Performance Engineer, you will work on highly parallel SoC architectures, leveraging deep understanding of GPU compute, memory...Principal
$272k - $431.25k
We are now looking for a Principal Hardware SoC Architect for our Tegra team! Do you want to be part of the Artificial Intelligence Revolution? Would you like to work with world-class systems architects and deep learning experts to define the next generation SoC?NVIDIA...PrincipalFull timeWork experience placement- ...between simulation and real-world robotic systems, minimizing the reality gap through domain randomization and adaptation techniques.Architect and deploy scalable simulation infrastructure on cloud platforms, ensuring seamless parallel simulation runs, distributed training...Principal
- ...we shape the future of AI and beyond. Together, we advance your career. THE ROLE: We are looking for a driven and diligent power architect to join the SOC Architecture team in the AMD Computing and Graphics group. Power architects guide SoCs from early concept stages through...Principal
- ...place to grow your career! Are you a system-on-chip visionary, an architect who thrives at the frontier where hardware and machine... ...accelerating Cirrus Logic’s diversification and strategic growth.As a Principal SoC Architect (ML Accelerators), you will play a pivotal role...PrincipalFull time
$272k - $431.25k
...directly into production AI stacks. The team's charter expands into emerging domains including quantum computing interconnects.This Principal Architect role leads the research agenda and architectural direction for how NVIDIA’s AI systems communicate at scale—across GPUs, DPUs...PrincipalFull timeRemote work- Job Overview:At Arm, the High-Speed I/O Architect defines and designs innovative on-chip interconnect architectures-coherent and non-coherent-for scalable SoC platforms. You will work across markets including mobile, automotive, datacenter, networking, and IoT, contributing...PrincipalWork at officeLocal areaShift work
- ...Sr. Principal Engineer, Mechanical Architect for PC NotebooksFrom applied research to advanced engineering, the Engineering Technologist team has the expertise to shape ground-breaking products, material and processes. It's a fascinating field of work. We're involved...Principal
$239k - $278.75k
...will be part of a culture that values trust, accountability, and shared success where your work truly matters.Job SummaryAs a Principal Architect you will serve as a trusted executive advisor responsible for influencing our clients’ cybersecurity transformation...PrincipalFull timeRemote workVisa sponsorshipWork visa$163.67k - $272.85k
...you need at LPL Financial to shape your success while helping clients pursue their financial goals. Job Overview The Principal M&A Architect is responsible for leading, analyzing, and designing the Enterprise Data Architecture (EDA) for the Merger and Acquisition...Principal- ...to become the world’s leading integrated design practice. Our architects, engineers, interior designers, consultants, sustainability specialists... ...your place with Stantec.Your OpportunityTo be successful as a Principal Architectural Lead in the Health Sector, a collaborative...PrincipalFull timeContract workTemporary workPart timeCasual workWork at officeLocal areaFlexible hours
$163.4k - $272.3k
...multiple Release Trains to identify dependencies and define integration strategies. Provide guidance to engineers and technical architects in clarifying requirements, allocating responsibilities to subsystems, defining interfaces, and stablishing delivery automation...PrincipalLocal area- Job DescriptionWe are looking for a Principal NPU Hardware Architect with 10 to 15 years of experience to drive the architectural definition and hardware implementation of high-performance Neural Processing Units (NPUs) targeted for microcontrollers and microprocessors...PrincipalWork at officeLocal areaRemote workFlexible hoursNight shift2 days per week
$139.9k - $233.1k
Senior Architect, Enterprise Architecture - Enterprise Servicing Career Architecture Title: IT Principal Architect, Enterprise Architecture Position Summary The Senior Architect, Enterprise Servicing shapes the enterprise architecture for an end-to-end member/patient...PrincipalLocal areaWork from home- ...reference architecture and other supporting collateral related specifically to storage solutions in support of AMD-based AI and HPC clustered systems at scale.Our team operates across industry verticals as subject matter experts in the AI stack and across the cluster. We’...
$171.6k - $338.3k
Position Summary Reliability & Maintenance Digital Transformation Architect Senior Manager - Downstream/Chemicals Position Summary We are a team of strategic advisors, architects, and implementers who drive business transformations. Our diverse talent energizes...Local areaVisa sponsorship- Key Responsibilities Architect, design, and evolve a scalable, secure AI platform and reference architecture that enables rapid development... ...Experience 8+ years of experience as a Senior, Lead, or Principal Engineer/Architect Hands-on experience with AI and ML...PrincipalVisa sponsorshipWork visa
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Cluster Reliability Architect. Be the first to apply!


