System Failure Analysis Engineer (GPU Servers / Data Center)
Advanced Micro Devices Inc
WHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. THE ROLE:Analysis (FA) Engineer, you will play a critical role in diagnosing, isolating, and resolving complex failures across GPU-accelerated server platforms deployed in rack-level and data center environments, including ODM factory support. This is a highly technical, hands-on role focused on server bring-up, system-level debug, and rack integration troubleshooting across CPU, GPU, memory, PCIe, networking, power delivery, and thermal subsystems. You will leverage advanced electrical, firmware, and platform-level diagnostic tools, along with AI-assisted debug and failure analysis techniques, to accelerate root cause identification and improve system reliability at scale. This role requires deep experience with server and rack-level platforms, including BIOS/BMC behavior and data center operations—not just component-level failure analysis. You will collaborate closely with platform design, firmware, validation, manufacturing, and quality teams to improve product robustness, shorten debug cycles, and drive corrective actions across deployed infrastructure.THE PERSON:The ideal candidate is analytical, detail-oriented, and thrives in high-visibility debug environments. You are comfortable owning complex investigations spanning hardware, firmware, power sequencing, and system interoperability in multi-node GPU server systems.You bring strong experience in server bring-up and rack-level troubleshooting and can translate technical findings into clear RCA/FMEA documentation and actionable design improvements. You are equally effective working independently in lab environments and cross-functionally with engineering, manufacturing, and quality teams.You are also comfortable adopting AI-assisted engineering workflows to enhance debug efficiency, log analysis, and knowledge sharing.KEY RESPONSIBILITIES:Platform & System Failure AnalysisPerform component-, system-, and rack-level failure analysis on GPU-accelerated server platformsDebug issues across CPU, GPU, memory, PCIe, networking, storage, power, and thermal subsystemsSupport server bring-up and manufacturing test failures (POST, BIOS/UEFI configuration, PCIe enumeration, firmware interactions)Analyze BIOS, BMC, IPMI, and system logs to identify hardware/firmware interaction issuesReproduce factory or field failures in lab environments to validate root cause and corrective actionsUtilize system-level debug tools (oscilloscopes, logic analyzers, protocol analyzers, power tools)ODM Factory Enablement & ExecutionSupport ODM manufacturing test, failure debug, and troubleshootingLead issue triage and structured debug activities with ODM partnersTrain ODM teams on debug procedures and failure isolation techniquesDevelop and maintain SOPs, debug guides, and troubleshooting documentationManage factory escalations with clear diagnosis, recommendations, and communicationCross-Functional CollaborationPartner with design, firmware, validation, manufacturing, and quality teams to resolve issuesDrive RCA/FMEA documentation with clear problem statements and data-backed actionsProvide feedback to improve DfR, DfT, and DfSAI-Assisted Debug & Knowledge DevelopmentUse AI tools to: Analyze logs and failure patternsImprove debug efficiencyContribute to scalable debug knowledge basesPREFERRED EXPERIENCE:Hands-on experience debugging server platforms in rack-level or data center environmentsExperience with GPU servers and multi-node systemsStrong understanding of power sequencing, PCIe, memory, and thermal systemsFamiliarity with BIOS/UEFI, BMC, IPMI, and firmware-level debuggingProficiency with lab debug tools (oscilloscopes, logic analyzers, protocol analyzers, power analyzers)Experience using AI tools for log analysis, debug workflows, or knowledge developmentStrong cross-functional collaboration and communication skillsExcellent technical documentation skillsWillingness to travel up to 25%ACADEMIC CREDENTIALSBachelor’s or master's degree in Electrical Engineering requiredLOCATION: Austin, TX (Onsite)This role is not eligible for visa sponsorship.#LI-CS1Benefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.
- ...Graphics Processing Units (GPU’s). Our team plays a... ...the architecture, design, analysis, and verification of power systems intended for powering high... ...performance CPUs/GPUs/APUs. The engineer will work with cross-... ...architectures and PDN in server, OAM, and PCIe...Suggested
$173.9k - $235.2k
...experienced Senior Systems Development Engineer to lead the... ...infrastructure for our server platforms. You... ...implement predictive failure detection systems... ..., sensor data, error trending,... ...storage, compute, GPU, networking in production... ...root cause analysis on hardware failures...SuggestedInternshipLocal areaWorldwideFlexible hours- ...forward.THE ROLE: As a Failure Analysis Engineer in Product... ...issues in silicon and system-level... ...RESPONSIBILITIES Bring up Server Platforms,... ...Analyze and present Data and contribute to... ...Experience in CPU/GPU Architecture, debug... ...Experience with GPU data center infrastructure is...Suggested
$148.7k - $201.2k
...We are seeking a Systems Development Engineer to develop automation... ...(AI/ML) server platforms. You will... ...implement predictive failure detection systems... ..., sensor data, error trending,... ...across compute, GPU, and networking in... ...Perform root cause analysis on hardware failures...SuggestedInternshipLocal areaWorldwideFlexible hours- ...experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of... ...a Lead Systems Debug Engineer to provide technical leadership... ...-generation datacenter GPU platforms. This role... ...while leading root-cause analysis and resolution efforts....Suggested
$183k - $247.6k
...opportunity to build the systems that define... ..., and network engineers, supply chain... ...performance server and/or accelerator... ...complex system failures in time... ...integrity, failure analysis, server components (e.g. CPU, GPU, SSDs, memory),... ...market- 5+ years of data center engineering or...Local areaFlexible hours- ...Senior Systems Engineer Austin, Texas, United States... ...platforms across lab and data center environments. This... ...platforms, including server blades, racks, and rack... ...system-level failures involving thermal behavior... ...to perform root cause analysis and propose corrective...Flexible hours
- ...experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a... ...a Senior Hardware Engineer to join the Data Center... ...Data Center GPU systems, working in... ...(DFT), reliability analysis, and component derating... ...understanding of AI server and rack-scale...
$110.5k - $160k
...with development centers in the U.S. and Israel... ...across silicon engineering, hardware design,... ...-to-chip), inter-system connections, and... ...boards and servers at scale, ensuring... ...deployed in AWS data centers. You'll collaborate... ...through silicon failure analysis and debugA day in...Flexible hours$143.7k - $194.4k
...with development centers in the U.S. and Israel... ...across silicon engineering, hardware design,... ...-to-chip), inter-system connections, and... ...boards and servers at scale, ensuring... ...deployed in AWS data centers. You'll collaborate... ...through silicon failure analysis and debugA day in...InternshipFlexible hours- ...experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a... ...Center Platform Engineering Group (DPEG) Thermal... ...and data center GPU platforms. Engineers... ...modeling, design analysis, and technology... ...computing, data center, server, or AI hardware...Internship
$183k - $247.6k
...a Senior Reliability Engineer you will engage with an... ...understanding of computer systems to influence design... ...prediction of failure mechanisms, products under... ...expectations.* Lead failure analysis for component quality... ...custom silicon and servers including the Nitro, Graviton...Work experience placementLocal areaFlexible hours- ...join our talented Team. Job Title: System Engineer Datacenter GPU Location(s): Austin, TX Client is... ...automated jobs per day on thousands of servers helping with the productivity of... ...design creative solutions, mine through data to uncover real problems and fix them...Worldwide
- ...experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a... ...execution, root-cause analysis, and deployment readiness... ..., and customer engineering teams.Influence future... ...EXPERIENCEStrong background in server, AI accelerator, GPU, or datacenter...
- ...experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a... ...Product Engineer in the Data Center... ...for AMD Data Center GPU products. This is... ...Experience working with server, rack, and data center... ...field-relevant failure analysis. The position is ideally...Contract workLocal area
- ...forward.THE ROLE:The Failure Analysis (FA) Planner and On-site Engineer is part of the... ...and L1 (Level 1) FA center. This role... ...knowledge of PCBA and GPU product diagnostics... ...failure analysis data, test logs, inspection... .... Experience with server systems, GPU-based...Contract work
- ...experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of... ...is a hands-on role for engineers who thrive on exploration... ..., or performance analysis Comparative analysis: experience... ...Kubernetes for HPC/AI (GPU operators, device...
$136.6k - $184.8k
..., and next-generation data center design can be merged?... ...knowledge of water treatment systems while providing input... ...with design engineering, construction, operations... ...such as water supply analysis and design for buildings... ...centers and all of the servers, storage, networking,...For contractorsLocal areaFlexible hours$145k - $187k
..., solutions, and data management our customers... ...work gets done. Engineers define intent,... ...solving complex systems challenges,... ...Computing & Data Center Knowledge: Strong... ...AI/HPC platforms, GPU servers, high-density compute... ....Reliability, Failure Analysis & Issue Resolution...$183k - $247.6k
...AI accelerator compute systems? Lead bleeding-edge HW... ...customers in our own Data Centers We are seeking experienced... ...Lead System Design Engineers to build the next generation of our cloud server infrastructure,... ...utilization.- Drive root cause analysis for hardware issues...Local areaFlexible hours- ...Summary The Staff Engineer, System Integration Test... ...complex, high-end servers and hyperscale... ...in storage/server, GPU and networking products... ..., execution, and analysis of test cycles.... ...and other detailed data. Occasional travel... ...critical data center infrastructure for...Work at officeLocal areaWorldwideShift work
- ...Responsibility: AMD, Inc. is hiring SMTS Systems Design Engineer to Research, design, develop, and... ...problems of moderate scope where analysis of situations or data requires a review of a variety of... ...the following: Enterprise/data center server systems foundations; Server...
- ...experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of... ...hiring Systems Design Engineers to research, design, develop... ...of moderate scope where analysis of situations or data... ...systems development; 7. Server systems and OS development...Internship
- ...experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of... ...a memory systems design engineer within the client and... ...level debug and root cause analysis to narrow down the... ...hardware architecture (GPU, CPU/APU, memory and bus...
$110.1k - $151.4k
...more sustainable data centers. Together, we’re not... ...to add a Sales Engineer located in Littleton... ...and other global system manufacturers. This... ...integrate into their server platforms and... ...quotes, configuration analysis, RFP/RFQ responses... ...(Dell PowerEdge, GPU platforms, AI...Full timeVisa sponsorshipFlexible hours$183k - $247.6k
...Mechanical Thermal Engineering team, you'll own... ...and interconnect systems Define thermal... ...concept, design, analysis, prototyping, validation... ...cold plates, and data center liquid cooling... ...and mechanical failures in production, implementing... ...on SoCs, Servers, or related...Local areaFlexible hours- ...experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of... ...As part of the AMD EPYC Server team, we are committed... ...team is looking for system engineering leaders who are... ...plans, and drive root cause analysis and resolution.Manages technical...
$173.9k - $235.2k
...your opportunity to build the systems that define what’s next for... ...join a diverse AWS Hardware Engineering team of software, hardware,... ...- vertically from baremetal server hardware up to the software... ...systems development in an IT or data center environment experience- 3+...InternshipLocal areaFlexible hours$136.6k - $184.8k
...We support all AWS data centers and all of the servers, storage, networking... ...team of design engineers, quality/reliability... ...resilient quality systems at our vendors and... ...Review historical failures, equipment design changes... ...(s) root cause analysis, working with suppliers...Flexible hours- ...and thermal management solutions for AI data centers and other mission-critical... ...growth, we’re looking to add a Principal Engineer, Systems Architecture Engineering located in Austin... ...architectures, models, specifications, and analysis yourself rather than only directing...Full timeVisa sponsorshipFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to System Failure Analysis Engineer (GPU Servers / Data Center). Be the first to apply!
- system engineer remote Austin, TX
- senior windows systems engineer Austin, TX
- systems engineer intern Austin, TX
- senior linux systems engineer Austin, TX
- ground systems engineer Austin, TX
- advanced systems engineer Austin, TX
- wireless systems engineer Austin, TX
- senior staff systems engineer Austin, TX
- mission system engineer Austin, TX
- sr systems engineer Austin, TX

