System Failure Analysis Engineer (GPU Servers / Data Center)
Advanced Micro Devices Inc
WHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. THE ROLE:Analysis (FA) Engineer, you will play a critical role in diagnosing, isolating, and resolving complex failures across GPU-accelerated server platforms deployed in rack-level and data center environments, including ODM factory support. This is a highly technical, hands-on role focused on server bring-up, system-level debug, and rack integration troubleshooting across CPU, GPU, memory, PCIe, networking, power delivery, and thermal subsystems. You will leverage advanced electrical, firmware, and platform-level diagnostic tools, along with AI-assisted debug and failure analysis techniques, to accelerate root cause identification and improve system reliability at scale. This role requires deep experience with server and rack-level platforms, including BIOS/BMC behavior and data center operations—not just component-level failure analysis. You will collaborate closely with platform design, firmware, validation, manufacturing, and quality teams to improve product robustness, shorten debug cycles, and drive corrective actions across deployed infrastructure.THE PERSON:The ideal candidate is analytical, detail-oriented, and thrives in high-visibility debug environments. You are comfortable owning complex investigations spanning hardware, firmware, power sequencing, and system interoperability in multi-node GPU server systems.You bring strong experience in server bring-up and rack-level troubleshooting and can translate technical findings into clear RCA/FMEA documentation and actionable design improvements. You are equally effective working independently in lab environments and cross-functionally with engineering, manufacturing, and quality teams.You are also comfortable adopting AI-assisted engineering workflows to enhance debug efficiency, log analysis, and knowledge sharing.KEY RESPONSIBILITIES:Platform & System Failure AnalysisPerform component-, system-, and rack-level failure analysis on GPU-accelerated server platformsDebug issues across CPU, GPU, memory, PCIe, networking, storage, power, and thermal subsystemsSupport server bring-up and manufacturing test failures (POST, BIOS/UEFI configuration, PCIe enumeration, firmware interactions)Analyze BIOS, BMC, IPMI, and system logs to identify hardware/firmware interaction issuesReproduce factory or field failures in lab environments to validate root cause and corrective actionsUtilize system-level debug tools (oscilloscopes, logic analyzers, protocol analyzers, power tools)ODM Factory Enablement & ExecutionSupport ODM manufacturing test, failure debug, and troubleshootingLead issue triage and structured debug activities with ODM partnersTrain ODM teams on debug procedures and failure isolation techniquesDevelop and maintain SOPs, debug guides, and troubleshooting documentationManage factory escalations with clear diagnosis, recommendations, and communicationCross-Functional CollaborationPartner with design, firmware, validation, manufacturing, and quality teams to resolve issuesDrive RCA/FMEA documentation with clear problem statements and data-backed actionsProvide feedback to improve DfR, DfT, and DfSAI-Assisted Debug & Knowledge DevelopmentUse AI tools to: Analyze logs and failure patternsImprove debug efficiencyContribute to scalable debug knowledge basesPREFERRED EXPERIENCE:Hands-on experience debugging server platforms in rack-level or data center environmentsExperience with GPU servers and multi-node systemsStrong understanding of power sequencing, PCIe, memory, and thermal systemsFamiliarity with BIOS/UEFI, BMC, IPMI, and firmware-level debuggingProficiency with lab debug tools (oscilloscopes, logic analyzers, protocol analyzers, power analyzers)Experience using AI tools for log analysis, debug workflows, or knowledge developmentStrong cross-functional collaboration and communication skillsExcellent technical documentation skillsWillingness to travel up to 25%ACADEMIC CREDENTIALSBachelor’s or master's degree in Electrical Engineering requiredLOCATION: Austin, TX (Onsite)This role is not eligible for visa sponsorship.#LI-CS1Benefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.
$173.9k - $235.2k
...experienced Senior Systems Development Engineer to lead the... ...infrastructure for our server platforms. You... ...implement predictive failure detection systems... ..., sensor data, error trending,... ...storage, compute, GPU, networking in production... ...root cause analysis on hardware failures...SuggestedInternshipLocal areaWorldwideFlexible hours- ...forward.THE ROLE: As a Failure Analysis Engineer in Product... ...issues in silicon and system-level... ...RESPONSIBILITIES Bring up Server Platforms,... ...Analyze and present Data and contribute to... ...Experience in CPU/GPU Architecture, debug... ...Experience with GPU data center infrastructure is...Suggested
- ...experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of... ...a Lead Systems Debug Engineer to provide technical leadership... ...-generation datacenter GPU platforms. This role... ...while leading root-cause analysis and resolution efforts....Suggested
$183k - $247.6k
...opportunity to build the systems that define... ..., and network engineers, supply chain... ...performance server and/or accelerator... ...complex system failures in time... ...integrity, failure analysis, server components (e.g. CPU, GPU, SSDs, memory),... ...market- 5+ years of data center engineering or...SuggestedLocal areaFlexible hours- ...experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a... ...a Senior Hardware Engineer to join the Data Center... ...Data Center GPU systems, working in... ...(DFT), reliability analysis, and component derating... ...understanding of AI server and rack-scale...Suggested
$143.7k - $194.4k
...with development centers in the U.S. and Israel... ...across silicon engineering, hardware design,... ...-to-chip), inter-system connections, and... ...boards and servers at scale, ensuring... ...deployed in AWS data centers. You'll collaborate... ...through silicon failure analysis and debugA day in...InternshipFlexible hours$110.5k - $160k
...with development centers in the U.S. and Israel... ...across silicon engineering, hardware design,... ...-to-chip), inter-system connections, and... ...boards and servers at scale, ensuring... ...deployed in AWS data centers. You'll collaborate... ...through silicon failure analysis and debugA day in...Flexible hours- ...experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a... ...Center Platform Engineering Group (DPEG) Thermal... ...and data center GPU platforms. Engineers... ...modeling, design analysis, and technology... ...computing, data center, server, or AI hardware...Internship
- ...join our talented Team. Job Title: System Engineer Datacenter GPU Location(s): Austin, TX Client is... ...automated jobs per day on thousands of servers helping with the productivity of... ...design creative solutions, mine through data to uncover real problems and fix them...Worldwide
- ...growth, we’re looking to add a System Integration Engineer located in Austin TX.... ...Integration Engineer will define data center solutions with multiple... ...customer requirements, analysis and validation of customer... ...of new product, failure analysis, and working closely...Full timeLocal areaFlexible hours
$100k - $147k
...We’re looking for a Senior Systems Engineer who brings deep infrastructure... ...you’ll own and advance our data center, Cloud, and Microsoft 365 environments... ..., with a strong focus on server infrastructure, identity,... ...through proactive analysis and innovation Collaborate and...Remote work- ...experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a... ...Product Engineer in the Data Center... ...for AMD Data Center GPU products. This is... ...Experience working with server, rack, and data center... ...field-relevant failure analysis. The position is ideally...Contract workLocal area
- ...forward.THE ROLE:The Failure Analysis (FA) Planner and On-site Engineer is part of the... ...and L1 (Level 1) FA center. This role... ...knowledge of PCBA and GPU product diagnostics... ...failure analysis data, test logs, inspection... .... Experience with server systems, GPU-based...Contract work
$136.6k - $184.8k
..., and next-generation data center design can be merged?... ...knowledge of water treatment systems while providing input... ...with design engineering, construction, operations... ...such as water supply analysis and design for buildings... ...centers and all of the servers, storage, networking,...For contractorsLocal areaFlexible hours- ...experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of... ...is a hands-on role for engineers who thrive on exploration... ..., or performance analysis Comparative analysis: experience... ...Kubernetes for HPC/AI (GPU operators, device...
$183k - $247.6k
...AI accelerator compute systems? Lead bleeding-edge HW... ...customers in our own Data Centers We are seeking experienced... ...Lead System Design Engineers to build the next generation of our cloud server infrastructure,... ...utilization.- Drive root cause analysis for hardware issues...Local areaFlexible hours- ...responsible for System and Silicon validation... ...of AMD EPYC Server & AMD Instinct products... ...for system level failures working with engineering teams across AMD.... ...spanning CPU, GPU, memory, BIOS,... ...investigation and root cause analysis of complex... ...workflows, and data-driven validation...
- ...experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of... ...a memory systems design engineer within the client and... ...level debug and root cause analysis to narrow down the... ...hardware architecture (GPU, CPU/APU, memory and bus...
$183k - $247.6k
...Mechanical Thermal Engineering team, you'll own... ...and interconnect systems Define thermal... ...concept, design, analysis, prototyping, validation... ...cold plates, and data center liquid cooling... ...and mechanical failures in production, implementing... ...on SoCs, Servers, or related...Local areaFlexible hours$110.1k - $151.4k
...more sustainable data centers. Together, we’re not... ...to add a Sales Engineer located in Littleton... ...and other global system manufacturers. This... ...integrate into their server platforms and... ...quotes, configuration analysis, RFP/RFQ responses... ...(Dell PowerEdge, GPU platforms, AI...Full timeVisa sponsorshipFlexible hours- ...experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of... ...As part of the AMD EPYC Server team, we are committed... ...team is looking for system engineering leaders who are... ...plans, and drive root cause analysis and resolution.Manages technical...
$136k - $184k
...our fleet of ML servers deployed around... ...are seeking an engineer who is... ...emergent problems in GPU and server hardware... ..., developing data infrastructure... ...responsible for system remediation, operational... ...cause hardware failures and identify... ...and systems analysis to identify and...Permanent employmentInternshipFlexible hours$173.9k - $235.2k
...your opportunity to build the systems that define what’s next for... ...join a diverse AWS Hardware Engineering team of software, hardware,... ...- vertically from baremetal server hardware up to the software... ...systems development in an IT or data center environment experience- 3+...InternshipLocal areaFlexible hours- ...enterprise client seeking a Systems Engineer 5 in Austin, TX. Responsibilities... ...via automation Design server monitoring and management solutions... ...at their root, looking for failure patterns amenable to long-... ...management and problem analysis skills Highly Desirable Attributes...Hourly payContract work
- ...experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a... ...Design Mechanical Engineer, you'll design and... ...from cutting-edge AI servers and high-density infrastructure... ...server and GPU products used in... ...in Creo, tolerance analysis, and system-level...Worldwide
$144k - $209k
Lead analysis of system hardware designs to enable proactive... ...integration sites to field (data centers) that help predict... ...and mitigate risk of failure early during new... ...Industrial or Mechanical Engineering, or equivalent... ...maintain hardware (like servers and its components) and...Contract workWorldwide- ...computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of... ...dynamic, energetic Systems Design Engineer to join our growing team. As a key... ...including delivery, sequencing, analysis, and optimization Knowledge of system...
$104.5k - $160k
...Stores Tech (WWGST) team is hiring a Systems Administrator/Engineer to join our POS Infrastructure (... ...for store devices (POS, SCO, servers, peripherals).- Perform strong root cause analysis on production incidents and device failures, identifying systemic issues and driving...Permanent employmentWork experience placementRemote workWorldwideFlexible hoursNight shift$157.25k - $203.5k
Principal Systems Development EngineerHelp architect and deliver... ...Principal Systems Development Engineer, you'll own the system-level... ...factories, and customer data centers, you'll partner with engineering... ...benchmarking, testing, and analysis across all phases of...$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to... ...ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer... ...tooling, and agentic optimization systems that improve GPU kernels at the assembly...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to System Failure Analysis Engineer (GPU Servers / Data Center). Be the first to apply!
- systems engineer Austin, TX
- system engineer contract Austin, TX
- system performance engineer Austin, TX
- system design engineer Austin, TX
- software system engineer Austin, TX
- computer system validation engineer Austin, TX
- unix linux systems engineer Austin, TX
- mission system engineer Austin, TX
- space systems engineer Austin, TX
- microsoft systems engineer Austin, TX

