Sr Systems Development Engineer, AWS Hardware Engineering Services, AI UltraServers
AmazonWebServices
AWS runs the world's largest fleet of AI/ML accelerator servers. When a model with billions of parameters trains across a large scale of GPUs, every minute of downtime costs real progress. We are building the automation, diagnostics, and predictive intelligence that keeps this fleet running at peak. If you want to work at the intersection of hardware, software, and scale — where your code directly prevents customer-impacting failures — this is the role.We are seeking a Systems Development Engineer to build automation software, diagnostic tooling, and fleet health infrastructure for our accelerated compute platforms. You will work across multiple teams and organizations to design scalable, reliable systems for our accelerated compute fleet.What You Will DoYou will tackle problems no one has fully defined yet — spanning hardware, firmware, kernel, and software simultaneously. You will own systems end to end, writing code that prevents failures rather than reacts to them, and building automation that replaces manual toil with intelligent self-healing. You will work across PCIe topology, GPU diagnostics, Linux drivers, and telemetry pipelines to correlate signals and isolate faults at fleet scale. When your system catches a failing GPU before a training job crashes, that is your impact.Why You Will Love ItYour automation runs at a large scale across servers in the cloud. When you ship, you see failure rates move within days. The team is small enough that your decisions shape the architecture, and large enough that you will always have experts to learn from across hardware, firmware, and software.The Ideal CandidateYou know the full stack from bare-metal to userland. You debug at the intersection of components, not just within them. You build at cloud scale and care how your systems decisions impact customers. You are an excellent communicator who can drive alignment across hardware, software, and operations teams.Key job responsibilitiesFleet Health & Predictive Infrastructure1. Build and own the automation infrastructure for accelerator (AI/ML) fleet health at a large scale of servers, driving toward zero-touch operations that detect, diagnose, triage, and remediate faults without human intervention2. Design and develop test frameworks, test coverage strategies, and diagnostic tooling to validate hardware functionality, detect faults, and ensure qualification coverage across the platform lifecycle.3. Design predictive failure detection using telemetry, sensor data, error trending, and log correlation to identify degrading components before customer impact4. Develop monitoring dashboards and alerting for real-time fleet health visibility across manufacturing, lab, and production environments5. Define and track fleet health metrics: failure rates, mean time to detect and resolve issues, first-time fix rate, test dwell time, and predictive accuracyDebugging & Troubleshooting1. Debug complex system-level issues across compute, GPU, and networking in production — including Linux boot/runtime failures, PCIe, power, NIC, NVMe, and GPU subsystems on x86 and ARM2. Perform root cause analysis correlating across firmware, kernel, driver, and physical layer; feed findings into manufacturing quality and design improvementsSystems Development & Automation1. Design scalable test automation for hardware bring-up, regression, and qualification — reducing manufacturing test cycle times without sacrificing coverage through intelligent test sequencing and parallel execution2. Build data pipelines correlating test results, sensor telemetry, and component-level data to identify systemic yield issues and drive upstream fixes3. Develop and maintain Linux device drivers on ARM and x86; work with OS internals and accelerator/GPU software stacks4. Build and manage tests covering all functional aspects of the system and CI/CD pipelines for rapid deployment to manufacturing lines and production fleetCross-Team Collaboration1. Work across engineering teams and internal customers to ensure new accelerated compute hardware meets data path, control path, and onboarding requirements2. Engage with ODMs and design partners on testability, diagnostic coverage, and automation requirements during hardware design and bring-up phases — influencing functional and performance readiness of the platform3. Partner with datacenter operations to close the loop between field failures, manufacturing escapes, and design improvementsOperational Excellence1. Participate in post-incident reviews, identify contributing causes and drive permanent fixes that eliminate whole classes of risk2. Produce clear, maintainable documentation for systems, runbooks, and automation to enable others to operate and extend your work3. Drive process improvements that increase team agility — reducing development friction, eliminating unnecessary gates, and improving delivery velocityMay require occasional (<10%) regional and international travel to Design and Manufacturing Partner sites.A day in the lifeYou start the day reviewing overnight validation run results, triaging a cluster of GPU errors that correlate with a specific firmware version. Mid-morning, you push a fix to your diagnostic automation pipeline and validate it catches the failure pattern in your test environment. In the afternoon, you join a hardware bring-up call with your ODM partner to debug a PCIe link training failure on a new EVT board, walking the team through kernel logs and signal integrity data. You end the day reviewing a pull request from a teammate on a new telemetry correlation engine, and updating your manufacturing test coverage dashboard with the latest yield data.About the teamThe Hardware Engineering AI/ML UltraServer platform team is a group of engineers and technical program managers directly responsible for launching GPU-accelerated servers into the AWS fleet. Located in Seattle, Austin, and Cupertino, we collaborate with global development teams and ODM partners to deliver next-generation AI/ML infrastructure deployed in datacenters worldwide. We move fast with small, empowered teams delivering end-to-end — from server conception through fleet-scale operations.Basic qualifications- 6+ years of programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, Ruby experience- 6+ years of non-internship professional software development experience- 6+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience- 5+ years of deploying and operating in a Linux/Unix environment experience- 6+ years of systems design, software development, operations, automation, and process improvement experience- Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent work experiencePreferred qualification - 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience- Master's degree or PhD in Computer Science, Electrical Engineering, or a related field- Track record leading delivery of complex software or computing infrastructure at production scale, while also delivering through others by force-multiplying engineering efforts.- 5+ years of experience building end-to-end automation for telemetry-based anomaly detection and intelligent remediation at fleet scale.- Hands-on experience troubleshooting and debugging system integration and low level issues across server and GPU/accelerator hardware, BMC/IPMI, firmware, PCIe topology, driver integration, and hardware-level fault isolation- Experience working with ODMs or hardware design partners on NPI validation- Experience working in large-scale datacenter or cloud environments and setting engineering best practices and operational standards across multiple teams.Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at .USA, CA, Cupertino - 173,900.00 - 235,200.00 USD annuallyUSA, TX, Austin - 151,200.00 - 204,600.00 USD annuallyUSA, WA, Seattle - 151,200.00 - 204,600.00 USD annually
$151.2k - $204.6k
...developing highly available and large-scale services? Would you like to invent and build... ...? Are you an experienced VMware Systems development engineer and a self-starter who is excited to build... ...run VMware-based workloads on AWS. The AWS Commercial Applications team...Amazon Web ServiceSeniorFlexible hours$159.2k - $215.3k
AWS Infrastructure Services owns the design, planning, delivery, and... ...Quality & Reliability Engineer to own the end-to-... ...next-generation server systems. This role will... ...reliably from initial development through deployment in... ...or high-tech hardware- Demonstrated ability...Amazon Web ServiceSeniorWork experience placementFlexible hours$183k - $247.6k
Amazon Web Services (AWS) Hardware Engineering Services (HWEngS) owns the new product development (NPI) and operation of all AWS global infrastructure. In other words, we’re the... ...the teamThe team is comprised of CHDE's, System Development Engineers and Technical Program...Amazon Web ServiceSeniorLocal areaFlexible hours$99.1k - $160k
...most customer centric experiences? As a System Development Engineer on the Endpoint Security Team, you'll... ...are integral to how Amazon's Customer Service Associates worldwide interact with... ...solutions will handle billions of records in AWS while maintaining Amazon's high...Amazon Web ServiceWorldwideFlexible hours$99.1k - $160k
...role is located in Seattle, WA. As a Systems Development Engineer on our team, you will design, build, and... ...stack, maintain high-availability services, and build observability through metrics... ...configuration and features in IAM, with AWS technology at the forefront.• Manage...Amazon Web ServiceFlexible hours$99.1k - $160k
Amazon Foundational Security Services (AFSS) owns the service portfolio that provides... ...IDPs, including Google Workspace and AWS Identity Center (IDC). We enable the authentication... ..., wiki pages, and Workdocs.As a Systems Development Engineer , you will design, build, and operate...Amazon Web ServicePermanent employmentFlexible hours$159.2k - $215.3k
Amazon Web Services’ Hardware Engineering team is looking for experienced professionals to help... ...closely with our ODMs, CMs, internal AWS hardware and software development teams, and with upstream... ...liquid component or air conditioning system manufacturing- 2 years of...Amazon Web ServiceSeniorContract workWork experience placementFlexible hours- Amazon Web Services (AWS) invites experienced software engineers to join a new hardware engineering team in Seattle. This full-time role focuses on distributed systems development across storage, HSM, AI accelerators architectures, using C++/Java/Python. As a member of...Amazon Web ServiceSeniorFull time
$183k - $247.6k
AWS Hardware Engineering Services owns the design, planning, delivery, and operation of all AWS global infrastructure... ...will own and lead the design and development of server products utilizing a wide... ..., or large-scale distributed systems experience- Experience with feature...Amazon Web ServiceSeniorLocal areaFlexible hours$129.2k - $174.8k
Amazon Supply Chain Services (ASCS) is looking for a Systems Development Engineer, focusing on Salesforce and AWS Development to join the team. This role will be responsible for identifying internal and external solutions to business requirements, designing and setting...Amazon Web ServiceFlexible hours$148.7k - $201.2k
AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS... ...join a diverse team of software, hardware, and network engineers, supply chain specialists,... ...to completion.We are seeking an Systems Development Engineer to join our Fleet Automation...Amazon Web ServiceInternshipFlexible hours$159.2k - $215.3k
AWS EC2 owns the design, planning, delivery, and operation of all AWS server... .... You’ll join a diverse team of hardware engineers, software engineers, system engineers, technical program managers... ...day in the lifeWhy AWS Amazon Web Services (AWS) is the world’s most comprehensive...Amazon Web ServiceSeniorFlexible hours$151.2k - $204.6k
...lowest latency and highest throughput services in all of AWS. CloudFront is a fast content... ...other factors. Amazon CloudFront’s Engineering team is looking for an experienced Systems Development Engineer to build and develop new hardware and provisioning platforms, and build...Amazon Web ServiceSeniorInternshipFlexible hours$152.2k - $205.9k
...projects to bring Generative AI to more developers... ...and adoption of the services. You'll use your passion... ...offering for customers.As a Sr. Product Manager... ...leadership community at AWS. This community plays a... ...of Product Planning and Development - Including customer goals...Amazon Web ServiceSeniorWorldwideFlexible hours$129.2k - $174.8k
The AWS Cross Domain Services team is seeking a Systems Development Engineer who is passionate about operational excellence. This engineer will help to operate and scale... ...tasks in the future.AWS Infrastructure Services (AIS)AWS Infrastructure Services owns the design,...Amazon Web ServiceFlexible hours$99.1k - $160k
...following locations:Seattle, WA, USAAs a Systems Development Engineer I you will have the opportunity to... ...infrastructure that powers Amazon's vast range of services, enabling them to deliver for... ...the productivity and savings for AWS and Amazon.Our ideal candidate will have...Amazon Web ServiceTemporary workFlexible hours$129.2k - $174.8k
...world, and we’ve designed the system with the capacity,... ...operate the organization’s services. You will own and operate... ...critical to Amazon Leo including engineering applications, AWS Cloud Infrastructure, and... ...in Agile/Scrum development processes, contributing to...Amazon Web ServiceLocal areaFlexible hours$231.2k - $312.8k
...position is part of the AWS Specialist and... ...strategy, recruiting, development, and growth of our... ...in communication services (email, SMS, WhatsApp... ...(WWSO) team as the Sr. Manager, Solutions... ...and agentic AI — where intelligent... ...Solutions Architects, engineers, or equivalent).- Experience...Amazon Web ServiceSeniorWork experience placementLocal areaWorldwideFlexible hours$168.1k - $227.4k
As part of the AWS Applied AI Solutions organization, we... ...applications.The Core Services team within AWS Applied... ..., pioneering the development of reusable AI components... ...requiring multiple engineers, balancing business goals... ...of new and existing systems experience-...Amazon Web ServiceSeniorInternshipWorldwideFlexible hoursDay shift$129.2k - $174.8k
...where it will be performed.Amazon Leo Productivity Engineering Application Services (ALPEAS) is looking for a System Development Engineer to own the deployment, administration,... .... You will build and operate infrastructure in AWS GovCloud to support Leo's systems engineering,...Amazon Web ServiceFlexible hours$129.2k - $174.8k
At AWS, we're working to be the most customer... ...AWS Infrastructure Services owns the design, planning... ...team of software, hardware, and network engineers, supply chain... ...responsibilitiesAs a System Development Engineer, you will:•... ...methods such as Agentic AI* Work with business...Amazon Web ServiceFlexible hours$137.3k - $185.7k
As a Sr .Hardware Development Engineer in Amazon’s Last Mile Automation and Robotics group... ..., reliability, and serviceability across global CM partners.... ...with AR (Advanced Robotics), Systems Engineering, and Manufacturing... ...(e.g., automation, AI driven testing) to enhance...SeniorContract workWorldwideFlexible hours$171.5k - $257.6k
...technology - hyperscalers, AI labs, the AI hardware supply chain, data... ...results-driven Consulting Systems Engineer for our Service Provider team to support... ...architectures ~ Familiarity with AWS and Azure cloud (public... ...the growth and development of every person, cultivating...Amazon Web ServiceLocal areaRemote workFlexible hours$120k - $170k
Sr. Manager/Manager Site Reliability Engineering Join to apply for the Sr. Manager/Manager... ...at Aritzia Get AI-powered advice on this... ...of the Quality and Service Delivery team and... ...progressive career development and an incredible employee... ...architectures (AWS, GCP, Kubernetes,...Amazon Web ServiceSeniorFull timeWork at officeRemote workFlexible hours- ...Governance & Data Access) and AI-powered workflow automation services Create and build product... ...Work closely with engineering teams to define technical... ...understanding of cloud services (AWS) Strong understanding of... ...with enterprise software development lifecycle Knowledge of...Amazon Web ServiceSenior
- ...Systems Development EngineerOne of the world's highest scale storage platforms... ...for a Systems Development Engineer with an obsession for... ...the world-class Amazon Web Services (AWS) S3 storage service. You have... ...patterns for the full software/hardware/networks development life...Amazon Web ServiceInternshipWorldwideFlexible hours
$173.9k - $235.2k
...position is part of the AWS Specialist and Partner... ...strategy, recruiting, development, and growth of our key... ....Within ASP, the AI & Strategic Partner Engineering team is seeking a Senior Systems Development Engineer to... ...BTP AI Core) and AWS AI services (Bedrock, Amazon Q), and...Amazon Web ServiceInternshipWorldwideFlexible hours$129.2k - $174.8k
AWS Infrastructure Services owns the design, planning, delivery, and operation... ...team of software, hardware, and network engineers, supply chain specialists... ...seeking a Linux/Unix System Engineer with the ability... ...in at least one development language. They must be...Amazon Web ServicePermanent employmentInternshipFlexible hours$183k - $247.6k
AWS operates the world's largest fleet... ...servers powering AI/ML training and inference... ..., drives the hardware designs, and owns... ...a Cloud Hardware Development Engineer to define server... ...to identify systemic issues and drive... ...are debuggable, serviceable, and automation-readyMay...Amazon Web ServiceSeniorWork at officeLocal areaWorldwideFlexible hoursDay shift$159.2k - $215.3k
Amazon Web Services (AWS) Hardware Engineering is a leading-edge product development team that creates enterprise compute and storage... ...solutions for advanced thermal systems? If you are a materials engineer... ....AWS is looking for a Sr. Materials Engineer to become...Amazon Web ServiceSeniorFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Sr Systems Development Engineer, AWS Hardware Engineering Services, AI UltraServers. Be the first to apply!
- application system engineer Seattle, WA
- sr systems engineer Seattle, WA
- software system engineer Seattle, WA
- distributed systems engineer Seattle, WA
- systems engineer intern Seattle, WA
- ground systems engineer Seattle, WA
- healthcare systems engineer Seattle, WA
- space systems engineer Seattle, WA
- senior windows systems engineer Seattle, WA
- electronic systems engineer Seattle, WA



