Sr Systems Development Engineer, AWS Hardware Engineering Services, AI UltraServers
$173.9k - $235.2kAmazonWebServices
AWS runs the world's largest fleet of AI/ML accelerator servers. When a model with billions of parameters trains across a large scale of GPUs, every minute of downtime costs real progress. We are building the automation, diagnostics, and predictive intelligence that keeps this fleet running at peak. If you want to work at the intersection of hardware, software, and scale — where your code directly prevents customer-impacting failures — this is the role.We are seeking a Systems Development Engineer to build automation software, diagnostic tooling, and fleet health infrastructure for our accelerated compute platforms. You will work across multiple teams and organizations to design scalable, reliable systems for our accelerated compute fleet.What You Will DoYou will tackle problems no one has fully defined yet — spanning hardware, firmware, kernel, and software simultaneously. You will own systems end to end, writing code that prevents failures rather than reacts to them, and building automation that replaces manual toil with intelligent self-healing. You will work across PCIe topology, GPU diagnostics, Linux drivers, and telemetry pipelines to correlate signals and isolate faults at fleet scale. When your system catches a failing GPU before a training job crashes, that is your impact.Why You Will Love ItYour automation runs at a large scale across servers in the cloud. When you ship, you see failure rates move within days. The team is small enough that your decisions shape the architecture, and large enough that you will always have experts to learn from across hardware, firmware, and software.The Ideal CandidateYou know the full stack from bare-metal to userland. You debug at the intersection of components, not just within them. You build at cloud scale and care how your systems decisions impact customers. You are an excellent communicator who can drive alignment across hardware, software, and operations teams.Key job responsibilitiesFleet Health & Predictive Infrastructure1. Build and own the automation infrastructure for accelerator (AI/ML) fleet health at a large scale of servers, driving toward zero-touch operations that detect, diagnose, triage, and remediate faults without human intervention2. Design and develop test frameworks, test coverage strategies, and diagnostic tooling to validate hardware functionality, detect faults, and ensure qualification coverage across the platform lifecycle.3. Design predictive failure detection using telemetry, sensor data, error trending, and log correlation to identify degrading components before customer impact4. Develop monitoring dashboards and alerting for real-time fleet health visibility across manufacturing, lab, and production environments5. Define and track fleet health metrics: failure rates, mean time to detect and resolve issues, first-time fix rate, test dwell time, and predictive accuracyDebugging & Troubleshooting1. Debug complex system-level issues across compute, GPU, and networking in production — including Linux boot/runtime failures, PCIe, power, NIC, NVMe, and GPU subsystems on x86 and ARM2. Perform root cause analysis correlating across firmware, kernel, driver, and physical layer; feed findings into manufacturing quality and design improvementsSystems Development & Automation1. Design scalable test automation for hardware bring-up, regression, and qualification — reducing manufacturing test cycle times without sacrificing coverage through intelligent test sequencing and parallel execution2. Build data pipelines correlating test results, sensor telemetry, and component-level data to identify systemic yield issues and drive upstream fixes3. Develop and maintain Linux device drivers on ARM and x86; work with OS internals and accelerator/GPU software stacks4. Build and manage tests covering all functional aspects of the system and CI/CD pipelines for rapid deployment to manufacturing lines and production fleetCross-Team Collaboration1. Work across engineering teams and internal customers to ensure new accelerated compute hardware meets data path, control path, and onboarding requirements2. Engage with ODMs and design partners on testability, diagnostic coverage, and automation requirements during hardware design and bring-up phases — influencing functional and performance readiness of the platform3. Partner with datacenter operations to close the loop between field failures, manufacturing escapes, and design improvementsOperational Excellence1. Participate in post-incident reviews, identify contributing causes and drive permanent fixes that eliminate whole classes of risk2. Produce clear, maintainable documentation for systems, runbooks, and automation to enable others to operate and extend your work3. Drive process improvements that increase team agility — reducing development friction, eliminating unnecessary gates, and improving delivery velocityMay require occasional (<10%) regional and international travel to Design and Manufacturing Partner sites.A day in the lifeYou start the day reviewing overnight validation run results, triaging a cluster of GPU errors that correlate with a specific firmware version. Mid-morning, you push a fix to your diagnostic automation pipeline and validate it catches the failure pattern in your test environment. In the afternoon, you join a hardware bring-up call with your ODM partner to debug a PCIe link training failure on a new EVT board, walking the team through kernel logs and signal integrity data. You end the day reviewing a pull request from a teammate on a new telemetry correlation engine, and updating your manufacturing test coverage dashboard with the latest yield data.About the teamThe Hardware Engineering AI/ML UltraServer platform team is a group of engineers and technical program managers directly responsible for launching GPU-accelerated servers into the AWS fleet. Located in Seattle, Austin, and Cupertino, we collaborate with global development teams and ODM partners to deliver next-generation AI/ML infrastructure deployed in datacenters worldwide. We move fast with small, empowered teams delivering end-to-end — from server conception through fleet-scale operations.Basic qualifications- 6+ years of programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, Ruby experience- 6+ years of non-internship professional software development experience- 6+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience- 5+ years of deploying and operating in a Linux/Unix environment experience- 6+ years of systems design, software development, operations, automation, and process improvement experience- Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent work experiencePreferred qualification - 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience- Master's degree or PhD in Computer Science, Electrical Engineering, or a related field- Track record leading delivery of complex software or computing infrastructure at production scale, while also delivering through others by force-multiplying engineering efforts.- 5+ years of experience building end-to-end automation for telemetry-based anomaly detection and intelligent remediation at fleet scale.- Hands-on experience troubleshooting and debugging system integration and low level issues across server and GPU/accelerator hardware, BMC/IPMI, firmware, PCIe topology, driver integration, and hardware-level fault isolation- Experience working with ODMs or hardware design partners on NPI validation- Experience working in large-scale datacenter or cloud environments and setting engineering best practices and operational standards across multiple teams.Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at .USA, CA, Cupertino - 173,900.00 - 235,200.00 USD annuallyUSA, TX, Austin - 151,200.00 - 204,600.00 USD annuallyUSA, WA, Seattle - 151,200.00 - 204,600.00 USD annually
$151.2k - $204.6k
...developing highly available and large-scale services? Would you like to invent and build... ...? Are you an experienced VMware Systems development engineer and a self-starter who is excited to build... ...run VMware-based workloads on AWS. The AWS Commercial Applications team...Amazon Web ServiceSeniorFlexible hours$151.2k - $204.6k
AWS Infrastructure Services owns the design, planning, delivery, and operation... ...team of software, hardware, and network engineers, supply chain specialists... ...will contribute to the development of a packet processor... ...responsibilities* Develop software systems and successfully...Amazon Web ServiceSeniorWorldwideFlexible hours$151.2k - $204.6k
...we’ve designed the system with the capacity, flexibility... ...Infrastructure & Engineering (ALPINE) owns and... ...across commercial AWS GovCloud... ...for a Senior Systems Development Engineer to serve as... ...across ALPINE-managed services while driving AI integration readiness...Amazon Web ServiceSeniorFlexible hoursShift workNight shiftDay shift$159.2k - $215.3k
AWS EC2 owns the design, planning, delivery, and operation of all AWS server... ...You’ll join a diverse team of hardware engineers, software engineers, system engineers, technical program managers... ...the life Why AWS Amazon Web Services (AWS) is the world’s most comprehensive...Amazon Web ServiceSeniorFlexible hours$159.2k - $215.3k
AWS Infrastructure Services owns the design, planning, delivery, and... ...Quality & Reliability Engineer to own the end-to-... ...next-generation server systems. This role will... ...reliably from initial development through deployment in... ...or high-tech hardware- Demonstrated ability...Amazon Web ServiceSeniorWork experience placementFlexible hours$183k - $247.6k
AWS Infrastructure Services owns the design, planning, delivery,... ...team of software, hardware, and network engineers, supply chain specialists... ...types). We solve systemic hardware issues... ...through development and production. You... ...Hardware Engineering AI / ML development team...Amazon Web ServiceSeniorLocal areaFlexible hours$159.2k - $215.3k
Amazon Web Services’ Hardware Engineering team is looking for experienced professionals to help build the world’s premier cloud computing... ...team works closely with our ODMs, CMs, internal AWS hardware and software development teams, and with upstream component vendors to...Amazon Web ServiceSeniorContract workWork experience placementFlexible hours$148.7k - $201.2k
AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS... ...join a diverse team of software, hardware, and network engineers, supply chain specialists,... ...to completion.We are seeking an Systems Development Engineer to join our Fleet Automation...Amazon Web ServiceInternshipFlexible hours$151.2k - $204.6k
...lowest latency and highest throughput services in all of AWS. CloudFront is a fast content... ...other factors. Amazon CloudFront’s Engineering team is looking for an experienced Systems Development Engineer to build and develop new hardware and provisioning platforms, and build...Amazon Web ServiceSeniorInternshipFlexible hours$129.2k - $174.8k
...where it will be performed.Amazon Leo Productivity Engineering Application Services (ALPEAS) is looking for a System Development Engineer to own the deployment, administration,... .... You will build and operate infrastructure in AWS GovCloud to support Leo's systems engineering,...Amazon Web ServiceFlexible hours$137.3k - $185.7k
As a Sr .Hardware Development Engineer in Amazon’s Last Mile Automation and Robotics group... ..., reliability, and serviceability across global CM partners.... ...with AR (Advanced Robotics), Systems Engineering, and Manufacturing... ...(e.g., automation, AI driven testing) to enhance...SeniorContract workWorldwideFlexible hours$129.2k - $174.8k
At AWS, we're working to be the most customer... ...AWS Infrastructure Services owns the design, planning... ...team of software, hardware, and network engineers, supply chain... ...responsibilitiesAs a System Development Engineer, you will:•... ...methods such as Agentic AI* Work with business...Amazon Web ServiceFlexible hours- ...Governance & Data Access) and AI-powered workflow automation services Create and build product... ...Work closely with engineering teams to define technical... ...understanding of cloud services (AWS) Strong understanding of... ...with enterprise software development lifecycle Knowledge of...Amazon Web ServiceSenior
$129.2k - $174.8k
Amazon Supply Chain Services (ASCS) is looking for a Systems Development Engineer, focusing on Salesforce and AWS Development to join the team. This role will be responsible for identifying internal and external solutions to business requirements, designing and setting...Amazon Web ServiceFlexible hours$148.7k - $201.2k
The AWS In Rack Power Team is driving rapid... ...in the power systems used by Amazon Web Services. Our designs are fundamentally... ..., and Outpost. Our engineers solve challenging... ...to lead the development, deployment and sustaining... ...customers, hardware and software development...Amazon Web ServiceSeniorFlexible hours$183k - $247.6k
AWS Hardware Engineering Services owns the design, planning, delivery, and operation of all AWS global infrastructure... ...will own and lead the design and development of server products utilizing a wide... ..., or large-scale distributed systems experience- Experience with feature...Amazon Web ServiceSeniorLocal areaFlexible hours$165k - $225.6k
...Every Identity, from AI to Human Identity is... ...) team delivers the systems, tools, and services that power internal operations... ...SITE RELIABILITY ENGINEER OPPORTUNITY... ...cloud environments and development tools while strictly... ...infrastructure projects. * AWS Expertise &...Amazon Web ServiceSeniorPermanent employmentFull timeLocal areaWorldwideFlexible hours- ...Senior System Engineer Reporting to the Manager, IT System Operations,... ...Microsoft Windows, VMware, Azure, AWS, and hybrid infrastructure... ...system upgrades, patching, hardware refreshes, and technology modernization... ...findings within established service level agreements (SLAs)....Amazon Web ServiceSenior
$151.2k - $204.6k
Imagine a system that stores petabytes of customer data and handles peaks of... ...looking for an experienced systems development engineer who is interested in building systems... ...at any scale. As a fast-growing service at the core of the AWS Cloud, our business and engineering...Amazon Web ServiceSeniorFlexible hours$173.9k - $235.2k
...position is part of the AWS Specialist and Partner... ...strategy, recruiting, development, and growth of our key... ....Within ASP, the AI & Strategic Partner Engineering team is seeking a Senior Systems Development Engineer to... ...BTP AI Core) and AWS AI services (Bedrock, Amazon Q), and...Amazon Web ServiceSeniorInternshipWorldwideFlexible hours$152.2k - $205.9k
As part of the AWS Applied AI Solutions organization, we... ...a cloud-based email service provider that enables... ...guidance, managing both development and business aspects... ...alongside a dedicated engineering team. Serve on cross-... ...the full software/hardware/networks development...Amazon Web ServiceSeniorWorldwideFlexible hours$152.2k - $205.9k
AI agents are reshaping how software is... ...and deployed — and AWS Marketplace is at... ...party software and services, including AI... ...and partner with engineering teams to deliver solutions... ...and fulfillment systems. Mid-morning, you... ...technical (software development, network...Amazon Web ServiceSeniorInternshipFlexible hoursShift workDay shift$168.1k - $227.4k
As part of the AWS Applied AI Solutions organization, we... ...applications.The Core Services team within AWS Applied... ..., pioneering the development of reusable AI components... ...requiring multiple engineers, balancing business goals... ...of new and existing systems experience-...Amazon Web ServiceSeniorInternshipWorldwideFlexible hoursDay shift$197k - $275.81k
## Sr. Manager - Software EngineeringApplylocations... ...vehicles and systems within a... ...- Software Engineering. You will lead... ...operates the AWS cloud... ...building self-service tools and automating... ...requests away. AI is a core tool... ...put AI-assisted development and automation...Amazon Web ServiceSeniorImmediate startShift work$264.1k - $369.74k
Sr Principal Software Engineering - Enterprise Technology page is loaded##... ...space vehicles and systems within a culture of... ...Origin’s Software Development organization. This... ...novel applications of AI to enhance user... ...and utilizing cloud services such as AWS.* Proven ability to...Amazon Web ServiceSeniorPermanent employmentTemporary workLocal areaRelocation$124k - $280k
...in data and analytics engineering focus on leveraging advanced... ...implementing advanced AI and ML solutions to... ...algorithms, models, and systems to enable intelligent decision... ...Cloud Platforms [e.g., AWS Certified Solutions... ...solutions using cloud services- Designing and managing...Amazon Web ServiceSeniorFull timeH1b$151.2k - $204.6k
We are looking for a System Development Engineer to design, build, and own the executive decision intelligence platform that serves senior leaders... ...SQL, Spark, or equivalent)- 3+ years working with cloud services (AWS preferred: Lambda, DynamoDB/RDS, S3, Redshift/Athena,...Amazon Web ServiceSeniorFlexible hours$191.4k - $252.72k
...accelerating the development of medical breakthroughs... ...'s best data and AI infrastructure... ...infrastructure and product engineering teams at... ...scale distributed systems, cloud infrastructure... ...providers (e.g., AWS, Azure, GCP), with... ...cloud‑native services. ~ Experience building...Amazon Web ServiceSeniorLocal areaWorldwide- ...shaping the future of AI-native business... ...intersect across AWS. You will influence... ...design, science, engineering, and product leaders... ...multiple AI services and solutions, ensuring... ...interact with AI systems, building trust and... ..., and Business Development teams to deliver compelling...Amazon Web ServiceWorldwide
$208.3k - $281.8k
...new projects to bring Agentic AI to more developers worldwide.... ...for revenue and adoption of the services. You'll use your passion for... ...product leadership community at AWS. This community plays a... ...Execution of Product Planning and Development - Including customer goals and...Amazon Web ServiceLocal areaWorldwideFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Sr Systems Development Engineer, AWS Hardware Engineering Services, AI UltraServers. Be the first to apply!
- systems engineer Seattle, WA
- system performance engineer Seattle, WA
- software system engineer Seattle, WA
- computer system validation engineer Seattle, WA
- mission system engineer Seattle, WA
- space systems engineer Seattle, WA
- microsoft systems engineer Seattle, WA
- healthcare systems engineer Seattle, WA
- ground systems engineer Seattle, WA
- distributed systems engineer Seattle, WA



