Sr Systems Development Engineer, AWS Hardware Engineering Services, AI UltraServers
$173.9k - $235.2kAmazon
AWS runs the world's largest fleet of AI/ML accelerator servers. When a model with billions of parameters trains across a large scale of GPUs, every minute of downtime costs real progress. We are building the automation, diagnostics, and predictive intelligence that keeps this fleet running at peak. If you want to work at the intersection of hardware, software, and scale — where your code directly prevents customer-impacting failures — this is the role.We are seeking a Systems Development Engineer to build automation software, diagnostic tooling, and fleet health infrastructure for our accelerated compute platforms. You will work across multiple teams and organizations to design scalable, reliable systems for our accelerated compute fleet.What You Will DoYou will tackle problems no one has fully defined yet — spanning hardware, firmware, kernel, and software simultaneously. You will own systems end to end, writing code that prevents failures rather than reacts to them, and building automation that replaces manual toil with intelligent self-healing. You will work across PCIe topology, GPU diagnostics, Linux drivers, and telemetry pipelines to correlate signals and isolate faults at fleet scale. When your system catches a failing GPU before a training job crashes, that is your impact.Why You Will Love ItYour automation runs at a large scale across servers in the cloud. When you ship, you see failure rates move within days. The team is small enough that your decisions shape the architecture, and large enough that you will always have experts to learn from across hardware, firmware, and software.The Ideal CandidateYou know the full stack from bare-metal to userland. You debug at the intersection of components, not just within them. You build at cloud scale and care how your systems decisions impact customers. You are an excellent communicator who can drive alignment across hardware, software, and operations teams.Key job responsibilitiesFleet Health & Predictive Infrastructure1. Build and own the automation infrastructure for accelerator (AI/ML) fleet health at a large scale of servers, driving toward zero-touch operations that detect, diagnose, triage, and remediate faults without human intervention2. Design and develop test frameworks, test coverage strategies, and diagnostic tooling to validate hardware functionality, detect faults, and ensure qualification coverage across the platform lifecycle.3. Design predictive failure detection using telemetry, sensor data, error trending, and log correlation to identify degrading components before customer impact4. Develop monitoring dashboards and alerting for real-time fleet health visibility across manufacturing, lab, and production environments5. Define and track fleet health metrics: failure rates, mean time to detect and resolve issues, first-time fix rate, test dwell time, and predictive accuracyDebugging & Troubleshooting1. Debug complex system-level issues across compute, GPU, and networking in production — including Linux boot/runtime failures, PCIe, power, NIC, NVMe, and GPU subsystems on x86 and ARM2. Perform root cause analysis correlating across firmware, kernel, driver, and physical layer; feed findings into manufacturing quality and design improvementsSystems Development & Automation1. Design scalable test automation for hardware bring-up, regression, and qualification — reducing manufacturing test cycle times without sacrificing coverage through intelligent test sequencing and parallel execution2. Build data pipelines correlating test results, sensor telemetry, and component-level data to identify systemic yield issues and drive upstream fixes3. Develop and maintain Linux device drivers on ARM and x86; work with OS internals and accelerator/GPU software stacks4. Build and manage tests covering all functional aspects of the system and CI/CD pipelines for rapid deployment to manufacturing lines and production fleetCross-Team Collaboration1. Work across engineering teams and internal customers to ensure new accelerated compute hardware meets data path, control path, and onboarding requirements2. Engage with ODMs and design partners on testability, diagnostic coverage, and automation requirements during hardware design and bring-up phases — influencing functional and performance readiness of the platform3. Partner with datacenter operations to close the loop between field failures, manufacturing escapes, and design improvementsOperational Excellence1. Participate in post-incident reviews, identify contributing causes and drive permanent fixes that eliminate whole classes of risk2. Produce clear, maintainable documentation for systems, runbooks, and automation to enable others to operate and extend your work3. Drive process improvements that increase team agility — reducing development friction, eliminating unnecessary gates, and improving delivery velocityMay require occasional (<10%) regional and international travel to Design and Manufacturing Partner sites.A day in the lifeYou start the day reviewing overnight validation run results, triaging a cluster of GPU errors that correlate with a specific firmware version. Mid-morning, you push a fix to your diagnostic automation pipeline and validate it catches the failure pattern in your test environment. In the afternoon, you join a hardware bring-up call with your ODM partner to debug a PCIe link training failure on a new EVT board, walking the team through kernel logs and signal integrity data. You end the day reviewing a pull request from a teammate on a new telemetry correlation engine, and updating your manufacturing test coverage dashboard with the latest yield data.About the teamThe Hardware Engineering AI/ML UltraServer platform team is a group of engineers and technical program managers directly responsible for launching GPU-accelerated servers into the AWS fleet. Located in Seattle, Austin, and Cupertino, we collaborate with global development teams and ODM partners to deliver next-generation AI/ML infrastructure deployed in datacenters worldwide. We move fast with small, empowered teams delivering end-to-end — from server conception through fleet-scale operations.Basic qualifications- 6+ years of programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, Ruby experience- 6+ years of non-internship professional software development experience- 6+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience- 6+ years of systems design, software development, operations, automation, and process improvement experience- 5+ years of deploying and operating in a Linux/Unix environment experience- Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent work experiencePreferred qualification - Master's degree or above in computer science, electrical engineering, or related field- 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience- Experience taking a leading role in building complex software or computing infrastructure that has been successfully delivered to customers- 5+ years of building scripts, tooling, and automation for large-scale computing environments experience- Experience with general troubleshooting/debugging of hardware, or experience with CUDA kernels or ML/low-level kernels- 5+ years of hardware design and validation of components, subsystems and systems experience- Experience leading technical teams through daily operations and maintenance evolutions, or experience in data center engineering or operationsAmazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at .USA, CA, Cupertino - 173,900.00 - 235,200.00 USD annuallyUSA, TX, Austin - 151,200.00 - 204,600.00 USD annuallyUSA, WA, Seattle - 151,200.00 - 204,600.00 USD annually
$173.9k - $235.2k
...Sr Systems Development Engineer, AWS Hardware Engineering Services, AI UltraServers Job ID: 10569346 | Amazon Data Services, Inc. AWS runs the world's largest fleet of AI/ML accelerator servers. When a model with billions of parameters trains across a large scale...Amazon Web ServiceSeniorPermanent employmentWork experience placementInternshipLocal areaWorldwideFlexible hoursNight shiftDay shift$151.2k - $204.6k
...OTS Supply Chain team as a Sr. Systems Development Engineer to transform technology infrastructure... ...through innovative Gen-AI powered solutions to... ...implement real-time validation services, and develop workflow... ...driven architectures using AWS services (Lambda, API Gateway...Amazon Web ServiceSeniorFull timeTemporary workSeasonal workFlexible hoursNight shift$183k - $247.6k
...shape the future of AI? Join the team... ...opportunity to build the systems that define what’s next for AWS — and for the... ...of software, hardware, and network engineers, supply chain specialists... ...system development on top of your server... ...AWSAmazon Web Services (AWS) is the world...Amazon Web ServiceSeniorLocal areaFlexible hours$173.9k - $235.2k
...future of Generative AI at AWS? Join the team building... ...to build the systems that define what’s next... ...ll join a diverse AWS Hardware Engineering team of software, hardware... ..., EC2, other AWS services) through server conception... ...software development experience- 4+ years...Amazon Web ServiceSeniorInternshipLocal areaFlexible hours$173.9k - $235.2k
...an experienced Senior Systems Development Engineer to lead the... ...edge and accelerated (AI/ML) compute fleet healthy... ...using a combination of hardware, software, system design... ...future designs for AWS server solutions. You... ...needed by dependent service teams- Work closely with...Amazon Web ServiceSeniorInternshipLocal areaWorldwideFlexible hours$159.2k - $215.3k
Shape the future of AI infrastructure! Join AWS as we scale our Manufacturing Test Engineering team developing comprehensive... ...AI acceleration hardware, deployed across our global... ...Organization and seeking talented Systems Test Engineers, Test Development Engineers, and...Amazon Web ServiceSeniorFlexible hours$159.2k - $215.3k
Amazon Web Services’ Hardware Engineering team is looking for experienced professionals to help build the world’s premier cloud computing... ...team works closely with our ODMs, CMs, internal AWS hardware and software development teams, and with upstream component vendors to...Amazon Web ServiceSeniorContract workWork experience placementFlexible hours$183k - $247.6k
...the Next Generation of AI accelerator compute systems? Lead bleeding-edge HW development projects? Have you heard of Amazon Web Services (AWS) Project Rainer? This... ...from our own Silicon to Hardware to Software and deploy... ...Lead System Design Engineers to build the next generation...Amazon Web ServiceSeniorLocal areaFlexible hours$129.2k - $174.8k
...rewarding, fulfilling and fun.As a Systems Development Engineer on the Customer Experience... ...to building innovative self-service infrastructure management... .... You'll help develop AWS-style services and automation... ...knowledge management systems, and AI-powered customer experiences...Amazon Web ServiceFull timeTemporary workInternshipSeasonal workWorldwideFlexible hours- ...Amazon Robotics is seeking a Systems Development Engineer II to help build self-service infrastructure management tools across our CXI platform. You will contribute to AWS-style services, automate deployments, and enable thousands of solution owners to discover, deploy...Amazon Web Service
$136k - $184k
...are at the forefront of hardware/software co-design not just in Amazon Web Services (AWS) but across the... ...world. We are seeking an engineer who is comfortable debugging... ...team responsible for system remediation,... ...the lifeAs a Platform Development Engineer, you are the...Amazon Web ServicePermanent employmentInternshipFlexible hours$148.7k - $201.2k
...backbone of Generative AI cloud at AWS? Do you want to build... ...release of newer AWS services and instance types that... ...builders like you.The AWS Hardware Engineering team creates server... ...building intelligent systems that drive the debug and development of next-generation cloud...Amazon Web ServiceInternshipLocal areaFlexible hours$183k - $247.6k
...Description AWS Infrastructure Services owns the design, planning, delivery... ...team of software, hardware, and network engineers, supply chain specialists... ...lead the design and development of server products utilizing... ...with a focus on system / server development in...Amazon Web ServiceSeniorLocal areaFlexible hours- ...Amazon Data Services, Inc. in Austin, TX is seeking a Sr Systems Development Engineer for the AI UltraServers program to build automation, diagnostics, and fleet health tooling... ...You will own end-to-end systems across hardware, firmware, kernel, and software, designing...Senior
$159.1k - $215.3k
...world.The Satellite Systems Integration and Test... ...networks, and cloud services. We integrate new capabilities... ....As a Senior Network Development Engineer, you will lead hands-... ...with systems, hardware, software, RF, networking... ....- Experience with AWS services, cloud...Amazon Web ServiceSeniorPermanent employmentWork experience placementLocal areaFlexible hours$153.6k - $207.8k
AWS Global Sales drives adoption of the AWS cloud worldwide... ...by providing tailored service, unmatched technology,... ...exciting advances in Generative AI and intelligent... ...domain areas (e.g. software development, cloud computing, systems engineering, infrastructure, security...Amazon Web ServiceSeniorWorldwideFlexible hours$184k - $287.5k
We are seeking a Linux Systems Engineer to join our... ...and Terraform.Tooling Development: Write high-quality, maintainable... ....Naming & Storage Services: Manage Linux Naming Services... ...managing GCP and AWS compute services Experience... ...vacancy. NVIDIA uses AI tools in its...Amazon Web ServiceSeniorFull timeWork experience placementRemote work$151.2k - $204.6k
...Operations Infrastructure Services (OIS) team is building... ..., analogous to AWS.com, enabling solution... ...and emerging agentic AI systems.Key job responsibilitiesOwn... ...-service.Partner with engineering to translate product... ...product marketing, business development or technology...Amazon Web ServiceSeniorFull timeContract workTemporary workSeasonal workFlexible hoursShift work- ...The Content & Industry Services team is the engine behind that data — we... ...mission critical systems that integrate with thousands... ...involve investing in AI-driven solutions to... ...in a product development process that is primarily... ...MySQL, PostgreSQL, and AWS services (EC2, S3,...Amazon Web ServiceSeniorWork at officeLocal area
$124k - $280k
...in data and analytics engineering focus on leveraging advanced... ...implementing advanced AI and ML solutions to... ...algorithms, models, and systems to enable intelligent decision... ...Cloud Platforms [e.g., AWS Certified Solutions... ...solutions using cloud services- Designing and managing...Amazon Web ServiceSeniorFull timeH1b- ...GlobalFoundries is a leading full-service semiconductor foundry... ...of design, development, and fabrication services... ...the technologies and systems that transform... ...integrating manufacturing and engineering data out of a high... ...teams. Experience with AWS cloud Services including...Amazon Web ServiceSeniorFull timeLocal area
$183k - $247.6k
The Annapurna AI Manufacturing, Quality... ...is part of AWS Annapurna Labs... ...largest Cloud Services provider. As a... ...Senior Reliability Engineer you will engage... ...of computer systems to influence design... ...products under development and products in... ..., firmware, hardware, and silicon...Amazon Web ServiceSeniorWork experience placementLocal areaFlexible hours- ...Product Designer to own end-to-end design for AWS Marketplace and Partner Services (AMPS), spanning traditional web consoles and AI interfaces. You will set design strategy,... ...tools, and collaborate with product, engineering, and business leaders to shape buyer and partner...Amazon Web ServiceSenior
$104k - $164k
...looking for a Senior Lead Systems Engineer who's passionate about... ...software engineering, AI, and intelligent... ...automation, and self-service.You'll lead complex technical... ...platforms such as AWS, GCP, or AzureExperience... ...leaveProfessional development supported by formal career...Amazon Web ServiceSeniorWork at officeLocal area$119.2k - $175.45k
...business analytics platform engineering team is at the... ...the power of Azure and AWS to deliver enterprise-grade... ...the company.As a Senior Systems Engineer, you'll play a... ...architecture with an AI‑enabled experience layer... ...expertise in Azure cloud services, particularly Cognos...Amazon Web ServiceSeniorFull timeH1bLocal areaWork from homeRelocation packageFlexible hours$140k - $215k
...world’s most advanced AI-native platform.... ...scale distributed systems, processing almost... ...libraries, services, and tooling that... ...directly with product engineering teams and their leadership... ...every feature development team and provide... ...on experience with AWS, Cassandra, Kafka,...Amazon Web ServiceSeniorFull timeWork experience placementWork at officeLocal area2 days per week3 days per week$129.2k - $174.8k
...Foundations org is looking for a passionate and innovative Systems Development Engineer to help tackle a diverse landscape of technical... ...technologies including Java, ASP.NET, C/C#, Windows, Linux, and AWS services.- Solve complex, novel technical challenges on a daily basis...Amazon Web ServiceWorldwideFlexible hours$129.2k - $174.8k
...network of over 4,000 field technicians and engineers. These professionals are vital in... ..., or areas where your team’s systems hinder the success or innovation of other... ...Bachelor's degree- Experience building services using AWS products- Experience utilizing AWS cloud...Amazon Web ServiceFull timeTemporary workSeasonal workFlexible hours$99.1k - $160k
The Device Tech team is seeking a Systems Development Engineer to build and maintain the platforms and automation... ...engineers to design and build self-service tools, automation workflows, and... ...globally. You will gain deep exposure to AWS services, hybrid cloud architecture,...Amazon Web ServiceFull timeTemporary workWork experience placementSeasonal workFlexible hoursDay shift$94.2k - $160k
An exciting opportunity for a Systems Development Engineer I to join our team and contribute to the transformation of our supply chain technology... ...methodologies (including scrum)- Knowledge of cloud services such as AWS or equivalentAmazon is an equal opportunity employer...Amazon Web ServiceFull timeTemporary workSeasonal workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Sr Systems Development Engineer, AWS Hardware Engineering Services, AI UltraServers. Be the first to apply!
- senior linux systems engineer Austin, TX
- system engineer contract Austin, TX
- data systems engineer Austin, TX
- senior staff systems engineer Austin, TX
- microsoft systems engineer Austin, TX
- operations support system engineer Austin, TX
- advanced systems engineer Austin, TX
- software system engineer Austin, TX
- space systems engineer Austin, TX
- system test engineer Austin, TX


