Sr Cloud Hardware Dev Engineer, AWS Generative AI & ML Servers
Amazon Locker
AWS operates the world's largest fleet of GPU-accelerated servers powering AI/ML training and inference at cloud scale. Our team defines the server architectures, drives the hardware designs, and owns the fleet quality for these platforms — from component selection through datacenter operations. If you want to shape the physical hardware that frontier models train on, this is the role.We are seeking a Cloud Hardware Development Engineer to define server architectures based on workload demand, translate them into detailed component specifications, and drive validation from PCBA bring-up through rack integration. You will lead ODM design partners through development and production, triage hardware issues across manufacturing and datacenters, and own fleet quality metrics post-launch.What You Will DoYou will define the hardware that runs the world's largest AI training workloads. Your designs span thermal, mechanical, power, and signal integrity across GPU-accelerated platforms. You will drive validation from first silicon through fleet-scale deployment, triage failures correlating across PCIe, power delivery, memory, and accelerator interconnects, and feed root cause findings back into design improvements. When a new server platform launches at a large scale, the architecture, component choices, and quality gates are yours.Key job responsibilitiesArchitecture & Design* Define server architectures based on workload demand and customer requirements, translating them into detailed designs and component specifications that enable high-performance AI training and inference at scale* Work with interdisciplinary teams of component, firmware, test, qualification, and integration engineers to deliver cohesive designs* Drive design reviews with ODM/JDM partners covering schematic, layout, BOM, and manufacturing DFx (Design for Test, Design for Manufacturing)Validation & Bring-up* Define and execute validation strategies from PCBA bring-up through server and rack integration — covering power sequencing, signal integrity, thermal characterization, and accelerator interconnect performance* Own hardware debug during EVT/DVT/PVT builds, correlating failures across PCIe, power rails, memory channels, and GPU subsystems* Triage hardware issues at both ODM facilities and datacenters, conduct root cause analysis, and implement corrective actionsFleet Quality & Continuous Improvement* Own fleet quality metrics post-launch: server-level annualized failure rates and component-level failure modes* Monitor operational telemetry to identify systemic issues and drive design or process changes for current and future platforms* Partner with test and automation teams to improve manufacturing yield and reduce test dwell timesCross-Team Collaboration* Work with EC2 architecture teams to align on instance definitions, workload requirements, and platform trade-offs* Drive ODM/JDM design partners through development milestones and production ramp* Collaborate with firmware, software, and operations teams to ensure designs are debuggable, serviceable, and automation-readyMay require occasional (<10%) regional and international travel to Design and Manufacturing Partner sites.The Ideal CandidateYou think across the full hardware stack — from silicon packaging and power delivery to rack-level thermal and mechanical design. You are as comfortable reviewing a schematic as you are analyzing fleet failure data. You drive quality through data, not assumption, and you hold design partners to the same standard you hold yourself. You mentor and develop junior engineers, contribute to hiring, and share your expertise to make the team stronger.Why You Will Love ItThe world's most advanced frontier models train on the hardware you design. You will see your architecture decisions scale to a large fleet of servers. The team is deeply technical and high-trust — you own platforms end to end from architecture definition through fleet operations.A day in the lifeYou start the day reviewing thermal and power validation data from an EVT build at your ODM partner. Mid-morning, you join a design review to close signal integrity findings on a high-speed accelerator interconnect. In the afternoon, you triage a fleet quality signal — correlating component-level failure data with manufacturing lot information to identify a systemic issue. You end the day aligning with architecture teams on requirements for the next-generation platform.About the teamThe Hardware Engineering AI/ML UltraServer platform team is a group of engineers and technical program managers directly responsible for launching GPU-accelerated servers into the AWS fleet. Located in Seattle, Austin, and Cupertino, we collaborate with global development teams and ODM partners to deliver next-generation AI/ML infrastructure deployed in datacenters worldwide. We move fast with small, empowered teams delivering end-to-end — from server conception through fleet-scale operations.Basic qualifications- Bachelor's degree in electrical engineering, computer engineering, or equivalent- Experience in developing functional specifications, design verification plans and functional test procedures- 7+ years of hardware design and development experience for server, compute, or large-scale infrastructure platforms- Experience in one or more server technologies: thermal/mechanical design, power delivery, high-speed signal integrity, or accelerator subsystems- Experience leading hardware development through full product lifecycle (concept through production ramp)Preferred qualification - Master's degree or PhD in Electrical Engineering, Computer Engineering, or a related field- 5+ years of experience working with ODMs through the product development and manufacturing lifecycle (EVT, DVT, PVT)- In-depth expertise in high-speed bus design, signal integrity analysis, or power delivery for GPU/accelerator platforms- 5+ years of experience with hardware bring-up, debug, and root cause analysis across PCIe, NVMe, memory, and accelerator interconnects- Experience owning fleet quality metrics and driving design improvements based on operational failure data- Experience with thermal/mechanical design for high-power-density compute platforms (liquid cooling, air cooling, or hybrid)- Experience working in large-scale datacenter or cloud environments- Track record of defining engineering standards and design best practices adopted across teams or partner organizationsAmazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at .USA, CA, Cupertino - 183,000.00 - 247,600.00 USD annuallyUSA, TX, Austin - 159,200.00 - 215,300.00 USD annuallyUSA, WA, Seattle - 159,200.00 - 215,300.00 USD annually
$183k - $247.6k
...the future of AI? Join the team... ...most advanced cloud for AI training... ...operate next-generation infrastructure... ...innovation in AI/ML and HPC... ...what’s next for AWS — and for the... ...of software, hardware, and network engineers, supply chain... ...high performance server and/or accelerator...Amazon Web ServiceCloudSeniorLocal areaFlexible hours$143k - $191k
...Eightfold is building the next generation of Agentic AI products that help... ...re looking for a Software Engineer to design and build highly... ...managers, designers, and AI/ML engineers to bring intelligent... ...Familiarity with cloud platforms (AWS, GCP, or Azure) Experience...Amazon Web ServiceCloudSeniorFull timeWork at officeRemote workFlexible hours$50k - $120k
...Mission Altimate AI, founded in 2022 in... ...the AI-powered data engineering revolution. You can... ...search of a Senior Generative AI Engineer who brings... ...+ years of hands-on ML/AI experience with... ...Develop and optimize cloud-native architectures (AWS, Kubernetes) for large...Amazon Web ServiceCloudFull timeWorldwide$162.7k - $220.2k
This position is part of the AWS Specialist and Partner Organization (... ...Specialist (PDS) to grow the AWS Generative Artificial Intelligence & Machine Learning (AI/ML) business through consulting, technology... ...and broadly adopted cloud platform. We pioneered cloud computing...Amazon Web ServiceCloudSeniorLocal areaFlexible hours$160k - $250k
...world’s most advanced AI-native platform. We... ...Development Engineer to join our AI Detection... ...Response (AIDR) Cloud team. In this role,... ...and build the next generation of AIDR services that... ...cloud environments (AWS/OCI/GCP/Azure)... ...understanding of AI/ML security challenges...Amazon Web ServiceCloudSeniorFull timeWork experience placementWork at officeLocal areaWorldwide2 days per week3 days per week- ...Job Title: Senior AI/ML Engineer Work Location with ZIP: Sunnyvale, CA 94085 (Hybrid... ...) - Hands?on experience with Generative AI / LLMs (OpenAI, Azure OpenAI, open... ...Experience building and deploying models in cloud environments (Azure / AWS / GCP) - Experience with MLOps (...Amazon Web ServiceCloudSenior
$183k - $247.6k
AWS Utility Computing (UC) provides product innovations... ...Amazon Elastic Compute Cloud (EC2), to consistently... ...world.We are seeking a Hardware Design Engineer with role in the definition... ...validation of AWS next generation ML Chips, Cards and server integration. As a senior...Amazon Web ServiceCloudSeniorLocal areaFlexible hours$150k - $250k
...interprets purchase intent, generates sourcing materials,... ...Role We’re seeking a Sr. AI Software Engineer with deep expertise in building... ...building production-grade AI/ML platforms or developer-... ...building and deploying on cloud platforms (AWS, Google Cloud, Azure); familiar...Amazon Web ServiceCloudSeniorFull timeWork at office$131k - $175k
...driven, client-to-cloud networking for large... ..., such as Best Engineering Team, Best Company... ...teams, including hardware, software, thermal... ...the world’s largest AI and cloud deployments... ...of next-generation rack architectures... ...or large-scale AI/ML cluster deploymentsExperience...CloudSeniorRemote workFlexible hours- ...Job Title: AI/ML Engineer Location: Sunnyvale, CA, USA... ...PayPal, Netflix, Meta, AWS, Product Companies, Startups... ...Do: Pilot next-generation technologies solving problems... ...Experience in cloud platforms AI Engineer... ...backend or full stack dev preferably in Python or...Amazon Web ServiceCloud
$220k - $350k
...enabling human life on Mars. SR. AI ENGINEER, SPECIAL PROGRAMS This... ...systems, APIs, or AI/ML applications (strong Python... ...with large language models, generative AI, or agentic systems—either... ...deployment Experience with cloud platforms (AWS, GCP, Azure),...Amazon Web ServiceCloudSeniorPermanent employmentFull timeTemporary workLocal areaImmediate startWeekend work$224k - $308k
...leader in materials engineering solutions used... ...world - like AI and IoT. If you... ...to create next generation technology,... ...Platform & Multi-Cloud AI InfrastructureEnable... ...AI Foundry, AWS Bedrock, and... ...Protocol) servers for internal tools... ...pipelines, CI/CD for ML, feature stores...Amazon Web ServiceCloudSeniorFull timeContract work- ...We are seeking expertise in Generative AI and Python to work directly... ...trusted advisor, hands-on engineer, and delivery lead — driving... ...in software engineering, ML engineering, or technical consulting... .... Familiarity with cloud platforms (AWS, Azure, GCP) and MLOps tools...Amazon Web ServiceCloudSeniorFull time
$169.8k - $233.5k
...of the largest B2B AI-native companies—decades... ...that combines Generative AI, Knowledge AI, Emotion... .... Job Description:SR AI Engineer Uniphore is a... .... Our Zero Data AI Cloud is built on a multimodal... ...machine learning (ML) and Generative AI... ...such as: Docker, AWS, or Kubernetes Good...Amazon Web ServiceCloudSeniorFull time$237.6k - $401.7k
...Sr Manager - Infrastructure, SRE, & AI Platforms - Services Special Projects... ...globally distributed engineering teams. In this role... .... Hybrid & Multi-Cloud Compute: Oversee internal... ...cloud platforms (AWS EKS, GCP GKE) and... ...scale systems for AI/ML training and inference...Amazon Web ServiceCloudSeniorRelocation$105k - $115k
...class end-to-end engineering solutions by leveraging... ...Do:Pilot next-generation technologies... ...by utilizing Gen AI or other machine... ...APIsExperience in cloud platformsAI Engineer/ ML Engineer with python... ...knowledge in aws or some cloud.Assess... ...backend or full stack dev preferably in...Amazon Web ServiceCloudTemporary work$203k - $258.6k
.../2026Meet the TeamCX AI Incubation team is part... ...work will power next-generation AI applications that... ...expertise in AI/ML, software development... ...product management and engineering teams to deliver impactful... ...on various AI cloud platforms such as AWS SageMaker, Google Cloud...Amazon Web ServiceCloudSeniorFull timeTemporary workLocal areaFlexible hours$100k - $200k
...and experienced backend engineer to join our growing... ...as the backbone for our generative AI-powered Android applications... ...platforms like Google Cloud or Azure to bring... ...closely with Android and ML teams on API contracts... ...major cloud platform (AWS, Azure, or GCP) ~ Experience...Amazon Web ServiceCloudFull time$227.5k - $300k
...transformation to AI-enabled software-defined... ...Senior Staff AI Engineer to join our team... ...innovations for next-generation software-defined... ...prototyping to global cloud deployment. This is... ..., traditional ML models, etc. (AI depth... ...cloud platforms (e.g., AWS, Azure, Google...Amazon Web ServiceCloudSeniorWork at officeLocal areaWorldwideFlexible hoursShift work$182k - $242k
Senior Software and AI Engineer Livingston, NJ / New York,... ...CoreWeave is The Essential Cloud for AI™. Built for... ...Solid familiarity with Generative AI frameworks ( LangChain... ...production-grade AI/ML/LLM-based applications,... ...of cloud environments (AWS, Azure, or GCP), container...Amazon Web ServiceCloudSeniorFull timeTemporary workCasual workWork at officeFlexible hours$193.3k - $261.5k
...experience as a mentor, tech lead, or engineering team lead We value a strong... ...someone to build virtualized, hardware-accelerated solutions for EC2... ...that powers modern cloud computing and AI/ML workloads Technologies: AI AWS Cloud EC2 Embedded...Amazon Web ServiceCloudSeniorFull timeInternshipLocal areaFlexible hours$148.7k - $201.2k
...compute capacity available to Generative AI customers? Do you want to solve... ...the boundary between physical hardware and software - at cloud scale?AWS Hardware Engineering is looking for a Systems... ...the health and development of server platforms at worldwide fleet scale...Amazon Web ServiceCloudInternshipLocal areaWorldwideFlexible hours$152k - $230k
...deep learning ignited modern AI — the next era of computing... ...us today.Design-for-Test Engineering at NVIDIA works on groundbreaking... ...methodologies for our next generation products using Gen AI... ...crucialHands-on experience with cloud platforms (AWS, Azure, GCP)Design and...Amazon Web ServiceCloudSeniorFull time- ...6+ years in Google Cloud Platform (GCP) and... ...enterprise-scale cloud and AI solutions. Job... ...implementation of Generative AI solutions. Key... ...and deploy AI/ML solutions leveraging... ...Mentor architects, engineers, and cloud practitioners... ...-cloud exposure (AWS / Azure) is an added...Amazon Web ServiceCloudSenior
$229.9k - $262.4k
...Sr. Lead AI Engineer (Gen AI Platform Services) Overview At... ...applications of AI and ML are bringing humanity... ...technologies such as AWS Ultraclusters, Huggingface... ...and your expertise in hardware, software, and AI... ...responsible AI solutions on cloud platforms (e.g. AWS,...Amazon Web ServiceCloudSeniorLocal area- ...is building the best AI systems for heavy industries... ...for Backend AI Engineers to design, build, and... ...that connect Generative AI, Computer Vision,... ...software engineering, ML infrastructure, and systems... ...Solid understanding of cloud infrastructure (AWS, GCP, or Azure) and...Amazon Web ServiceCloudFull time
$193.3k - $261.5k
We develop AWS Neuron, the complete software stack for Trainium, Amazon's custom cloud-scale machine learning accelerators. Join us to optimize the... ...to run really fast on the Trainium hardware.As a Sr. Software Development Engineer on the Inference Model Enablement team...Amazon Web ServiceCloudSeniorInternshipLocal areaFlexible hours$153.6k - $207.8k
AWS Global Sales drives adoption of the AWS cloud worldwide, enabling customers of all sizes to innovate and expand... ...their strategic platform for Generative AI and ML workloads? Do you enjoy... ...services. You will engage with senior engineers, architects, product leaders, data...Amazon Web ServiceCloudSeniorLocal areaWorldwideFlexible hours$193.3k - $261.5k
...to be part of AI revolution? At AWS our vision is to... ...innovative software and hardware solutions that... ...of complex ML models executed... ...software engineer in the Compiler... ...building next generation Neuron compiler... ...Trainium based servers in the Amazon cloud. You will be responsible...Amazon Web ServiceCloudSeniorLocal areaFlexible hours$162.7k - $220.2k
Amazon Web Services (AWS) is leading the... ...Data, Analytics and AI ISV partners. As a... ...the ISV's preferred cloud computing partner across... ...not limited to, AI/ML (Artificial... ...Learning), GenAI (Generative AI), Analytics, Database... ...administration, finance, engineering, human resources,...Amazon Web ServiceCloudSeniorContract workLocal areaFlexible hoursShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Sr Cloud Hardware Dev Engineer, AWS Generative AI & ML Servers. Be the first to apply!
- senior cloud data engineer Cupertino, CA
- cloud engineer Cupertino, CA
- aws cloud architect Cupertino, CA
- aws cloud security engineer Cupertino, CA
- cloud developer Cupertino, CA
- informatica cloud developer Cupertino, CA
- senior cloud security engineer Cupertino, CA
- software engineer - cloud services Cupertino, CA
- senior principal cloud computing engineer Cupertino, CA
- senior cloud network engineer Cupertino, CA



