Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Sr Cloud Hardware Dev Engineer, AWS Generative AI & ML Servers

Amazon Locker

AWS operates the world's largest fleet of GPU-accelerated servers powering AI/ML training and inference at cloud scale. Our team defines the server architectures, drives the hardware designs, and owns the fleet quality for these platforms — from component selection through datacenter operations. If you want to shape the physical hardware that frontier models train on, this is the role.We are seeking a Cloud Hardware Development Engineer to define server architectures based on workload demand, translate them into detailed component specifications, and drive validation from PCBA bring-up through rack integration. You will lead ODM design partners through development and production, triage hardware issues across manufacturing and datacenters, and own fleet quality metrics post-launch.What You Will DoYou will define the hardware that runs the world's largest AI training workloads. Your designs span thermal, mechanical, power, and signal integrity across GPU-accelerated platforms. You will drive validation from first silicon through fleet-scale deployment, triage failures correlating across PCIe, power delivery, memory, and accelerator interconnects, and feed root cause findings back into design improvements. When a new server platform launches at a large scale, the architecture, component choices, and quality gates are yours.Key job responsibilitiesArchitecture & Design* Define server architectures based on workload demand and customer requirements, translating them into detailed designs and component specifications that enable high-performance AI training and inference at scale* Work with interdisciplinary teams of component, firmware, test, qualification, and integration engineers to deliver cohesive designs* Drive design reviews with ODM/JDM partners covering schematic, layout, BOM, and manufacturing DFx (Design for Test, Design for Manufacturing)Validation & Bring-up* Define and execute validation strategies from PCBA bring-up through server and rack integration — covering power sequencing, signal integrity, thermal characterization, and accelerator interconnect performance* Own hardware debug during EVT/DVT/PVT builds, correlating failures across PCIe, power rails, memory channels, and GPU subsystems* Triage hardware issues at both ODM facilities and datacenters, conduct root cause analysis, and implement corrective actionsFleet Quality & Continuous Improvement* Own fleet quality metrics post-launch: server-level annualized failure rates and component-level failure modes* Monitor operational telemetry to identify systemic issues and drive design or process changes for current and future platforms* Partner with test and automation teams to improve manufacturing yield and reduce test dwell timesCross-Team Collaboration* Work with EC2 architecture teams to align on instance definitions, workload requirements, and platform trade-offs* Drive ODM/JDM design partners through development milestones and production ramp* Collaborate with firmware, software, and operations teams to ensure designs are debuggable, serviceable, and automation-readyMay require occasional (<10%) regional and international travel to Design and Manufacturing Partner sites.The Ideal CandidateYou think across the full hardware stack — from silicon packaging and power delivery to rack-level thermal and mechanical design. You are as comfortable reviewing a schematic as you are analyzing fleet failure data. You drive quality through data, not assumption, and you hold design partners to the same standard you hold yourself. You mentor and develop junior engineers, contribute to hiring, and share your expertise to make the team stronger.Why You Will Love ItThe world's most advanced frontier models train on the hardware you design. You will see your architecture decisions scale to a large fleet of servers. The team is deeply technical and high-trust — you own platforms end to end from architecture definition through fleet operations.A day in the lifeYou start the day reviewing thermal and power validation data from an EVT build at your ODM partner. Mid-morning, you join a design review to close signal integrity findings on a high-speed accelerator interconnect. In the afternoon, you triage a fleet quality signal — correlating component-level failure data with manufacturing lot information to identify a systemic issue. You end the day aligning with architecture teams on requirements for the next-generation platform.About the teamThe Hardware Engineering AI/ML UltraServer platform team is a group of engineers and technical program managers directly responsible for launching GPU-accelerated servers into the AWS fleet. Located in Seattle, Austin, and Cupertino, we collaborate with global development teams and ODM partners to deliver next-generation AI/ML infrastructure deployed in datacenters worldwide. We move fast with small, empowered teams delivering end-to-end — from server conception through fleet-scale operations.Basic qualifications- Bachelor's degree in electrical engineering, computer engineering, or equivalent- Experience in developing functional specifications, design verification plans and functional test procedures- 7+ years of hardware design and development experience for server, compute, or large-scale infrastructure platforms- Experience in one or more server technologies: thermal/mechanical design, power delivery, high-speed signal integrity, or accelerator subsystems- Experience leading hardware development through full product lifecycle (concept through production ramp)Preferred qualification - Master's degree or PhD in Electrical Engineering, Computer Engineering, or a related field- 5+ years of experience working with ODMs through the product development and manufacturing lifecycle (EVT, DVT, PVT)- In-depth expertise in high-speed bus design, signal integrity analysis, or power delivery for GPU/accelerator platforms- 5+ years of experience with hardware bring-up, debug, and root cause analysis across PCIe, NVMe, memory, and accelerator interconnects- Experience owning fleet quality metrics and driving design improvements based on operational failure data- Experience with thermal/mechanical design for high-power-density compute platforms (liquid cooling, air cooling, or hybrid)- Experience working in large-scale datacenter or cloud environments- Track record of defining engineering standards and design best practices adopted across teams or partner organizationsAmazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at .USA, CA, Cupertino - 183,000.00 - 247,600.00 USD annuallyUSA, TX, Austin - 159,200.00 - 215,300.00 USD annuallyUSA, WA, Seattle - 159,200.00 - 215,300.00 USD annually

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Sr Cloud Hardware Dev Engineer, AWS Generative AI & ML Servers in Cupertino, CA vacancy
  • $183k - $247.6k

     ...the future of AI? Join the team...  ...most advanced cloud for AI training...  ...operate next-generation infrastructure...  ...innovation in AI/ML and HPC...  ...what’s next for AWS — and for the...  ...of software, hardware, and network engineers, supply chain...  ...high performance server and/or accelerator... 
    Amazon Web Service
    Cloud
    Senior
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $143k - $191k

     ...Eightfold is building the next generation of Agentic AI products that help...  ...re looking for a Software Engineer to design and build highly...  ...managers, designers, and AI/ML engineers to bring intelligent...  ...Familiarity with cloud platforms (AWS, GCP, or Azure) Experience... 
    Amazon Web Service
    Cloud
    Senior
    Full time
    Work at office
    Remote work
    Flexible hours

    Eightfold

    Santa Clara, CA
    a month ago
  • $50k - $120k

     ...Mission Altimate AI, founded in 2022 in...  ...the AI-powered data engineering revolution. You can...  ...search of a Senior Generative AI Engineer who brings...  ...+ years of hands-on ML/AI experience with...  ...Develop and optimize cloud-native architectures (AWS, Kubernetes) for large... 
    Amazon Web Service
    Cloud
    Full time
    Worldwide

    Pa Early Stage Partners

    Sunnyvale, CA
    more than 2 months ago
  • $162.7k - $220.2k

    This position is part of the AWS Specialist and Partner Organization (...  ...Specialist (PDS) to grow the AWS Generative Artificial Intelligence & Machine Learning (AI/ML) business through consulting, technology...  ...and broadly adopted cloud platform. We pioneered cloud computing... 
    Amazon Web Service
    Cloud
    Senior
    Local area
    Flexible hours

    AmazonWebServices

    Mountain View, CA
    a month ago
  • $160k - $250k

     ...world’s most advanced AI-native platform. We...  ...Development Engineer to join our AI Detection...  ...Response (AIDR) Cloud team. In this role,...  ...and build the next generation of AIDR services that...  ...cloud environments (AWS/OCI/GCP/Azure)...  ...understanding of AI/ML security challenges... 
    Amazon Web Service
    Cloud
    Senior
    Full time
    Work experience placement
    Work at office
    Local area
    Worldwide
    2 days per week
    3 days per week

    CrowdStrike

    Sunnyvale, CA
    a month ago
  •  ...Job Title: Senior AI/ML Engineer Work Location with ZIP: Sunnyvale, CA 94085 (Hybrid...  ...) - Hands?on experience with Generative AI / LLMs (OpenAI, Azure OpenAI, open...  ...Experience building and deploying models in cloud environments (Azure / AWS / GCP) - Experience with MLOps (... 
    Amazon Web Service
    Cloud
    Senior

    eTeam

    Sunnyvale, CA
    1 day ago
  • $183k - $247.6k

    AWS Utility Computing (UC) provides product innovations...  ...Amazon Elastic Compute Cloud (EC2), to consistently...  ...world.We are seeking a Hardware Design Engineer with role in the definition...  ...validation of AWS next generation ML Chips, Cards and server integration. As a senior... 
    Amazon Web Service
    Cloud
    Senior
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  • $150k - $250k

     ...interprets purchase intent, generates sourcing materials,...  ...Role We’re seeking a Sr. AI Software Engineer with deep expertise in building...  ...building production-grade AI/ML platforms or developer-...  ...building and deploying on cloud platforms (AWS, Google Cloud, Azure); familiar... 
    Amazon Web Service
    Cloud
    Senior
    Full time
    Work at office

    Globality, Inc

    Palo Alto, CA
    26 days ago
  • $131k - $175k

     ...driven, client-to-cloud networking for large...  ..., such as Best Engineering Team, Best Company...  ...teams, including hardware, software, thermal...  ...the world’s largest AI and cloud deployments...  ...of next-generation rack architectures...  ...or large-scale AI/ML cluster deploymentsExperience... 
    Cloud
    Senior
    Remote work
    Flexible hours

    Arista Networks

    Santa Clara, CA
    a month ago
  •  ...Job Title: AI/ML Engineer Location: Sunnyvale, CA, USA...  ...PayPal, Netflix, Meta, AWS, Product Companies, Startups...  ...Do: Pilot next-generation technologies solving problems...  ...Experience in cloud platforms AI Engineer...  ...backend or full stack dev preferably in Python or... 
    Amazon Web Service
    Cloud

    RIT Solutions, Inc.

    Sunnyvale, CA
    13 hours ago
  • $220k - $350k

     ...enabling human life on Mars. SR. AI ENGINEER, SPECIAL PROGRAMS This...  ...systems, APIs, or AI/ML applications (strong Python...  ...with large language models, generative AI, or agentic systems—either...  ...deployment Experience with cloud platforms (AWS, GCP, Azure),... 
    Amazon Web Service
    Cloud
    Senior
    Permanent employment
    Full time
    Temporary work
    Local area
    Immediate start
    Weekend work

    Spacex

    Palo Alto, CA
    more than 2 months ago
  • $224k - $308k

     ...leader in materials engineering solutions used...  ...world - like AI and IoT. If you...  ...to create next generation technology,...  ...Platform & Multi-Cloud AI InfrastructureEnable...  ...AI Foundry, AWS Bedrock, and...  ...Protocol) servers for internal tools...  ...pipelines, CI/CD for ML, feature stores... 
    Amazon Web Service
    Cloud
    Senior
    Full time
    Contract work

    Applied Materials

    Santa Clara, CA
    a month ago
  •  ...We are seeking expertise in Generative AI and Python to work directly...  ...trusted advisor, hands-on engineer, and delivery lead — driving...  ...in software engineering, ML engineering, or technical consulting...  .... Familiarity with cloud platforms (AWS, Azure, GCP) and MLOps tools... 
    Amazon Web Service
    Cloud
    Senior
    Full time

    Infinite Computer Solutions

    Sunnyvale, CA
    6 days ago
  • $169.8k - $233.5k

     ...of the largest B2B AI-native companies—decades...  ...that combines Generative AI, Knowledge AI, Emotion...  .... Job Description:SR AI Engineer Uniphore is a...  .... Our Zero Data AI Cloud is built on a multimodal...  ...machine learning (ML) and Generative AI...  ...such as: Docker, AWS, or Kubernetes Good... 
    Amazon Web Service
    Cloud
    Senior
    Full time

    Uniphore

    Palo Alto, CA
    a month ago
  • $237.6k - $401.7k

     ...Sr Manager - Infrastructure, SRE, & AI Platforms - Services Special Projects...  ...globally distributed engineering teams. In this role...  .... Hybrid & Multi-Cloud Compute: Oversee internal...  ...cloud platforms (AWS EKS, GCP GKE) and...  ...scale systems for AI/ML training and inference... 
    Amazon Web Service
    Cloud
    Senior
    Relocation

    Apple

    Cupertino, CA
    6 days ago
  • $105k - $115k

     ...class end-to-end engineering solutions by leveraging...  ...Do:Pilot next-generation technologies...  ...by utilizing Gen AI or other machine...  ...APIsExperience in cloud platformsAI Engineer/ ML Engineer with python...  ...knowledge in aws or some cloud.Assess...  ...backend or full stack dev preferably in... 
    Amazon Web Service
    Cloud
    Temporary work

    Quest Global Services

    Sunnyvale, CA
    a month ago
  • $203k - $258.6k

     .../2026Meet the TeamCX AI Incubation team is part...  ...work will power next-generation AI applications that...  ...expertise in AI/ML, software development...  ...product management and engineering teams to deliver impactful...  ...on various AI cloud platforms such as AWS SageMaker, Google Cloud... 
    Amazon Web Service
    Cloud
    Senior
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    5 days ago
  • $100k - $200k

     ...and experienced backend engineer to join our growing...  ...as the backbone for our generative AI-powered Android applications...  ...platforms like Google Cloud or Azure to bring...  ...closely with Android and ML teams on API contracts...  ...major cloud platform (AWS, Azure, or GCP) ~ Experience... 
    Amazon Web Service
    Cloud
    Full time

    OPPO US Research Center

    Palo Alto, CA
    a month ago
  • $227.5k - $300k

     ...transformation to AI-enabled software-defined...  ...Senior Staff AI Engineer   to join our team...  ...innovations for next-generation software-defined...  ...prototyping to global cloud deployment. This is...  ..., traditional ML models, etc. (AI depth...  ...cloud platforms (e.g., AWS, Azure, Google... 
    Amazon Web Service
    Cloud
    Senior
    Work at office
    Local area
    Worldwide
    Flexible hours
    Shift work

    Sonatus

    Sunnyvale, CA
    26 days ago
  • $182k - $242k

    Senior Software and AI Engineer Livingston, NJ / New York,...  ...CoreWeave is The Essential Cloud for AI™. Built for...  ...Solid familiarity with Generative AI frameworks ( LangChain...  ...production-grade AI/ML/LLM-based applications,...  ...of cloud environments (AWS, Azure, or GCP), container... 
    Amazon Web Service
    Cloud
    Senior
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    2 days ago
  • $193.3k - $261.5k

     ...experience as a mentor, tech lead, or engineering team lead We value a strong...  ...someone to build virtualized, hardware-accelerated solutions for EC2...  ...that powers modern cloud computing and AI/ML workloads Technologies: AI AWS Cloud EC2 Embedded... 
    Amazon Web Service
    Cloud
    Senior
    Full time
    Internship
    Local area
    Flexible hours

    Amazon.com Services LLC

    Santa Clara, CA
    2 days ago
  • $148.7k - $201.2k

     ...compute capacity available to Generative AI customers? Do you want to solve...  ...the boundary between physical hardware and software - at cloud scale?AWS Hardware Engineering is looking for a Systems...  ...the health and development of server platforms at worldwide fleet scale... 
    Amazon Web Service
    Cloud
    Internship
    Local area
    Worldwide
    Flexible hours

    Amazon

    Cupertino, CA
    15 days ago
  • $152k - $230k

     ...deep learning ignited modern AI — the next era of computing...  ...us today.Design-for-Test Engineering at NVIDIA works on groundbreaking...  ...methodologies for our next generation products using Gen AI...  ...crucialHands-on experience with cloud platforms (AWS, Azure, GCP)Design and... 
    Amazon Web Service
    Cloud
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  •  ...6+ years in Google Cloud Platform (GCP) and...  ...enterprise-scale cloud and AI solutions. Job...  ...implementation of Generative AI solutions. Key...  ...and deploy AI/ML solutions leveraging...  ...Mentor architects, engineers, and cloud practitioners...  ...-cloud exposure (AWS / Azure) is an added... 
    Amazon Web Service
    Cloud
    Senior

    Diamondpick

    Milpitas, CA
    4 days ago
  • $229.9k - $262.4k

     ...Sr. Lead AI Engineer (Gen AI Platform Services) Overview At...  ...applications of AI and ML are bringing humanity...  ...technologies such as AWS Ultraclusters, Huggingface...  ...and your expertise in hardware, software, and AI...  ...responsible AI solutions on cloud platforms (e.g. AWS,... 
    Amazon Web Service
    Cloud
    Senior
    Local area

    Capital One National Association

    San Jose, CA
    2 days ago
  •  ...is building the best AI systems for heavy industries...  ...for Backend AI Engineers to design, build, and...  ...that connect Generative AI, Computer Vision,...  ...software engineering, ML infrastructure, and systems...  ...Solid understanding of cloud infrastructure (AWS, GCP, or Azure) and... 
    Amazon Web Service
    Cloud
    Full time

    Nexxa.ai

    Sunnyvale, CA
    a month ago
  • $193.3k - $261.5k

    We develop AWS Neuron, the complete software stack for Trainium, Amazon's custom cloud-scale machine learning accelerators. Join us to optimize the...  ...to run really fast on the Trainium hardware.As a Sr. Software Development Engineer on the Inference Model Enablement team... 
    Amazon Web Service
    Cloud
    Senior
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    a month ago
  • $153.6k - $207.8k

    AWS Global Sales drives adoption of the AWS cloud worldwide, enabling customers of all sizes to innovate and expand...  ...their strategic platform for Generative AI and ML workloads? Do you enjoy...  ...services. You will engage with senior engineers, architects, product leaders, data... 
    Amazon Web Service
    Cloud
    Senior
    Local area
    Worldwide
    Flexible hours

    AmazonWebServices

    Mountain View, CA
    26 days ago
  • $193.3k - $261.5k

     ...to be part of AI revolution? At AWS our vision is to...  ...innovative software and hardware solutions that...  ...of complex ML models executed...  ...software engineer in the Compiler...  ...building next generation Neuron compiler...  ...Trainium based servers in the Amazon cloud. You will be responsible... 
    Amazon Web Service
    Cloud
    Senior
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    a month ago
  • $162.7k - $220.2k

    Amazon Web Services (AWS) is leading the...  ...Data, Analytics and AI ISV partners. As a...  ...the ISV's preferred cloud computing partner across...  ...not limited to, AI/ML (Artificial...  ...Learning), GenAI (Generative AI), Analytics, Database...  ...administration, finance, engineering, human resources,... 
    Amazon Web Service
    Cloud
    Senior
    Contract work
    Local area
    Flexible hours
    Shift work

    AmazonWebServices

    Mountain View, CA
    a month ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Sr Cloud Hardware Dev Engineer, AWS Generative AI & ML Servers. Be the first to apply!