Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Sr Cloud Hardware Dev Engineer, AWS Generative AI & ML Servers

$183k - $247.6k

Amazon

AWS operates the world's largest fleet of GPU-accelerated servers powering AI/ML training and inference at cloud scale. Our team defines the server architectures, drives the hardware designs, and owns the fleet quality for these platforms — from component selection through datacenter operations. If you want to shape the physical hardware that frontier models train on, this is the role.We are seeking a Cloud Hardware Development Engineer to define server architectures based on workload demand, translate them into detailed component specifications, and drive validation from PCBA bring-up through rack integration. You will lead ODM design partners through development and production, triage hardware issues across manufacturing and datacenters, and own fleet quality metrics post-launch.What You Will DoYou will define the hardware that runs the world's largest AI training workloads. Your designs span thermal, mechanical, power, and signal integrity across GPU-accelerated platforms. You will drive validation from first silicon through fleet-scale deployment, triage failures correlating across PCIe, power delivery, memory, and accelerator interconnects, and feed root cause findings back into design improvements. When a new server platform launches at a large scale, the architecture, component choices, and quality gates are yours.Key job responsibilitiesArchitecture & Design* Define server architectures based on workload demand and customer requirements, translating them into detailed designs and component specifications that enable high-performance AI training and inference at scale* Work with interdisciplinary teams of component, firmware, test, qualification, and integration engineers to deliver cohesive designs* Drive design reviews with ODM/JDM partners covering schematic, layout, BOM, and manufacturing DFx (Design for Test, Design for Manufacturing)Validation & Bring-up* Define and execute validation strategies from PCBA bring-up through server and rack integration — covering power sequencing, signal integrity, thermal characterization, and accelerator interconnect performance* Own hardware debug during EVT/DVT/PVT builds, correlating failures across PCIe, power rails, memory channels, and GPU subsystems* Triage hardware issues at both ODM facilities and datacenters, conduct root cause analysis, and implement corrective actionsFleet Quality & Continuous Improvement* Own fleet quality metrics post-launch: server-level annualized failure rates and component-level failure modes* Monitor operational telemetry to identify systemic issues and drive design or process changes for current and future platforms* Partner with test and automation teams to improve manufacturing yield and reduce test dwell timesCross-Team Collaboration* Work with EC2 architecture teams to align on instance definitions, workload requirements, and platform trade-offs* Drive ODM/JDM design partners through development milestones and production ramp* Collaborate with firmware, software, and operations teams to ensure designs are debuggable, serviceable, and automation-readyMay require occasional (<10%) regional and international travel to Design and Manufacturing Partner sites.The Ideal CandidateYou think across the full hardware stack — from silicon packaging and power delivery to rack-level thermal and mechanical design. You are as comfortable reviewing a schematic as you are analyzing fleet failure data. You drive quality through data, not assumption, and you hold design partners to the same standard you hold yourself. You mentor and develop junior engineers, contribute to hiring, and share your expertise to make the team stronger.Why You Will Love ItThe world's most advanced frontier models train on the hardware you design. You will see your architecture decisions scale to a large fleet of servers. The team is deeply technical and high-trust — you own platforms end to end from architecture definition through fleet operations.A day in the lifeYou start the day reviewing thermal and power validation data from an EVT build at your ODM partner. Mid-morning, you join a design review to close signal integrity findings on a high-speed accelerator interconnect. In the afternoon, you triage a fleet quality signal — correlating component-level failure data with manufacturing lot information to identify a systemic issue. You end the day aligning with architecture teams on requirements for the next-generation platform.About the teamThe Hardware Engineering AI/ML UltraServer platform team is a group of engineers and technical program managers directly responsible for launching GPU-accelerated servers into the AWS fleet. Located in Seattle, Austin, and Cupertino, we collaborate with global development teams and ODM partners to deliver next-generation AI/ML infrastructure deployed in datacenters worldwide. We move fast with small, empowered teams delivering end-to-end — from server conception through fleet-scale operations.Basic qualifications- Bachelor's degree in electrical engineering, computer engineering, or equivalent- Experience in developing functional specifications, design verification plans and functional test procedures- 7+ years of hardware design and development experience for server, compute, or large-scale infrastructure platforms- Experience in one or more server technologies: thermal/mechanical design, power delivery, high-speed signal integrity, or accelerator subsystems- Experience leading hardware development through full product lifecycle (concept through production ramp)Preferred qualification - Master's degree or PhD in Electrical Engineering, Computer Engineering, or a related field- 5+ years of experience working with ODMs through the product development and manufacturing lifecycle (EVT, DVT, PVT)- In-depth expertise in high-speed bus design, signal integrity analysis, or power delivery for GPU/accelerator platforms- 5+ years of experience with hardware bring-up, debug, and root cause analysis across PCIe, NVMe, memory, and accelerator interconnects- Experience owning fleet quality metrics and driving design improvements based on operational failure data- Experience with thermal/mechanical design for high-power-density compute platforms (liquid cooling, air cooling, or hybrid)- Experience working in large-scale datacenter or cloud environments- Track record of defining engineering standards and design best practices adopted across teams or partner organizationsAmazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at .USA, CA, Cupertino - 183,000.00 - 247,600.00 USD annuallyUSA, TX, Austin - 159,200.00 - 215,300.00 USD annuallyUSA, WA, Seattle - 159,200.00 - 215,300.00 USD annually

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Sr Cloud Hardware Dev Engineer, AWS Generative AI & ML Servers in Cupertino, CA vacancy
  • $157.3k - $212.8k

     ...to build the backbone of Generative AI cloud at AWS? Do you want to build the...  ...performance and scalability in AI/ML and HPC workloads.Utility...  ...team of software, hardware, and network engineers, supply chain specialists...  ...segment of accelerated servers.You will work closely with... 
    Amazon Web Service
    Cloud
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    3 hours ago
  • $122.6k - $185k

     ...you want to build the backbone of Generative AI cloud at AWS? Do you want to build the future...  ...performance and scalability in AI/ML and HPC workloads.You are...  ...looking for builders like you. The AWS Hardware Engineering team creates server designs for Amazon’s innovative web... 
    Amazon Web Service
    Cloud
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    4 days ago
  • $183k - $247.6k

     ...basisAWS Compute & ML Services owns the...  ...operation of all AWS global...  ...people who keep the cloud running. We support...  ...centers and all of the servers, storage,...  ...team of software, hardware, and network engineers, supply chain specialists...  ...of next generation storage (SSD) for... 
    Amazon Web Service
    Cloud
    Senior
    Local area
    Flexible hours

    AmazonWebServices

    Cupertino, CA
    1 day ago
  • $148.7k - $201.2k

     ...build the backbone of Generative AI at AWS? Do you want to build the future of the cloud for AI training and...  ...Systems Development Engineer to develop automation...  ...our accelerated (AI/ML) server platforms. You will work...  ...using a combination of hardware, software, system... 
    Amazon Web Service
    Cloud
    Internship
    Local area
    Worldwide
    Flexible hours

    Amazon

    Cupertino, CA
    3 hours ago
  • $143k - $191k

     ...Eightfold is building the next generation of Agentic AI products that help...  ...re looking for a Software Engineer to design and build highly...  ...managers, designers, and AI/ML engineers to bring intelligent...  ...Familiarity with cloud platforms (AWS, GCP, or Azure) Experience... 
    Amazon Web Service
    Cloud
    Senior
    Full time
    Work at office
    Remote work
    Flexible hours

    Eightfold

    Santa Clara, CA
    1 day ago
  • $162.7k - $220.2k

    This position is part of the AWS Specialist and Partner Organization (...  ...Specialist (PDS) to grow the AWS Generative Artificial Intelligence & Machine Learning (AI/ML) business through consulting, technology...  ...and broadly adopted cloud platform. We pioneered cloud computing... 
    Amazon Web Service
    Cloud
    Senior
    Local area
    Flexible hours

    AmazonWebServices

    Mountain View, CA
    4 days ago
  • $192.2k - $260k

    Come join the AWS AI science team in building the next generation models for intelligent automation....  ...world-leading provider of cloud services, has fostered...  ...resources, and to world-class engineers and developers that can...  ...our dynamic team of AI/ML practitioners, applied... 
    Amazon Web Service
    Cloud
    Senior
    Local area
    Immediate start
    Worldwide
    Flexible hours

    Amazon

    Santa Clara, CA
    1 day ago
  • $175k - $236.8k

    AWS Trainium is deployed at scale, with...  ...deep learning and generative AI workloads with optimal...  ...the training AI/ML ecosystem and what...  ...will partner with engineering teams building...  ...'t achieve in the cloud.About Amazon Annapurna...  ...engineering, hardware design, software and... 
    Amazon Web Service
    Cloud
    Senior
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    3 hours ago
  • $183k - $247.6k

    AWS Utility Computing (UC) provides product innovations...  ...Amazon Elastic Compute Cloud (EC2), to consistently...  ...world.We are seeking a Hardware Design Engineer with role in the definition...  ...validation of AWS next generation ML Chips, Cards and server integration. As a senior... 
    Amazon Web Service
    Cloud
    Senior
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $193.3k - $261.5k

     ...serving as a tech lead, or leading an engineering team ~ Masters degree in computer science...  ...inference stack Technologies: AWS C# Cloud Java Machine Learning vLLM Architect DevOps GitHub Hardware Support More: We develop AWS Neuron... 
    Amazon Web Service
    Cloud
    Senior
    Full time
    Internship

    Annapurna Labs (U.S.) Inc.

    Cupertino, CA
    4 days ago
  • $140k - $215k

     ...world’s most advanced AI-native platform. We...  ...Development Engineer to join our AI Detection...  ...Response (AIDR) Cloud team. In this role,...  ...and build the next generation of AIDR services that...  ...cloud environments (AWS/OCI/GCP/Azure)...  ...understanding of AI/ML security challenges... 
    Amazon Web Service
    Cloud
    Senior
    Full time
    Work experience placement
    Work at office
    Local area
    Worldwide
    2 days per week
    3 days per week

    CrowdStrike

    Sunnyvale, CA
    3 hours ago
  • $50k - $120k

     ...Mission Altimate AI, founded in 2022 in...  ...the AI-powered data engineering revolution. You can...  ...search of a Senior Generative AI Engineer who brings...  ...+ years of hands-on ML/AI experience with...  ...Develop and optimize cloud-native architectures (AWS, Kubernetes) for large... 
    Amazon Web Service
    Cloud
    Full time
    Worldwide

    Pa Early Stage Partners

    Sunnyvale, CA
    1 day ago
  • $131k - $175k

     ...driven, client-to-cloud networking for large...  ..., such as Best Engineering Team, Best Company...  ...teams, including hardware, software, thermal...  ...the world’s largest AI and cloud deployments...  ...of next-generation rack architectures...  ...or large-scale AI/ML cluster deploymentsExperience... 
    Cloud
    Senior
    Remote work
    Flexible hours

    Arista Networks

    Santa Clara, CA
    3 days ago
  • $139.23k - $163.8k

     ...DescriptionLead Software Architect Engineer (Generative AI Platforms) is responsible for...  ...ready AI applications across cloud environments while driving...  ...platforms across Azure and AWS, ensuring high availability,...  ...Experience deploying and managing AI/ML workloads in Azure and/or AWS... 
    Amazon Web Service
    Cloud
    Full time
    Work experience placement
    Local area
    3 days per week

    US Bank

    Cupertino, CA
    2 days ago
  • $162.7k - $220.2k

     ...necessary to help position AWS as the cloud provider of choice for...  ...? Join the Data & AI team as a Go-To-Market...  ..., retrieval-augmented generation (RAG), and generative...  ...closely with product and engineering teams to translate...  ...search technologies and AI/ML to create best-in-... 
    Amazon Web Service
    Cloud
    Senior
    Local area
    Worldwide
    Flexible hours
    Day shift

    Amazon Web Services, Inc.

    Mountain View, CA
    3 hours ago
  • $184k - $287.5k

     ...motivated software engineers to join us and build AI inference systems...  ...-node, and multi-cloud environments. You...  ...NVIDIA GPU hardware features; profile...  ...tuned and compiler-generated) using techniques...  ...for the field of ML Systems; survey recent...  ...cloud platforms (AWS/GCP/Azure),... 
    Amazon Web Service
    Cloud
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $224k - $308k

     ...leader in materials engineering solutions used...  ...world - like AI and IoT. If you...  ...to create next generation technology,...  ...Platform & Multi-Cloud AI InfrastructureEnable...  ...AI Foundry, AWS Bedrock, and...  ...Protocol) servers for internal tools...  ...pipelines, CI/CD for ML, feature stores... 
    Amazon Web Service
    Cloud
    Senior
    Full time
    Contract work

    Applied Materials

    Santa Clara, CA
    4 days ago
  • $150k - $250k

     ...Sr. AI Engineer Palo Alto, CA Globality is the autonomous sourcing...  ...purchase intent, generates sourcing materials, evaluates...  ...building production-grade AI/ML platforms or developer-centric...  ...building and deploying on cloud platforms (AWS, Google Cloud, Azure); familiar... 
    Amazon Web Service
    Cloud
    Senior
    Work at office

    Globality Inc

    Palo Alto, CA
    2 days ago
  • $148.7k - $240.53k

     ...Inclusion. We weave AI into the fabric of...  ...building innovative, AI/ML-powered security...  ...private and public clouds, directly shaping...  ...Collaborate proactively with engineering and DevOps teams to...  ....Drive the next generation of Prisma AIRS...  ...such as AWS bedrock, Azure Foundry... 
    Amazon Web Service
    Cloud
    Senior
    Full time
    Work at office

    Palo Alto Networks

    Santa Clara, CA
    3 hours ago
  • $169.8k - $233.5k

     ...of the largest B2B AI-native companies—decades...  ...that combines Generative AI, Knowledge AI, Emotion...  .... Job Description:SR AI Engineer Uniphore is a...  .... Our Zero Data AI Cloud is built on a multimodal...  ...machine learning (ML) and Generative AI...  ...such as: Docker, AWS, or Kubernetes Good... 
    Amazon Web Service
    Cloud
    Senior
    Full time

    Uniphore

    Palo Alto, CA
    1 day ago
  • $220k - $350k

     ...enabling human life on Mars.SR. AI ENGINEER, SPECIAL PROGRAMSThis team...  ...scalable systems, APIs, or AI/ML applications (strong Python...  ...large language models, generative AI, or agentic systems—either...  ...deploymentExperience with cloud platforms (AWS, GCP, Azure), containerization... 
    Amazon Web Service
    Cloud
    Senior
    Permanent employment
    Temporary work
    Local area
    Immediate start
    Weekend work

    SpaceX

    Palo Alto, CA
    3 hours ago
  •  ...We are seeking expertise in Generative AI and Python to work directly...  ...trusted advisor, hands-on engineer, and delivery lead — driving...  ...in software engineering, ML engineering, or technical consulting...  .... Familiarity with cloud platforms (AWS, Azure, GCP) and MLOps tools... 
    Amazon Web Service
    Cloud
    Senior
    Full time

    Infinite Computer Solutions

    Sunnyvale, CA
    1 day ago
  • $227.5k - $300k

     ...transformation to AI-enabled software-defined...  ...a Senior Staff AI Engineer with a combination...  ...to audit AI-generated actions before executionConduct...  ..., traditional ML models, etc. (AI...  ...whiteboards to global cloud deployments.Strong...  ...platforms (e.g., AWS, Azure, Google Cloud... 
    Amazon Web Service
    Cloud
    Senior
    Work at office
    Worldwide
    Flexible hours
    Shift work

    Sonatus

    Sunnyvale, CA
    1 day ago
  • $105k - $115k

     ...class end-to-end engineering solutions by leveraging...  ...Do:Pilot next-generation technologies...  ...by utilizing Gen AI or other machine...  ...APIsExperience in cloud platformsAI Engineer/ ML Engineer with python...  ...knowledge in aws or some cloud.Assess...  ...backend or full stack dev preferably in... 
    Amazon Web Service
    Cloud
    Temporary work

    Quest Global Services

    Sunnyvale, CA
    4 days ago
  • $153.2k - $234.1k

     ...intelligent automation, AI-enabled engineering workflows, and...  ..., QA, and AI/ML teams to design and...  ...simulation and hardware-in-loop testing,...  ...observability platforms, and cloud services into...  ...to evaluate AI-generated outputs for...  ...metrics pipelines in AWS, GCP, Azure, or... 
    Amazon Web Service
    Cloud
    Senior
    Full time
    Local area
    Work from home
    Flexible hours

    General Motors

    Sunnyvale, CA
    2 days ago
  •  ...Position: Agentic AI-Sr Architect Location: Santa Clara...  ...degree in Computer Science, AI/ML, Data Science, or related...  ...5 years of experience in cloud AI platforms (AWS Sagemaker, Azure ML, GCP AI)...  ...requirements Collaborate with data engineers, LLM Ops, and software teams... 
    Amazon Web Service
    Cloud
    Senior
    Full time

    Noblesoft Technologies

    Santa Clara, CA
    3 days ago
  • $203k - $258.6k

     .../2026Meet the TeamCX AI Incubation team is part...  ...work will power next-generation AI applications that...  ...expertise in AI/ML, software development...  ...product management and engineering teams to deliver impactful...  ...on various AI cloud platforms such as AWS SageMaker, Google Cloud... 
    Amazon Web Service
    Cloud
    Senior
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    3 days ago
  • $193.3k - $261.5k

     ..., tech lead, or engineering manager. We require...  ...Experience with ML communications...  ...plus. Prior AI/ML experience is...  ...accelerators, servers, and...  ...performance toward the hardware roofline by profiling...  ...workloads on AWS. We collaborate...  ...Network RDMA Cloud LLM More:... 
    Amazon Web Service
    Cloud
    Senior
    Full time
    Internship
    Flexible hours

    Amazon.com Services LLC

    Cupertino, CA
    3 days ago
  • $105k - $115k

     ...class end-to-end engineering solutions by...  ...Pilot next-generation technologies solving...  ...by utilizing Gen AI or other machine...  ...Experience in cloud platforms AI Engineer/ ML Engineer with python...  ...some knowledge in aws or some cloud....  ...backend or full stack dev preferably in... 
    Amazon Web Service
    Cloud
    Temporary work

    QuEST Global

    Sunnyvale, CA
    3 days ago
  • $206.9k - $279.9k

    AWS Neuron is looking for an experienced Technical...  ...best-in-class ML performance in the cloud. You will lead NKI requirements...  ...optimization, and hardware acceleration.The...  ...to and influence engineering discussions around...  ...'s growing suite of generative AI services and other cloud... 
    Amazon Web Service
    Cloud
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Sr Cloud Hardware Dev Engineer, AWS Generative AI & ML Servers. Be the first to apply!