Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Sr Hardware Development Engineer, High Performance AI & ML Servers

$183k - $247.6k

Amazon

Description

Do you want to shape the future of AI? Join the team building the foundation of the world’s most advanced cloud for AI training and inference — where multi-billion-parameter models come to life at scale. Here, you’ll design, deliver, and operate next-generation infrastructure that powers breakthrough innovation in AI/ML and HPC workloads. If you’re passionate about pushing the limits of performance, efficiency, and scalability in the cloud, this is your opportunity to build the systems that define what’s next for AWS — and for the entire AI industry.

You’ll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts, operations managers, and other vital roles. You’ll collaborate with people across AWS to help us deliver the highest standards for safety and security while providing seemingly infinite capacity at the lowest possible cost for our customers. And you’ll experience an inclusive culture that welcomes bold ideas and empowers you to own them to completion.

Key job responsibilities

  • Lead technical solutions for complex high performance server and/or accelerator server and rack system architectural challenges

  • Own end-to-end system reliability, proactively identifying and resolving deficiencies before customer impact

  • Design and implement solutions to address system-level issues at large scale

  • Decompose complex server system problems (testability, reliability, diagnostics) into deliverable tasks and features

  • Apply expertise across hardware, software, system design, x86 architecture, processes, and operations

  • Collaborate with hardware, software, manufacturing, supply chain and product management teams

  • Develop and implement diagnostic tools and monitoring solutions for production systems

  • Debug complex system failures in time sensitive settings

A day in the life

Your day to day responsibilities will include interfacing with our internal and external customers to understand project requirements and facilitate system development on top of your server design. You will be responsible for solving operational challenges to our existing fleet with the goal of improving the current customer experience as well as developing improved systems for future designs. You will work directly with vendors and ODM/JDM design teams to develop and manufacture your product at scale.

About the team

The team is comprised of both Hardware Design Engineers, System Design Engineers, Software Development Engineers and Technical Program Managers, all with the common goal of delivering the best specialized server fleet possible to our customers.

Why AWS

Amazon Web Services (AWS) is the world’s most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating — that’s why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses.

Inclusive Team Culture

Here at AWS, it’s in our nature to learn and be curious. Our employee-led affinity groups foster a culture of inclusion that empower us to be proud of our differences. Ongoing events and learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences, inspire us to never stop embracing our uniqueness.

Work/Life Balance

We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there’s nothing we can’t achieve in the cloud.

Mentorship and Career Growth

We’re continuously raising our performance bar as we strive to become Earth’s Best Employer. That’s why you’ll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional.

Diverse Experiences

Amazon values diverse experiences. Even if you do not meet all of the preferred qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn’t followed a traditional path, or includes alternative experiences, don’t let it stop you from applying.

Basic Qualifications

  • Experience in developing functional specifications, design verification plans and functional test procedures

  • Experience in server technologies such as, thermal, mechanical, power, and signal integrity

  • Bachelor's degree or above in electrical engineering, computer engineering, or equivalent

  • 5+ years of Design/Innovation, research & development, manufacturing, process, industrial engineering, or related experience

  • 5+ years of process development experience

  • Experience in English-language communication skills, both written and verbal

  • In depth expertise in one or more server technologies such as Thermal / Mechanical design, high speed bus design and signal integrity, failure analysis, server components (e.g. CPU, GPU, SSDs, memory), BIOS, BMC, and networking

Preferred Qualifications

  • Master's degree or above in electrical engineering, computer engineering, or equivalent

  • Experience working with interdisciplinary teams to execute product design from concept to production

  • 10+ years of server, storage, networking, or large-scale distributed systems experience

  • Experience working with engineering and product teams to define a product and bring it to market

  • 5+ years of data center engineering or operations experience

  • Experience in Linux/RHEL, or experience with programming/scripting (Batch, VB, PowerShell, Java, C#, Chef, Perl, Ruby and/or PHP) and experience that includes strong analytical skills, attention to detail, and effective communication abilities

  • Experience with server validation and issue root causing

  • Experience with leading hardware and software development engineering teams

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at .

USA, CA, Cupertino - 183,000.00 - 247,600.00 USD annually

USA, TX, Austin - 159,200.00 - 215,300.00 USD annually

USA, WA, Seattle - 159,200.00 - 215,300.00 USD annually

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Sr Hardware Development Engineer, High Performance AI & ML Servers in Cupertino, CA vacancy
  • $148.7k - $201.2k

     ...backbone of Generative AI at AWS? Do you...  ...continuous price performance improvements for...  ...a Systems Development Engineer to develop automation...  ...accelerated (AI/ML) server platforms. You will...  ...a combination of hardware, software, system...  ...operations — driving high quality and... 
    Performance
    Internship
    Local area
    Worldwide
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $131k - $175k

     ...awards, such as Best Engineering Team, Best Company...  ...of quality and performance in everything we do...  ...teams, including hardware, software, thermal...  ...the world’s largest AI and cloud...  ...power, cooling, into high-density GPU environments...  ...or large-scale AI/ML cluster... 
    Senior
    Performance
    Remote work
    Flexible hours

    Arista Networks

    Santa Clara, CA
    5 days ago
  • $122.6k - $185k

     ...backbone of Generative AI cloud at AWS? Do...  ...price performance improvements in...  ...offerings that enable high performance and...  ...in AI/ML and HPC workloads...  ...like you. The AWS Hardware Engineering team creates server designs for Amazon...  ...Engineering AI / ML development team is a group... 
    Performance
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $183k - $247.6k

     ...you’ll support the development and management of...  ....We are seeking a Hardware Design Engineer with role in the definition...  ...next generation ML Chips, Cards and server integration. As a...  .... You’ll have high standards for...  ...improve your products performance, quality and cost.... 
    Senior
    Performance
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  • $183k - $247.6k

     ...basisAWS Compute & ML Services owns...  ...and all of the servers, storage, networking...  ...of software, hardware, and network engineers, supply chain specialists...  ...-leading in performance, frugality and...  ...sound, high-quality server systems...  ...and support the development of automated monitoring... 
    Senior
    Performance
    Local area
    Flexible hours

    AmazonWebServices

    Cupertino, CA
    3 days ago
  • $183k - $247.6k

     ...next-generation server components. You will...  ...architecture and development of components...  ...and programmable hardware. You will interact...  ...interdisciplinary team of engineers to design,...  ...areas. You have high standards for yourself...  ...ways to improve performance, quality, and... 
    Senior
    Performance
    Local area
    Flexible hours

    AmazonWebServices

    Cupertino, CA
    3 days ago
  • $157.3k - $212.8k

     ...of Generative AI cloud at AWS?...  ...price performance improvements...  ...that enable high performance and...  ...scalability in AI/ML and HPC workloads...  ...support the development and...  ...of software, hardware, and network engineers, supply chain...  ...accelerated servers.You will work... 
    Performance
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $210k - $290k

     ...interconnect solutions, including High-Speed Board-to-Board, High-...  ...to optimize both the performance and cost of a system from the...  ...optoelectronic transceiver design and development engineers own the design of our high-...  ...hyperscale data center and AI/ML interconnects, 100G/400G/800... 
    Performance
    Flexible hours

    Samtec

    Santa Clara, CA
    a month ago
  • $175k - $225k

    The AI infrastructure market...  ...looking for a Sr. Technical Product...  ....This is a high-impact, high-...  ...working with engineering to define the...  ...across hardware (NVIDIA Blackwell...  ..., high-performance storage protocols...  ...storage, or AI/ML...  ...joint solution development, and partner-... 
    Senior
    Performance
    Shift work

    DataDirect Networks

    Santa Clara, CA
    2 days ago
  • $150k - $225k

     ...We are building a highly-advanced low...  .... As a Senior AI / Embedded Engineer, you will be responsible...  ...-constrained hardware. This includes data...  ...ingestion, model development, optimization,...  ...power, real-time ML systems that operate...  ...compare model performance • TinyML and Embedded... 
    Senior
    Performance
    Full time
    Work at office
    Immediate start
    Visa sponsorship
    Night shift

    E-Space

    Saratoga, CA
    22 hours ago
  • $144k - $208k

     ...system-level test development, managing...  ...Partner with design engineering teams to drive...  ...updates and performance dashboards.Author...  ...Author and maintain high-quality...  ...generation of hardware experiences, delivering...  ...Introduction (ML-NPI) team is...  ...behind Google’s AI infrastructure,... 
    Senior
    Performance
    Contract work
    Worldwide

    Google

    Sunnyvale, CA
    3 days ago
  •  ...experiences—from AI and data centers,...  ...solutions that combine hardware and software into...  ...credible with Engineering and TME; effective...  ...executable plansA high ownership mindset...  ...judgment in balancing performance, supportability,...  ...infrastructureAI/ML platformsHPC systemsEnterprise... 
    Senior
    Performance
    Work at office

    AMD

    Santa Clara, CA
    22 hours ago
  •  ...computing experiences—from AI and data centers, to...  ...are hiring AI / ML Platform Engineers to build the platform...  ...Applied AI Engineers, and hardware domain experts to...  ..., observability, performance, and developer experience...  ...GPU clusters or high-performance ML infrastructure... 
    Senior
    Performance

    AMD

    Santa Clara, CA
    4 days ago
  • $183k - $247.6k

     ...Services’ Network Products Development team is looking for experienced hardware engineering professionals to help...  ...experience in high speed digital design as...  ...centers and all of the servers, storage, networking,...  ...continuously raising our performance bar as we strive to become... 
    Senior
    Performance
    Local area
    Flexible hours

    AmazonWebServices

    Cupertino, CA
    3 days ago
  • $183k - $247.6k

     ...centers and all of the servers, storage, networking...  ...team of software, hardware, and network engineers, supply chain...  ...set the standards on performance, quality, cost, and...  ...lead the design and development of server products utilizing...  ...areas. You’ll have high standards for... 
    Senior
    Performance
    Local area
    Flexible hours

    AmazonWebServices

    Cupertino, CA
    1 day ago
  • $193.3k - $261.5k

     ..., the software development kit used to accelerate...  ...includes an ML compiler,...  ...their training performance. Working across...  ...JAX down to the hardware and software boundary, our engineers build the infrastructure...  ..., and tune high-performance...  ...the direction of AI acceleration... 
    Senior
    Performance
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $143.3k

     ...Networking and Outpost hardware team. We are a start-...  ...in building high performant hardware used to run...  ...drives changes back into development and builds mechanisms...  ...computing.We aim to hire engineers who will thrive in a...  ...development on top of your server design. You will be... 
    Senior
    Performance
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  • Cerebras Systems Inc. in Sunnyvale, CA is seeking a Hardware Analytics Engineer to design and operate scalable data pipelines for telemetry, reliability analytics, and performance optimization across AI server platforms. You will develop frameworks using Python, SQL, Tableau... 
    Performance
    Remote job

    Cerebras

    Sunnyvale, CA
    5 days ago
  •  ...millions of patients worldwide.We’re a team of engineers, clinicians, and innovators united by one...  .... Every day, our work helps care teams perform with greater precision and patients...  ...generation robotic platforms. As a Senior AI/ML Research Engineer, you will develop and fine... 
    Senior
    Performance
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    Sunnyvale, CA
    1 day ago
  • $140k - $215k

     ...world’s most advanced AI-native platform....  ...work closely with engineering teams to expand...  ...ensures reliability and performance as we deploy AI...  ...combined software development and test...  ...Experience testing AI/ML systems, LLM applications...  ...products handling high throughput,... 
    Senior
    Performance
    Full time
    Contract work
    Work experience placement
    Work at office
    Local area

    CrowdStrike

    Sunnyvale, CA
    1 day ago
  • $198k - $326k

     ...it will be performed both from home...  ...LinkedIn's AI model...  ...training, feature engineering and serving...  ..., and hardware to harness...  ...implement high performance...  ...guide the development of containerized...  ...of new ML models per...  ...scale.As a Sr. Staff Software...  ...and client-server architecturesExperience... 
    Senior
    Performance
    For contractors
    Work at office
    Flexible hours

    Linkedin

    Sunnyvale, CA
    2 days ago
  • $129.3k - $193.9k

     ...Inc. Job Area: Engineering Group, Engineering Group...  ...skilled and motivated AI Model Training Engineer...  ...scalable models that meet performance, efficiency, and...  ...engineering teams to ensure high-quality, well-labeled,...  ...Python and familiar with ML libraries such as... 
    Senior
    Performance
    Work experience placement
    Work from home

    Qualcomm

    Santa Clara, CA
    22 hours ago
  • $111.61k - $131.3k

     ...analysis, design, testing, development and maintenance of...  ..., and maintain core AI-SDLC platform capabilities...  ...adoption.Partner with engineering and delivery teams to...  ...reliability and performance.Build and maintain secure...  ...closed earlier due to high volume of applicants.SummaryLocation... 
    Senior
    Performance
    Full time
    Work experience placement
    Local area
    3 days per week

    US Bank

    Cupertino, CA
    1 day ago
  •  ...world’s largest AI chip, 56 times...  ...workloads with ultra high-speed inference...  ...' Wafer-Scale Engine (WSE). We build...  ..., high-performance kernel enablement...  ...distributed systems, hardware architecture,...  ...for ML model compilation...  ...optimization as well as development of high-... 
    Senior
    Performance

    Cerebras

    Sunnyvale, CA
    4 days ago
  • $210k - $220k

     ...comprehensive Data Transformation/AI/ML and automation Vision/Strategy aligned...  .... Lead and manage a small but high-performing team of data engineers, machine learning engineers, and automation specialists. Oversee the development, testing, and deployment of Data Transformation... 
    Senior
    Performance
    Part time
    Work experience placement
    Remote work
    Flexible hours

    Danaher Corporation

    Sunnyvale, CA
    3 days ago
  • $100k

     ...industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use...  ...technologists have developed a high performance RISC-V CPU...  ...of Tenstorrent hardware using an open-source, performant...  ...AreA passionate software engineer eager to work on compiler... 
    Senior
    Performance
    Permanent employment

    Tenstorrent

    Santa Clara, CA
    3 days ago
  •  ...experiences—from AI and data centers,...  ...solutions that combine hardware and software into...  ...credible with Engineering and TME; effective...  ...executable plans A high ownership mindset...  ...judgment in balancing performance, supportability,...  ...AI/ML platforms HPC systems... 
    Senior
    Performance
    Work at office

    Advanced Micro Devices

    Santa Clara, CA
    3 days ago
  • $100k

     ...industry on cutting-edge AI technology, revolutionizing performance expectations, ease...  ...have developed a high performance RISC-V...  ...skilled Software Engineer with expertise in compilers...  ...work closely with hardware engineers, software...  ..., and software development in C/C++. Python... 
    Senior
    Performance
    Permanent employment
    Full time

    Tenstorrent

    Santa Clara, CA
    22 hours ago
  • $185.39k - $277.7k

     ...Across enterprise, cloud and AI, and carrier architectures, our...  ...ImpactWe are seeking a Director of Hardware Engineering to lead the design, analysis, and validation of high-speed board-level hardware for...  ..., and delivering high-performance hardware that meets timing, reliability... 
    Performance
    Permanent employment
    Full time
    Internship
    Work from home

    Marvell

    Santa Clara, CA
    3 days ago
  • $241k - $344.5k

     ...the unique demands of AI. You will also lead programs...  ...), balancing high-level business acumen with...  ...You will partner with engineering to define what to build...  ...familiarity with the AI/ML lifecycle and the governance...  ...RAG pipelines, monitor performance with evals and have a... 
    Senior
    Performance
    Full time
    Work at office
    Local area

    Palo Alto Networks

    Santa Clara, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Sr Hardware Development Engineer, High Performance AI & ML Servers. Be the first to apply!