Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Systems Development Engineer, AWS Generative AI & ML Servers

$148.7k - $201.2k

Amazon Locker

Do you want to build the backbone of Generative AI at AWS? Do you want to build the future of the cloud for AI training and inference, delivering continuous price performance improvements for multi-billion variable LLMs at cloud scale? Come join us. We are seeking a Systems Development Engineer to develop automation software, diagnostic tooling, and fleet health infrastructure for our accelerated (AI/ML) server platforms. You will work across multiple teams and organizations to build scalable, reliable systems that keep our fleet healthy — with a vision toward zero-touch operations where automation detects, diagnoses, and resolves issues without human intervention.What You Will DoYou will solve complex architectural problems that may not be well-defined in advance. You will own your team's systems, proactively identify deficiencies, and write scalable, robust code to solve issues before they impact customers. You will decompose large, difficult server testability, reliability, and diagnosis problems into straightforward tasks and components — delivering yourself and through others in parallel — using a combination of hardware, software, system design, processor architecture, diagnostics, and operations knowledge.Key job responsibilitiesFleet Health & Predictive Infrastructure1. Build and own the automation infrastructure responsible for the health of the accelerator (AI/ML) compute server fleet2. Design and implement predictive failure detection systems using telemetry, sensor data, error trending, and log correlation to identify hardware issues before they cause customer impact3. Drive toward zero-touch operations — building automation that detects, diagnoses, triages, and remediates hardware and software faults without human intervention4. Develop monitoring tools, dashboards, and alerting systems to provide real-time visibility into fleet health across lab and production environments5. Define and track fleet health metrics (failure rates, mean time to detect, mean time to repair, first-time fix rate, predictive accuracy)Debugging & Troubleshooting1. Debug and resolve complex system-level issues across compute, GPU, and networking in production environments2. Troubleshoot Linux boot and runtime failures across x86 and ARM architectures, including PCIe, power, NIC, NVMe, and GPU subsystems3. Perform root cause analysis on hardware failures — correlating across firmware, kernel, driver, and physical layer to isolate faults4. Build diagnostic tooling that automates root cause identification and reduces reliance on manual triage5. Improve manufacturing throughput and yield through test optimizationSystems Development & Automation1. Define and develop software, automation, and enabling tools for server hardware programs; track and report progress2. Design and build scalable system-level software with focus on durability, availability, security, and diagnostics3. Develop and maintain device drivers for Linux on ARM and x86 architectures4. Build automation solutions using modern programming languages (Python, Ruby, Java, C/C++, etc.)5. Work with OS internals and accelerator/GPU software stacks in Linux-based environments6. Build, manage, and deploy CI/CD pipelines for rapid deployment of code changes to org-owned and customer-owned systemsCross-Team Collaboration1. Work across internal HWEng teams to ensure new server hardware addresses data path and control path functionality needed by dependent service teams2. Work closely with internal customers to identify early any potential problems onboarding new accelerated compute servers into their ecosystem3. Engage with ODMs and design partners on testability, diagnostic, and automation requirements during hardware design and development (NPI)4. Contribute to server design to improve robustness, testability, diagnosability, and reliability5. Partner with datacenter operations teams to close the loop between field failures and design improvementsA day in the lifeYou will collaborate with a variety of roles (SDEs, SDETs, Mechanical/Electrical/Hardware Engineers, TPMs, Managers, Principals) and organizations through server conception, test validation, qualification, launch, and operations — driving high quality and reliability into current and future designs for AWS accelerated server solutions. From orchestration tooling development to hardware integration to kernel driver debugging, you dive deep into problems across the breadth of AWS.About the teamThe Hardware Engineering AI/ML development team is a group of engineers and technical program managers directly responsible for launching and maintaining server hardware in the fleet — including AI/ML accelerator servers with GPUs. Located in Seattle, Cupertino, and Austin, we work with internal development teams, ODMs, and design partners to deliver servers deployed in datacenters worldwide.Basic qualifications- 2+ years of non-internship professional software development experience- 1+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience- Experience programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, RubyPreferred qualification - Familiarity with server hardware architecture, BMC/IPMI, firmware, PCIe topology, and hardware diagnostics- Experience working with ODMs or hardware design partners- Exposure to zero-touch or self-healing automation concepts for large-scale infrastructure- Experience working in large-scale datacenter or cloud environments- Experience with hardware bring-up, validation, or fleet-wide deployment- Familiarity with telemetry pipelines, anomaly detection, or operational metrics at scaleAmazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at .USA, CA, Cupertino - 148,700.00 - 201,200.00 USD annuallyUSA, TX, Austin - 129,200.00 - 174,800.00 USD annuallyUSA, WA, Seattle - 129,200.00 - 174,800.00 USD annually

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Systems Development Engineer, AWS Generative AI & ML Servers in Cupertino, CA vacancy
  • $148.7k - $201.2k

     ...massively scalable systems that are used...  ...impacts all AWS systems globally...  ...deliver healthy servers for their...  ...build the next generation of platform level...  ...Deeply technical engineers, who stay close...  ...rack that runs AI/ML workloads, every...  ...professional software development experience- 3+... 
    Amazon Web Service
    Contract work
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  •  ...We are seeking a motivated AI / Machine Learning Engineer with hands-on experience in Intelligent Systems and Generative AI to join our growing...  ...AI solutions and scalable ML models. Experience: 6...  ...recommendation systems Knowledge of AWS, Azure, or GCP... 
    Amazon Web Service
    Full time
    Internship
    Relocation

    Hudson Manpower

    San Jose, CA
    1 day ago
  • $169k - $338k

     ...OfficePosition Summary...As a Distinguished AI/ML Engineer within Walmart Global Tech's Site...  ..., you will lead the technical development of next-generation agentic AI systems and intelligent automation...  ...experience (Azure, GCP, AWS) with deep knowledge of cloud-native... 
    Amazon Web Service
    Full time
    Temporary work
    Part time

    Walmart

    Sunnyvale, CA
    1 day ago
  • $50k - $120k

     ...Altimate AI, founded in 202...  ...AI-powered data engineering revolution. You...  ...search of a Senior Generative AI Engineer who...  ...models and AI systems at scale. This...  ...years of hands-on ML/AI experience with...  ...API development expertise (FastAPI...  ...architectures (AWS, Kubernetes) for... 
    Amazon Web Service
    Full time
    Worldwide

    Pa Early Stage Partners

    Sunnyvale, CA
    1 day ago
  • $184k - $287.5k

     ...motivated software engineers to join us and build AI inference systems that serve large-...  ...tuned and compiler-generated) using techniques...  ...for the field of ML Systems; survey recent...  ...cloud platforms (AWS/GCP/Azure),...  ...advance AI research and development to create... 
    Amazon Web Service
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $157.3k - $212.8k

     ...the backbone of Generative AI cloud at AWS? Do you want to...  ...scalability in AI/ML and HPC...  ...’ll support the development and management of...  ...hardware, and network engineers, supply chain specialists...  ...of accelerated servers.You will work...  ...and facilitate system development... 
    Amazon Web Service
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  • $185k - $215k

     ...machine learning and software development space who are interested in...  ...you will operate as a Staff AI Systems Engineer, setting technical direction...  ...-on expertise across the ML lifecycle—including dataset...  ...orchestration tools (e.g., AWS, Azure, GCP, Airflow, Argo Workflows... 
    Amazon Web Service
    Odd job
    Work experience placement
    Internship
    Local area
    Worldwide

    Robert Bosch

    Sunnyvale, CA
    2 days ago
  • $122.6k - $185k

     ...build the backbone of Generative AI cloud at AWS? Do you want to...  ...scalability in AI/ML and HPC workloads....  ...The AWS Hardware Engineering team creates server designs for Amazon...  ...scale and curious how systems and software...  ...Engineering AI / ML development team is a group of... 
    Amazon Web Service
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $135.2k - $306.4k

     ...highly skilled Linux Systems Engineer with deep expertise across...  ...should be comfortable generating Linux images,...  ...experience with RPM package development and maintenance, including...  ...saving care. And with AI embedded across our...  ...Amazon Web Services (AWS)o Microsoft Azureo Google... 
    Amazon Web Service
    Temporary work
    Flexible hours

    Oracle Corporation

    Santa Clara, CA
    2 days ago
  •  ...teams. 3+ yrs distributed systems / ML infra. About Orbifold AI Orbifold AI is building...  ...infrastructure that the next generation of physical AI runs on ....  ...a Machine Learning Engineer to scale and optimize the...  ...and major cloud providers (AWS, GCP, Azure) Hardware accelerator... 
    Amazon Web Service

    Orbifold AI

    Palo Alto, CA
    1 day ago
  • $99k - $121k

     ...of scientists, engineers, and physicians...  ...power of next-generation sequencing (NGS...  ...Engineer - Software Development Engineer -...  ...Intelligence (AI) is an experienced...  ...platforms (AWS, Azure, or GCP)...  ...source control systems, Agile development...  ...environments, AI/ML frameworks,... 
    Amazon Web Service
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    Shift work

    Jobzhr

    Sunnyvale, CA
    2 days ago
  • $148.7k - $201.2k

    Amazon Web Services (AWS) Hardware Engineering team creates compute and storage server designs for Amazon’s...  ...lead the design and development of server products utilizing...  ...to create next-generation hardware. You will have...  ...big difficult server system testability,... 
    Amazon Web Service
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $162.7k - $220.2k

     ...position is part of the AWS Specialist and Partner...  ...strategy, recruiting, development, and growth of our key...  ...(PDS) to grow the AWS Generative Artificial Intelligence & Machine Learning (AI/ML) business through consulting...  ...with or evaluating AI systems experiencePreferred... 
    Amazon Web Service
    Local area
    Flexible hours

    AmazonWebServices

    Mountain View, CA
    2 days ago
  • $122.6k - $174.6k

     ...inventive research and development company that designs and engineer’s high-profile...  ..., software systems and integrations...  ...intelligence, pioneering AI-based design review and design generation capabilities...  ...to integrate AI/ML technologies...  ...frameworks (AWS Bedrock, TensorFlow... 
    Amazon Web Service
    Local area
    Flexible hours

    Amazon

    Sunnyvale, CA
    4 days ago
  • $148.7k - $201.2k

    Join the AWS EC2 Nitro team building the foundation of cloud computing at unprecedented...  ...is tasked with developing the next generation of EC2 Supercomputers, optimized for...  ....We are looking for an experienced System development engineer to drive development for new EC2... 
    Amazon Web Service
    Internship
    Local area
    Work from home
    Worldwide
    Flexible hours

    Amazon

    Santa Clara, CA
    1 day ago
  • $160k - $198k

     ...members.What You’ll DoAs a Senior AI Systems Engineer, you will architect, deploy,...  ...and optimize the AI development lifecycle.Compute & Inference...  ...with a dedicated focus on AI/ML systems, high-performance computing...  ...-scaler infrastructure (AWS) alongside specialized AI-... 
    Amazon Web Service
    Local area

    Archer Aviation

    San Jose, CA
    1 day ago
  • $190k - $237k

     ...members.What You’ll DoAs a Staff AI Systems Engineer, you will architect, deploy,...  ...and optimize the AI development lifecycle.Compute & Inference...  ...with a dedicated focus on AI/ML systems, high-performance computing...  ...-scaler infrastructure (AWS) alongside specialized AI-... 
    Amazon Web Service
    Local area

    Archer Aviation

    San Jose, CA
    4 days ago
  •  ...Machine Learning Engineer, you will be responsible...  ...learning models/systems and innovative web...  ...the power of Generative AI to our customers....  ...design, specification, development, testing, and...  ...building and evolving ML Training and...  ...AzureML, GCP Vertex, AWS Sagemaker or similar... 
    Amazon Web Service
    Local area

    Typeface

    Palo Alto, CA
    4 days ago
  •  ...Job Title: Senior AI/ML Engineer Work Location with ZIP: Sunnyvale,...  ...- Hands?on experience with Generative AI / LLMs (OpenAI, Azure OpenAI...  ...cloud environments (Azure / AWS / GCP) - Experience with...  ...environments - Familiarity with API development and microservices... 
    Amazon Web Service

    eTeam

    Sunnyvale, CA
    5 days ago
  •  ...Job Title: Senior AI Search & Agentic Systems Engineer Location: Onsite - Santa Clara...  ...Retrieval-Augmented Generation (RAG) pipelines to power...  ...with React.js for front-end development. Deep expertise in Elasticsearch...  ...(Azure preferred; AWS/GCP a plus). Strong... 
    Amazon Web Service

    ReqRoute,Inc

    Santa Clara, CA
    2 days ago
  • $131k - $175k

     ...prestigious awards, such as Best Engineering Team, Best Company for...  ...the world’s largest AI and cloud deployments....  ...and deployment of next-generation rack architectures...  ...rack layouts and cabling systems, including fiber/copper...  ...or large-scale AI/ML cluster deploymentsExperience... 
    Remote work
    Flexible hours

    Arista Networks

    Santa Clara, CA
    1 day ago
  • $100k - $200k

     ...experienced backend engineer to join our...  ...backbone for our generative AI-powered Android applications...  ...Core Development & Infrastructure...  ...Performance & Systems Optimize API performance...  ...with Android and ML teams on API contracts...  ...major cloud platform (AWS, Azure, or GCP) ~... 
    Amazon Web Service
    Full time

    OPPO US Research Center

    Palo Alto, CA
    3 days ago
  • $105k - $115k

     ...world-class end-to-end engineering solutions by...  ...will Do:Pilot next-generation technologies solving...  ...with by utilizing Gen AI or other machine...  ...platformsAI Engineer/ ML Engineer with...  ...some knowledge in aws or some cloud.Assess...  ...delivering NLP or LLM-based systems, with knowledge of... 
    Amazon Web Service
    Temporary work

    Quest Global Services

    Sunnyvale, CA
    2 days ago
  • $152k - $208.5k

     ...leader in materials engineering solutions used to...  ...our world - like AI and IoT. If you want...  ...to create next generation technology, join us...  ...expertise in intricate systems, deciphering code,...  ...AI assisted development workflows across Applied...  ...integrating AI/ML models into... 
    Full time

    Applied Materials

    Santa Clara, CA
    1 day ago
  •  ...Tech SolutionsJob Title: AI / ML engineerLocation:...  ...Senior Machine Learning Engineer / AI Solutions Architect...  ...scalable, production-ready AI systems.Key Responsibilities:...  ...(YOLOv8), and generative AI (Stable Diffusion, SAM...  ...with cloud platforms (AWS, Azure, GCP) for scalable... 
    Amazon Web Service
    Full time

    SRI Tech

    Sunnyvale, CA
    4 days ago
  • $200k - $322k

     ...learning ignited modern AI — the next era of...  ...today.Design-for-X Engineering at NVIDIA works on...  ...that requires ML & Gen AI expertise....  ...methodologies for our next generation products using Gen...  ...cloud platforms (AWS, Azure, GCP)...  ...globally distributed systems to meet our high standards... 
    Amazon Web Service
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $250k - $344.5k

     ...Security (NetSec) Engineering – Our team is at the...  ...customer-facing AI-enabled solutions...  ...dedicated to integrating Generative AI and automation...  ...you will lead the development of a diverse...  ...that leverage AI/ML to solve real-world...  ...native infrastructure (AWS/GCP, Kubernetes,... 
    Amazon Web Service

    Palo Alto Networks, Inc.

    Santa Clara, CA
    3 days ago
  •  ...skilled, and motivated engineers to join our founding team...  ...Develop frontend for AI model interaction Use...  ...solutions for portal development and leverage SaaS software...  ...with user management systems (Auth0, Stack Auth, Firebase...  ...platform experience (AWS, Azure, GCP) VLM and... 
    Amazon Web Service
    Full time

    Dexmate

    Santa Clara, CA
    1 day ago
  • $207k - $300k

     ...architecture, design, and development of high-performance server platforms,...  ...hardware, software, and system engineering teams to drive pre-...  ...develop the next-generation technologies that...  ...compute servers and ML headnodes. You will...  ...future solutions.The AI and Infrastructure... 
    Worldwide

    Google

    Sunnyvale, CA
    21 hours ago
  •  ...Full Stack AI Engineer/Developer - OnsiteNTT...  ...backend systems, real-time streaming...  ...powered application development....  ...workflowsCollaborate with AI/ML engineers to...  ...of MCP server designExperience...  ...structured JSON generation systemsHands-on...  ...APIsExposure to AWS, Azure, or GCP... 
    Amazon Web Service
    Work experience placement

    NTT DATA

    Santa Clara, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Systems Development Engineer, AWS Generative AI & ML Servers. Be the first to apply!