Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Systems Development Engineer, AWS Generative AI & ML Servers

$148.7k - $201.2k

Amazon Locker

Do you want to build the backbone of Generative AI at AWS? Do you want to build the future of the cloud for AI training and inference, delivering continuous price performance improvements for multi-billion variable LLMs at cloud scale? Come join us. We are seeking a Systems Development Engineer to develop automation software, diagnostic tooling, and fleet health infrastructure for our accelerated (AI/ML) server platforms. You will work across multiple teams and organizations to build scalable, reliable systems that keep our fleet healthy — with a vision toward zero-touch operations where automation detects, diagnoses, and resolves issues without human intervention.What You Will DoYou will solve complex architectural problems that may not be well-defined in advance. You will own your team's systems, proactively identify deficiencies, and write scalable, robust code to solve issues before they impact customers. You will decompose large, difficult server testability, reliability, and diagnosis problems into straightforward tasks and components — delivering yourself and through others in parallel — using a combination of hardware, software, system design, processor architecture, diagnostics, and operations knowledge.Key job responsibilitiesFleet Health & Predictive Infrastructure1. Build and own the automation infrastructure responsible for the health of the accelerator (AI/ML) compute server fleet2. Design and implement predictive failure detection systems using telemetry, sensor data, error trending, and log correlation to identify hardware issues before they cause customer impact3. Drive toward zero-touch operations — building automation that detects, diagnoses, triages, and remediates hardware and software faults without human intervention4. Develop monitoring tools, dashboards, and alerting systems to provide real-time visibility into fleet health across lab and production environments5. Define and track fleet health metrics (failure rates, mean time to detect, mean time to repair, first-time fix rate, predictive accuracy)Debugging & Troubleshooting1. Debug and resolve complex system-level issues across compute, GPU, and networking in production environments2. Troubleshoot Linux boot and runtime failures across x86 and ARM architectures, including PCIe, power, NIC, NVMe, and GPU subsystems3. Perform root cause analysis on hardware failures — correlating across firmware, kernel, driver, and physical layer to isolate faults4. Build diagnostic tooling that automates root cause identification and reduces reliance on manual triage5. Improve manufacturing throughput and yield through test optimizationSystems Development & Automation1. Define and develop software, automation, and enabling tools for server hardware programs; track and report progress2. Design and build scalable system-level software with focus on durability, availability, security, and diagnostics3. Develop and maintain device drivers for Linux on ARM and x86 architectures4. Build automation solutions using modern programming languages (Python, Ruby, Java, C/C++, etc.)5. Work with OS internals and accelerator/GPU software stacks in Linux-based environments6. Build, manage, and deploy CI/CD pipelines for rapid deployment of code changes to org-owned and customer-owned systemsCross-Team Collaboration1. Work across internal HWEng teams to ensure new server hardware addresses data path and control path functionality needed by dependent service teams2. Work closely with internal customers to identify early any potential problems onboarding new accelerated compute servers into their ecosystem3. Engage with ODMs and design partners on testability, diagnostic, and automation requirements during hardware design and development (NPI)4. Contribute to server design to improve robustness, testability, diagnosability, and reliability5. Partner with datacenter operations teams to close the loop between field failures and design improvementsA day in the lifeYou will collaborate with a variety of roles (SDEs, SDETs, Mechanical/Electrical/Hardware Engineers, TPMs, Managers, Principals) and organizations through server conception, test validation, qualification, launch, and operations — driving high quality and reliability into current and future designs for AWS accelerated server solutions. From orchestration tooling development to hardware integration to kernel driver debugging, you dive deep into problems across the breadth of AWS.About the teamThe Hardware Engineering AI/ML development team is a group of engineers and technical program managers directly responsible for launching and maintaining server hardware in the fleet — including AI/ML accelerator servers with GPUs. Located in Seattle, Cupertino, and Austin, we work with internal development teams, ODMs, and design partners to deliver servers deployed in datacenters worldwide.Basic qualifications- 2+ years of non-internship professional software development experience- 1+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience- Experience programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, RubyPreferred qualification - Familiarity with server hardware architecture, BMC/IPMI, firmware, PCIe topology, and hardware diagnostics- Experience working with ODMs or hardware design partners- Exposure to zero-touch or self-healing automation concepts for large-scale infrastructure- Experience working in large-scale datacenter or cloud environments- Experience with hardware bring-up, validation, or fleet-wide deployment- Familiarity with telemetry pipelines, anomaly detection, or operational metrics at scaleAmazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at .USA, CA, Cupertino - 148,700.00 - 201,200.00 USD annuallyUSA, TX, Austin - 129,200.00 - 174,800.00 USD annuallyUSA, WA, Seattle - 129,200.00 - 174,800.00 USD annually

Vacancy posted 14 hours ago
Similar jobs that could be interesting for youBased on the Systems Development Engineer, AWS Generative AI & ML Servers in Cupertino, CA vacancy
  • $183k - $247.6k

     ...the future of AI? Join the...  ...operate next-generation infrastructure...  ...innovation in AI/ML and HPC...  ...to build the systems that define what...  ...what’s next for AWS — and for the...  ...and network engineers, supply chain...  ...performance server and/or accelerator...  ...system development on top of your... 
    Amazon Web Service
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $183k - $247.6k

    AWS operates the world's largest fleet of GPU-accelerated servers powering AI/ML training and inference at cloud scale...  ...a Cloud Hardware Development Engineer to define server architectures...  ...to identify systemic issues and drive...  ...for the next-generation platform.About the... 
    Amazon Web Service
    Local area
    Worldwide
    Flexible hours
    Day shift

    Amazon

    Cupertino, CA
    2 days ago
  • $173.9k - $235.2k

     ...massively scalable systems that are used...  ...impacts all AWS systems globally...  ...deliver healthy servers for their...  ...build the next generation of platform level...  ...Deeply technical engineers, who stay close...  ...rack that runs AI/ML workloads, every...  ...professional software development experience- 3+... 
    Amazon Web Service
    Contract work
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    4 days ago
  • $148.7k - $201.2k

     ...Intelligence compute capacity available to Generative AI customers? Do you want to solve...  ...and software - at cloud scale?AWS Hardware Engineering is looking for a Systems Development Engineer to own the health and development of server platforms at worldwide fleet scale.... 
    Amazon Web Service
    Internship
    Local area
    Worldwide
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $169k - $338k

     ...OfficePosition Summary...As a Distinguished AI/ML Engineer within Walmart Global Tech's Site...  ..., you will lead the technical development of next-generation agentic AI systems and intelligent automation...  ...experience (Azure, GCP, AWS) with deep knowledge of cloud-native... 
    Amazon Web Service
    Full time
    Temporary work
    Part time

    Walmart

    Sunnyvale, CA
    2 days ago
  • $160k - $198k

     ...across air taxis, UAS, AI and powertrain development - building real-world systems that combine software intelligence...  ...'re seeking exceptional engineers, operators and builders...  ...a dedicated focus on AI/ML systems, high-...  ...-scaler infrastructure (AWS) alongside specialized AI... 
    Amazon Web Service
    Local area
    Visa sponsorship
    Night shift

    Archer

    San Jose, CA
    1 day ago
  • $177.54k - $295.9k

     ...The Solution Engineer is a core contributor...  ...knowledge of test systems and networking protocols...  ...demands of AI/ML data center networking...  ..., Kubernetes, AWS, GCP, and Azure, is...  ...innovation in solution development and deployment....  ...experience with traffic generation / network test... 
    Amazon Web Service
    Work experience placement
    Free visa
    Shift work

    Keysight Technologies

    Santa Clara, CA
    2 days ago
  •  ...Job Title: Senior AI/ML Engineer Work Location with ZIP: Sunnyvale,...  ...- Hands?on experience with Generative AI / LLMs (OpenAI, Azure OpenAI...  ...cloud environments (Azure / AWS / GCP) - Experience with...  ...environments - Familiarity with API development and microservices... 
    Amazon Web Service

    eTeam

    Sunnyvale, CA
    1 day ago
  • $50k - $120k

     ...Altimate AI, founded in 202...  ...AI-powered data engineering revolution. You...  ...search of a Senior Generative AI Engineer who...  ...models and AI systems at scale. This...  ...years of hands-on ML/AI experience with...  ...API development expertise (FastAPI...  ...architectures (AWS, Kubernetes) for... 
    Amazon Web Service
    Full time
    Worldwide

    Pa Early Stage Partners

    Sunnyvale, CA
    16 hours ago
  • $184k - $287.5k

     ...motivated software engineers to join us and build AI inference systems that serve large-...  ...tuned and compiler-generated) using techniques...  ...for the field of ML Systems; survey recent...  ...cloud platforms (AWS/GCP/Azure),...  ...advance AI research and development to create... 
    Amazon Web Service
    Full time

    Nvidia

    Santa Clara, CA
    16 hours ago
  •  ...Job Title: AI/ML Engineer Location: Sunnyvale, CA, USA Duration...  ...Apple, PayPal, Netflix, Meta, AWS, Product Companies, Startups...  ...You will Do: Pilot next-generation technologies solving problems...  ...delivering NLP or LLM-based systems, with knowledge of multi-modal... 
    Amazon Web Service

    RIT Solutions, Inc.

    Sunnyvale, CA
    16 hours ago
  • $139.23k - $163.8k

     ...DescriptionLead Software Architect Engineer (Generative AI Platforms) is...  ..., and distributed system architectures.Implement...  ...across Azure and AWS, ensuring high availability...  ...proficiency in Python development, API design, and...  ...deploying and managing AI/ML workloads in Azure and... 
    Amazon Web Service
    Full time
    Work experience placement
    Local area
    3 days per week

    US Bank

    Cupertino, CA
    1 day ago
  • $100k - $200k

     ...experienced backend engineer to join our...  ...backbone for our generative AI-powered Android applications...  ...Core Development & Infrastructure...  ...Performance & Systems Optimize API performance...  ...with Android and ML teams on API contracts...  ...major cloud platform (AWS, Azure, or GCP) ~... 
    Amazon Web Service
    Full time

    OPPO US Research Center

    Palo Alto, CA
    29 days ago
  • $122.6k - $185k

     ...build the backbone of Generative AI cloud at AWS? Do you want to...  ...scalability in AI/ML and HPC workloads....  ...The AWS Hardware Engineering team creates server designs for Amazon...  ...scale and curious how systems and software...  ...Engineering AI / ML development team is a group of... 
    Amazon Web Service
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  •  ...with 5–7 years of experience in AI/ML, Generative AI, and Agentic AI systems to join our Advanced Supply...  ...frameworks , and cloud-based ML engineering (AWS) . You will design scalable, production...  .../ML, Generative AI & Agentic AI Development  Design, build, and deploy AI/... 
    Amazon Web Service
    Full time
    Temporary work
    Remote work
    Flexible hours
    Shift work

    Sandisk

    Milpitas, CA
    1 day ago
  •  ...AI Engineer Client: UST/Applied Materials Location...  ...chunking, embedding generation, and retrieval systems About the Role:...  ...our team and lead the development of intelligent AI...  ...computer science, AI/ML, or related field....  ...with cloud platforms (AWS, Azure, GCP) and containerization... 
    Amazon Web Service
    Local area

    Kasmo Global

    Santa Clara, CA
    3 days ago
  • $168k - $270.25k

    Site Reliability Engineering (SRE) at NVIDIA is...  ...scale production systems with high efficiency...  ...Much of our software development focuses on...  ...power NVIDIA’s next-generation AI-driven enterprise...  ...Security, and AI/ML teams to implement...  ...infrastructure-as-code such as AWS CDK, AWS... 
    Amazon Web Service
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...Senior Data Scientist – ML, GenAI & Agentic AI Location: Santa...  ...closely with product and engineering teams, and can...  ...productionize LLM and Generative AI solutions. Develop...  ...graphs. Work with AWS Bedrock or comparable...  ...improvement of AI/ML systems. Mentor junior data... 
    Amazon Web Service
    Full time

    Umanist Staffing LLC

    Santa Clara, CA
    6 days ago
  •  ...Overview We are seeking expertise in Generative AI and Python to work directly...  ...a trusted advisor, hands-on engineer, and delivery lead — driving...  ...experience in software engineering, ML engineering, or technical...  ...with cloud platforms (AWS, Azure, GCP) and MLOps tools is... 
    Amazon Web Service
    Full time

    Infinite Computer Solutions

    Sunnyvale, CA
    5 days ago
  • $180k - $210k

     ...available to any downstream query engine and use case (from traditional analytics to real-time AI / ML). We are a team of self-...  ...have created large-scale data systems and globally distributed platforms...  ...including Uber, Snowflake, AWS, Linkedin, Confluent and many more... 
    Amazon Web Service
    Work at office
    Remote work

    Onehouse

    Sunnyvale, CA
    21 days ago
  •  ...Responsibilities AI Integration:...  ...and implementing generative AI features...  ...Llama. Frontend Development: Building responsive...  ...models. Prompt Engineering & RAG:...  ...Generation (RAG) systems and crafting effective...  ...deploying to AWS, Azure, or GCP....  ...1 year focused on AI/ML. VBeyond
    Amazon Web Service

    VBeyond

    Sunnyvale, CA
    3 days ago
  • $150.4k - $277.6k

     ...Machine Learning and AI Imagine building...  ...at Apple — systems that power how products are made, how engineers think, and ultimately...  ...shipping Apple's next-generation sensing technologies...  ...cloud AI services (AWS Bedrock, Azure OpenAI...  ...projects in AI/ML infrastructure Comfort... 
    Amazon Web Service
    Relocation

    Apple

    Cupertino, CA
    1 day ago
  •  ...and influential engineering leader who has...  ...ability to scale AI from early...  ...machine learning and Generative AI, including...  ...deploying AI/ML applications at...  ...experience with AI systems, including data...  ...scale (e.g., AWS, Azure, GCP)....  ...Enabling the development of high-performance... 
    Amazon Web Service

    Synopsys Inc

    Sunnyvale, CA
    a month ago
  • $200k - $210k

    Get AI-powered advice on this job and more...  ...roadmap for AI/Generative AI platforms. Prioritize...  ...to align AI/ML initiatives with company...  ...the design and development of scalable,...  ...infrastructure. Work with engineering teams to build LLM...  ...AI platforms (AWS SageMaker, GCP... 
    Amazon Web Service
    Full time
    Local area
    Flexible hours

    HCLTech

    Santa Clara, CA
    5 days ago
  •  ...Generative AI Engineer The project is focused on building production-grade GenAI solutions with emphasis...  ...leveraging LLMs Prompt engineering (system/tool prompts, function calling,...  ...Cloud LLM providers (Azure OpenAI, AWS Bedrock, Vertex AI) Workflow orchestration... 
    Amazon Web Service

    InterSources

    Mountain View, CA
    4 days ago
  •  ...Job Title: AI Engineer Location: Santa Clara CA...  ...develop, and implement AI/ML models and algorithms...  ...cloud platforms (e.g., AWS, Azure, GCP). Collaborate...  ...models into existing systems. Optimize AI...  ...Participate in the full software development lifecycle, including... 
    Amazon Web Service
    Local area

    United IT Solutions

    Santa Clara, CA
    3 days ago
  • $248k - $396.75k

     ...Site Reliability Engineering (SRE) at NVIDIA...  ...production systems with exceptional...  ...direction of NVIDIA’s AI Platform...  ...NVIDIA’s next-generation AI-driven products...  ...the design and development of AI agents,...  ...Networking, and AI/ML organizations...  ...such as AWS, Azure, or GCP.... 
    Amazon Web Service
    Full time

    NVIDIA

    Santa Clara, CA
    3 days ago
  • $182k - $242k

     ...Senior Software and AI Engineer Livingston, NJ...  ...AI Engineer, IT Systems, you'll design and...  ...Codex) to accelerate development, testing, and...  ...familiarity with Generative AI frameworks ( LangChain...  ...-grade AI/ML/LLM-based applications...  ...cloud environments (AWS, Azure, or GCP),... 
    Amazon Web Service
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    1 day ago
  • $215k - $250k

     ...Data Infrastructure Engineer Onehouse is a...  ...analytics to real-time AI / ML). We are a team...  ...large-scale data systems and globally...  ...Uber, Snowflake, AWS, Linkedin, Confluent...  ...productionize the next generation of our data tech...  ...across feature development and tech debt with... 
    Amazon Web Service
    Odd job
    Work at office
    Local area
    Remote work
    Relocation
    Relocation package

    OneHouse LLC

    Sunnyvale, CA
    4 days ago
  • $173.9k - $235.2k

    AWS Hardware Engineering Services owns the design, planning, delivery...  ...AWS global compute systems. In other words, we’...  ...and all of the servers, storage, networking...  ...lead the design and development of server products utilizing...  ...to create next-generation hardware. You will... 
    Amazon Web Service
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    7 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Systems Development Engineer, AWS Generative AI & ML Servers. Be the first to apply!