Systems Development Engineer, AWS Generative AI & ML Servers
$148.7k - $201.2kAmazon Locker
Do you want to build the backbone of Generative AI at AWS? Do you want to build the future of the cloud for AI training and inference, delivering continuous price performance improvements for multi-billion variable LLMs at cloud scale? Come join us. We are seeking a Systems Development Engineer to develop automation software, diagnostic tooling, and fleet health infrastructure for our accelerated (AI/ML) server platforms. You will work across multiple teams and organizations to build scalable, reliable systems that keep our fleet healthy — with a vision toward zero-touch operations where automation detects, diagnoses, and resolves issues without human intervention.What You Will DoYou will solve complex architectural problems that may not be well-defined in advance. You will own your team's systems, proactively identify deficiencies, and write scalable, robust code to solve issues before they impact customers. You will decompose large, difficult server testability, reliability, and diagnosis problems into straightforward tasks and components — delivering yourself and through others in parallel — using a combination of hardware, software, system design, processor architecture, diagnostics, and operations knowledge.Key job responsibilitiesFleet Health & Predictive Infrastructure1. Build and own the automation infrastructure responsible for the health of the accelerator (AI/ML) compute server fleet2. Design and implement predictive failure detection systems using telemetry, sensor data, error trending, and log correlation to identify hardware issues before they cause customer impact3. Drive toward zero-touch operations — building automation that detects, diagnoses, triages, and remediates hardware and software faults without human intervention4. Develop monitoring tools, dashboards, and alerting systems to provide real-time visibility into fleet health across lab and production environments5. Define and track fleet health metrics (failure rates, mean time to detect, mean time to repair, first-time fix rate, predictive accuracy)Debugging & Troubleshooting1. Debug and resolve complex system-level issues across compute, GPU, and networking in production environments2. Troubleshoot Linux boot and runtime failures across x86 and ARM architectures, including PCIe, power, NIC, NVMe, and GPU subsystems3. Perform root cause analysis on hardware failures — correlating across firmware, kernel, driver, and physical layer to isolate faults4. Build diagnostic tooling that automates root cause identification and reduces reliance on manual triage5. Improve manufacturing throughput and yield through test optimizationSystems Development & Automation1. Define and develop software, automation, and enabling tools for server hardware programs; track and report progress2. Design and build scalable system-level software with focus on durability, availability, security, and diagnostics3. Develop and maintain device drivers for Linux on ARM and x86 architectures4. Build automation solutions using modern programming languages (Python, Ruby, Java, C/C++, etc.)5. Work with OS internals and accelerator/GPU software stacks in Linux-based environments6. Build, manage, and deploy CI/CD pipelines for rapid deployment of code changes to org-owned and customer-owned systemsCross-Team Collaboration1. Work across internal HWEng teams to ensure new server hardware addresses data path and control path functionality needed by dependent service teams2. Work closely with internal customers to identify early any potential problems onboarding new accelerated compute servers into their ecosystem3. Engage with ODMs and design partners on testability, diagnostic, and automation requirements during hardware design and development (NPI)4. Contribute to server design to improve robustness, testability, diagnosability, and reliability5. Partner with datacenter operations teams to close the loop between field failures and design improvementsA day in the lifeYou will collaborate with a variety of roles (SDEs, SDETs, Mechanical/Electrical/Hardware Engineers, TPMs, Managers, Principals) and organizations through server conception, test validation, qualification, launch, and operations — driving high quality and reliability into current and future designs for AWS accelerated server solutions. From orchestration tooling development to hardware integration to kernel driver debugging, you dive deep into problems across the breadth of AWS.About the teamThe Hardware Engineering AI/ML development team is a group of engineers and technical program managers directly responsible for launching and maintaining server hardware in the fleet — including AI/ML accelerator servers with GPUs. Located in Seattle, Cupertino, and Austin, we work with internal development teams, ODMs, and design partners to deliver servers deployed in datacenters worldwide.Basic qualifications- 2+ years of non-internship professional software development experience- 1+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience- Experience programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, RubyPreferred qualification - Familiarity with server hardware architecture, BMC/IPMI, firmware, PCIe topology, and hardware diagnostics- Experience working with ODMs or hardware design partners- Exposure to zero-touch or self-healing automation concepts for large-scale infrastructure- Experience working in large-scale datacenter or cloud environments- Experience with hardware bring-up, validation, or fleet-wide deployment- Familiarity with telemetry pipelines, anomaly detection, or operational metrics at scaleAmazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at .USA, CA, Cupertino - 148,700.00 - 201,200.00 USD annuallyUSA, TX, Austin - 129,200.00 - 174,800.00 USD annuallyUSA, WA, Seattle - 129,200.00 - 174,800.00 USD annually
$183k - $247.6k
...the future of AI? Join the... ...operate next-generation infrastructure... ...innovation in AI/ML and HPC... ...to build the systems that define what... ...what’s next for AWS — and for the... ...and network engineers, supply chain... ...performance server and/or accelerator... ...system development on top of your...Amazon Web ServiceLocal areaFlexible hours$183k - $247.6k
AWS operates the world's largest fleet of GPU-accelerated servers powering AI/ML training and inference at cloud scale... ...a Cloud Hardware Development Engineer to define server architectures... ...to identify systemic issues and drive... ...for the next-generation platform.About the...Amazon Web ServiceLocal areaWorldwideFlexible hoursDay shift$173.9k - $235.2k
...massively scalable systems that are used... ...impacts all AWS systems globally... ...deliver healthy servers for their... ...build the next generation of platform level... ...Deeply technical engineers, who stay close... ...rack that runs AI/ML workloads, every... ...professional software development experience- 3+...Amazon Web ServiceContract workInternshipLocal areaFlexible hours$148.7k - $201.2k
...Intelligence compute capacity available to Generative AI customers? Do you want to solve... ...and software - at cloud scale?AWS Hardware Engineering is looking for a Systems Development Engineer to own the health and development of server platforms at worldwide fleet scale....Amazon Web ServiceInternshipLocal areaWorldwideFlexible hours$169k - $338k
...OfficePosition Summary...As a Distinguished AI/ML Engineer within Walmart Global Tech's Site... ..., you will lead the technical development of next-generation agentic AI systems and intelligent automation... ...experience (Azure, GCP, AWS) with deep knowledge of cloud-native...Amazon Web ServiceFull timeTemporary workPart time$160k - $198k
...across air taxis, UAS, AI and powertrain development - building real-world systems that combine software intelligence... ...'re seeking exceptional engineers, operators and builders... ...a dedicated focus on AI/ML systems, high-... ...-scaler infrastructure (AWS) alongside specialized AI...Amazon Web ServiceLocal areaVisa sponsorshipNight shift$177.54k - $295.9k
...The Solution Engineer is a core contributor... ...knowledge of test systems and networking protocols... ...demands of AI/ML data center networking... ..., Kubernetes, AWS, GCP, and Azure, is... ...innovation in solution development and deployment.... ...experience with traffic generation / network test...Amazon Web ServiceWork experience placementFree visaShift work- ...Job Title: Senior AI/ML Engineer Work Location with ZIP: Sunnyvale,... ...- Hands?on experience with Generative AI / LLMs (OpenAI, Azure OpenAI... ...cloud environments (Azure / AWS / GCP) - Experience with... ...environments - Familiarity with API development and microservices...Amazon Web Service
$50k - $120k
...Altimate AI, founded in 202... ...AI-powered data engineering revolution. You... ...search of a Senior Generative AI Engineer who... ...models and AI systems at scale. This... ...years of hands-on ML/AI experience with... ...API development expertise (FastAPI... ...architectures (AWS, Kubernetes) for...Amazon Web ServiceFull timeWorldwide$184k - $287.5k
...motivated software engineers to join us and build AI inference systems that serve large-... ...tuned and compiler-generated) using techniques... ...for the field of ML Systems; survey recent... ...cloud platforms (AWS/GCP/Azure),... ...advance AI research and development to create...Amazon Web ServiceFull time- ...Job Title: AI/ML Engineer Location: Sunnyvale, CA, USA Duration... ...Apple, PayPal, Netflix, Meta, AWS, Product Companies, Startups... ...You will Do: Pilot next-generation technologies solving problems... ...delivering NLP or LLM-based systems, with knowledge of multi-modal...Amazon Web Service
$139.23k - $163.8k
...DescriptionLead Software Architect Engineer (Generative AI Platforms) is... ..., and distributed system architectures.Implement... ...across Azure and AWS, ensuring high availability... ...proficiency in Python development, API design, and... ...deploying and managing AI/ML workloads in Azure and...Amazon Web ServiceFull timeWork experience placementLocal area3 days per week$100k - $200k
...experienced backend engineer to join our... ...backbone for our generative AI-powered Android applications... ...Core Development & Infrastructure... ...Performance & Systems Optimize API performance... ...with Android and ML teams on API contracts... ...major cloud platform (AWS, Azure, or GCP) ~...Amazon Web ServiceFull time$122.6k - $185k
...build the backbone of Generative AI cloud at AWS? Do you want to... ...scalability in AI/ML and HPC workloads.... ...The AWS Hardware Engineering team creates server designs for Amazon... ...scale and curious how systems and software... ...Engineering AI / ML development team is a group of...Amazon Web ServiceLocal areaFlexible hours- ...with 5–7 years of experience in AI/ML, Generative AI, and Agentic AI systems to join our Advanced Supply... ...frameworks , and cloud-based ML engineering (AWS) . You will design scalable, production... .../ML, Generative AI & Agentic AI Development Design, build, and deploy AI/...Amazon Web ServiceFull timeTemporary workRemote workFlexible hoursShift work
- ...AI Engineer Client: UST/Applied Materials Location... ...chunking, embedding generation, and retrieval systems About the Role:... ...our team and lead the development of intelligent AI... ...computer science, AI/ML, or related field.... ...with cloud platforms (AWS, Azure, GCP) and containerization...Amazon Web ServiceLocal area
$168k - $270.25k
Site Reliability Engineering (SRE) at NVIDIA is... ...scale production systems with high efficiency... ...Much of our software development focuses on... ...power NVIDIA’s next-generation AI-driven enterprise... ...Security, and AI/ML teams to implement... ...infrastructure-as-code such as AWS CDK, AWS...Amazon Web ServiceFull time- ...Senior Data Scientist – ML, GenAI & Agentic AI Location: Santa... ...closely with product and engineering teams, and can... ...productionize LLM and Generative AI solutions. Develop... ...graphs. Work with AWS Bedrock or comparable... ...improvement of AI/ML systems. Mentor junior data...Amazon Web ServiceFull time
- ...Overview We are seeking expertise in Generative AI and Python to work directly... ...a trusted advisor, hands-on engineer, and delivery lead — driving... ...experience in software engineering, ML engineering, or technical... ...with cloud platforms (AWS, Azure, GCP) and MLOps tools is...Amazon Web ServiceFull time
$180k - $210k
...available to any downstream query engine and use case (from traditional analytics to real-time AI / ML). We are a team of self-... ...have created large-scale data systems and globally distributed platforms... ...including Uber, Snowflake, AWS, Linkedin, Confluent and many more...Amazon Web ServiceWork at officeRemote work- ...Responsibilities AI Integration:... ...and implementing generative AI features... ...Llama. Frontend Development: Building responsive... ...models. Prompt Engineering & RAG:... ...Generation (RAG) systems and crafting effective... ...deploying to AWS, Azure, or GCP.... ...1 year focused on AI/ML. VBeyondAmazon Web Service
$150.4k - $277.6k
...Machine Learning and AI Imagine building... ...at Apple — systems that power how products are made, how engineers think, and ultimately... ...shipping Apple's next-generation sensing technologies... ...cloud AI services (AWS Bedrock, Azure OpenAI... ...projects in AI/ML infrastructure Comfort...Amazon Web ServiceRelocation- ...and influential engineering leader who has... ...ability to scale AI from early... ...machine learning and Generative AI, including... ...deploying AI/ML applications at... ...experience with AI systems, including data... ...scale (e.g., AWS, Azure, GCP).... ...Enabling the development of high-performance...Amazon Web Service
$200k - $210k
Get AI-powered advice on this job and more... ...roadmap for AI/Generative AI platforms. Prioritize... ...to align AI/ML initiatives with company... ...the design and development of scalable,... ...infrastructure. Work with engineering teams to build LLM... ...AI platforms (AWS SageMaker, GCP...Amazon Web ServiceFull timeLocal areaFlexible hours- ...Generative AI Engineer The project is focused on building production-grade GenAI solutions with emphasis... ...leveraging LLMs Prompt engineering (system/tool prompts, function calling,... ...Cloud LLM providers (Azure OpenAI, AWS Bedrock, Vertex AI) Workflow orchestration...Amazon Web Service
- ...Job Title: AI Engineer Location: Santa Clara CA... ...develop, and implement AI/ML models and algorithms... ...cloud platforms (e.g., AWS, Azure, GCP). Collaborate... ...models into existing systems. Optimize AI... ...Participate in the full software development lifecycle, including...Amazon Web ServiceLocal area
$248k - $396.75k
...Site Reliability Engineering (SRE) at NVIDIA... ...production systems with exceptional... ...direction of NVIDIA’s AI Platform... ...NVIDIA’s next-generation AI-driven products... ...the design and development of AI agents,... ...Networking, and AI/ML organizations... ...such as AWS, Azure, or GCP....Amazon Web ServiceFull time$182k - $242k
...Senior Software and AI Engineer Livingston, NJ... ...AI Engineer, IT Systems, you'll design and... ...Codex) to accelerate development, testing, and... ...familiarity with Generative AI frameworks ( LangChain... ...-grade AI/ML/LLM-based applications... ...cloud environments (AWS, Azure, or GCP),...Amazon Web ServiceFull timeTemporary workCasual workWork at officeFlexible hours$215k - $250k
...Data Infrastructure Engineer Onehouse is a... ...analytics to real-time AI / ML). We are a team... ...large-scale data systems and globally... ...Uber, Snowflake, AWS, Linkedin, Confluent... ...productionize the next generation of our data tech... ...across feature development and tech debt with...Amazon Web ServiceOdd jobWork at officeLocal areaRemote workRelocationRelocation package$173.9k - $235.2k
AWS Hardware Engineering Services owns the design, planning, delivery... ...AWS global compute systems. In other words, we’... ...and all of the servers, storage, networking... ...lead the design and development of server products utilizing... ...to create next-generation hardware. You will...Amazon Web ServiceInternshipLocal areaFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Systems Development Engineer, AWS Generative AI & ML Servers. Be the first to apply!
- healthcare systems engineer Cupertino, CA
- operations support system engineer Cupertino, CA
- systems engineer Cupertino, CA
- computer system validation engineer Cupertino, CA
- system performance engineer Cupertino, CA
- advanced systems engineer Cupertino, CA
- operating system engineer Cupertino, CA
- software engineer - web development Cupertino, CA
- senior development engineer Cupertino, CA
- research and development engineer Cupertino, CA



