Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Sr. System Development Engineer, Edge & High Performance Accelerator Servers for AI/ML

$173.9k - $235.2k

Amazon Locker

We are seeking an experienced Senior Systems Development Engineer to lead the development of automation software, diagnostic tooling, and fleet health infrastructure for our server platforms. You will work across multiple teams and organizations to build scalable, reliable systems that keep our edge and accelerated (AI/ML) compute fleet healthy — with a vision toward zero-touch operations where automation detects, diagnoses, and resolves issues without human intervention.You will be a technical leader solving complex architectural problems that may not be well-defined in advance. You will own your team's systems, proactively identify deficiencies, write scalable and robust code to solve issues before they impact customers. You will decompose large, difficult server testability, reliability, and diagnosis problems into straightforward tasks and components — leading delivery yourself and through others in parallel — using a combination of hardware, software, system design, processor architecture, diagnostics, and operations knowledge.You will collaborate with a variety of roles (SDEs, SDETs, Mechanical/Electrical/Hardware Engineers, TPMs, Managers, Principals) and organizations through server conception, test validation, qualification, launch, and operations — driving high quality and reliability into current and future designs for AWS server solutions. You will also work closely with ODMs and Design Partners to ensure our tooling, diagnostics, and automation requirements are met throughout the hardware development lifecycle (NPI).Key job responsibilitiesFleet Health & Predictive Infrastructure- Build and own the automation infrastructure responsible for the health of the server fleet across edge and accelerator (AI/ML) compute platforms- Design and implement predictive failure detection systems using telemetry, sensor data, error trending, and log correlation to identify hardware issues before they cause customer impact- Drive toward zero-touch operations — building automation that detects, diagnoses, triages, and remediates hardware and software faults without human intervention- Develop monitoring tools, dashboards, and alerting systems to provide real-time visibility into fleet health across lab and production environments- Define and track fleet health metrics (failure rates, mean time to detect, mean time to repair, first-time fix rate, predictive accuracy)Debugging & Troubleshooting- Debug and resolve complex system-level issues across storage, compute, GPU, networking in production environments- Troubleshoot Linux boot and runtime failures across x86 and ARM architectures, including PCIe, power, NIC, NVMe, and GPU subsystems- Perform root cause analysis on hardware failures — correlating across firmware, kernel, driver, and physical layer to isolate faults- Build diagnostic tooling that automates root cause identification and reduces reliance on manual triage- Improve manufacturing throughput and yield through test optimization Systems Development & Automation- Lead the definition and development of software, automation, and enabling tools for server hardware programs; track and report progress- Design and build scalable system-level software with focus on durability, availability, security, and diagnostics- Develop and maintain device drivers for Linux on ARM and x86 architectures- Build automation solutions using modern programming languages (Python, Ruby, Java, C/C++, etc.)- Work with OS internals, storage subsystems, and accelerator/GPU software stacks in Linux-based environments- Build, manage, and deploy CI/CD pipelines for rapid deployment of code changes to org-owned and customer-owned systemsCross-Team Collaboration- Work across internal HWEng teams to ensure new server hardware addresses data path and control path functionality needed by dependent service teams- Work closely with internal customers to identify early any potential problems onboarding new servers — edge or accelerated compute — into their ecosystem- Engage with ODMs and design partners on testability, diagnostic, and automation requirements during hardware design and development- Contribute to server design to improve robustness, testability, diagnosability, and reliability- Partner with datacenter operations teams to close the loop between field failures and design improvementsA day in the lifeSystems Development Engineers in AWS Hardware Engineering wear many hats. From orchestration tooling development to hardware integration to kernel driver debugging, we dive deep into problems across the breadth of AWS. Our teams are directly responsible for launching and maintaining server hardware in the fleet — including edge servers and AI/ML accelerator servers with GPUs. Located in Seattle and Cupertino, we work with internal development teams, ODMs, and design partners to deliver servers deployed in datacenters worldwide.Basic qualifications- 6+ years of non-internship professional software development experience- 6+ years of systems design, software development, operations, automation, and process improvement experience- 6+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience- 5+ years of programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, Ruby experience- Experience with Linux/Unix- Experience leading the design, build and deployment of complex and performant (reliable and scalable) software solutions in productionPreferred qualification - Knowledge of engineering practices and patterns for the full software/hardware/networks development life cycle, including coding standards, code reviews, source control management, build processes, testing, certification, and livesite operations- Experience taking a leading role in building complex software or computing infrastructure that has been successfully delivered to customers- Experience building predictive failure detection or proactive remediation systems at fleet scale- Experience with Linux kernel driver development- Experience with storage, compute, GPU/accelerator platforms, including driver integration, diagnostics, or performance validation- Familiarity with server hardware architecture, BMC/IPMI, firmware, PCIe topology and hardware diagnostics- Experience working with ODMs or hardware design partners through the product development lifecycle- Experience building zero-touch or self-healing automation for large-scale infrastructure- Experience working in large-scale datacenter or cloud environments- Track record of rapidly coming up to speed on new engineering disciplines and making impactful decisions- Experience with hardware bring-up, validation, and fleet-wide deployment- Familiarity with telemetry pipelines, anomaly detection, and operational metrics at scale- Familiarity with manufacturing workflows and yield improvement optimizationAmazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at .USA, CA, Cupertino - 173,900.00 - 235,200.00 USD annuallyUSA, TX, Austin - 151,200.00 - 204,600.00 USD annuallyUSA, WA, Seattle - 151,200.00 - 204,600.00 USD annually

Vacancy posted 7 hours ago
Similar jobs that could be interesting for youBased on the Sr. System Development Engineer, Edge & High Performance Accelerator Servers for AI/ML in Austin, TX vacancy
  • $173.9k - $235.2k

     ...Generative AI at AWS?...  ...innovation in AI/ML and HPC...  ...limits of performance,...  ...build the systems that define...  ...Hardware Engineering team of software...  ...baremetal server hardware up...  .... Driving high quality and...  ...designs for AWS Accelerated server...  ...software development experience... 
    Senior
    Performance
    Internship
    Local area
    Flexible hours

    Amazon

    Austin, TX
    7 hours ago
  • $171k - $231.4k

     ...all of the servers, storage, networking...  ...and network engineers, supply...  ...leading edge technologies...  ...standards on performance, quality, cost...  ...design and development of server products...  ...You’ll have high standards...  ...deprecating systems that are no...  ..., gpu (AI/ML), or storage... 
    Senior
    Performance
    Local area
    Flexible hours

    Amazon

    Austin, TX
    7 hours ago
  • $183k - $247.6k

     ...the future of AI? Join the...  ...innovation in AI/ML and HPC...  ...the limits of performance, efficiency,...  ...build the systems that define...  ...and network engineers, supply chain...  ...for complex high performance server and/or accelerator server and rack...  ...system development on top of your... 
    Senior
    Performance
    Local area
    Flexible hours

    Amazon

    Austin, TX
    7 hours ago
  •  ...everywhere — powering the AI edge revolution with...  ..., research, development, production,...  ...Edge AI Applied ML Engineer with deep experience...  ...optimize, and deploy highly efficient on-...  ...This role will help accelerate the shift to on-...  ...correctness, performance, and usability.... 
    Senior
    Performance
    Full time
    Shift work

    Ambiq Micro, Inc

    Austin, TX
    11 hours ago
  • $183k - $247.6k

     ...the Next Generation of AI accelerator compute systems? Lead bleeding-edge HW development projects? Have you...  ...experienced Lead System Design Engineers to build the next generation of our cloud server infrastructure,...  ...chips, designed for high-performance AI training. As a... 
    Senior
    Performance
    Local area
    Flexible hours

    Amazon

    Austin, TX
    4 days ago
  •  ...products that accelerate next-generation...  ...—from AI and data centers...  ...and embedded systems. Grounded in a...  ...boundaries of high-performance computing, graphics...  ...track record of development and leadership...  ..., and product engineers to deliver...  ...Mathematics AI/ML Experience with... 
    Senior
    Performance

    AMD

    Austin, TX
    2 days ago
  • $129.2k - $174.8k

    We're seeking a Systems Development Engineer to join the Unified Workcell...  ...This is a hands-on, high-impact role where you...  ...that manage Amazon's edge device fleet — over a...  ...fleets — from AI perception systems to...  ...contributing to the resilience, performance, and cost efficiency... 
    Performance
    Full time
    Temporary work
    Seasonal work
    Worldwide
    Flexible hours
    Day shift

    Amazon

    Austin, TX
    2 days ago
  • $124k - $280k

     ...data and analytics engineering focus on...  ...implementing advanced AI and ML solutions to...  ...algorithms, models, and systems to enable...  ...develop and sustain high performing, diverse, and...  ...you will lead the development of AI, GenAI,...  ...Claude code to accelerate development and... 
    Senior
    Performance
    Full time
    H1b

    PwC

    Austin, TX
    3 days ago
  • $178.4k - $267.6k

     ...Inc.Job Area:Engineering Group, Engineering...  ...Summary:AI Performance Engineer (Cloud...  ...Inference Acceleration.  We are hiring...  ...from cutting-edge research and development to commercial...  ...implement high-level kernels...  ...architecture, ML accelerators,...  ...distributed systems. Strong communication... 
    Senior
    Performance
    Work experience placement
    Work from home

    Qualcomm

    Austin, TX
    2 days ago
  •  ...do.Schwab’s AI Strategy & Transformation...  ...product, engineering, strategy...  ..., and accelerate delivery...  ...opportunity to join a high-profile team...  ...for cutting-edge GenAI...  ...critical AI systems. Above all,...  ...early in the development lifecycle, promoting...  ..., performance, and scalability... 
    Senior
    Performance
    Full time

    The Charles Schwab Corporation

    Austin, TX
    1 day ago
  • $182.75k - $236.5k

     ...Senior Principal Systems Development Engineer Our customers’ system requirements...  ...are usually highly complex. Bringing together...  ...at the very cutting edge of technology to meet...  ..., scalable, and high‑performance data ecosystems....  ...governance frameworks, and AI/ML environments. Drive... 
    Senior
    Performance

    Dell Technologies

    Austin, TX
    3 days ago
  •  ...The AI Engineering and Productivity...  ..., accelerate decision making...  ...support Product Development...  ...requirements and system specifications...  ...quality, performance, and...  ...), and AI/ML integrations...  ...Write high-quality, performant...  ...e.g., SQL Server, Oracle,...  ...Edge (Preferred... 
    Senior
    Performance
    Local area
    Work from home
    Relocation package

    General Motors

    Austin, TX
    5 days ago
  • $224k - $356.5k

     ...gaming, and accelerated computing for...  ...potential of AI to define the...  ...efficiency on NVIDIA edge AI hardware....  ...platform — performance, CI/CD...  ...inference recipe development, performance...  ...Science, Computer Engineering, Electrical...  ...computing, ML systems, or high-performance... 
    Senior
    Performance
    Full time
    Local area

    Nvidia

    Austin, TX
    1 day ago
  •  ...supercomputing, high-performance computing, cloud, and AI. Whether you...  ...and Edge AI platforms...  ...world-class engineering teams, and drive...  ...intelligent systems across a wide...  ...and career development programs.Foster...  ...of AI/ML frameworks,...  ..., NPUs, and accelerators.Excellent communication... 
    Senior
    Performance

    AMD

    Austin, TX
    2 days ago
  •  ...products that accelerate next-generation...  ...experiences—from AI and data...  ...gaming and embedded systems. Grounded in a...  ...credible with Engineering and TME; effective...  ...executable plansA high ownership...  ...judgment in balancing performance, supportability...  .../ML platformsHPC systemsEnterprise... 
    Senior
    Performance

    AMD

    Austin, TX
    3 days ago
  • $100k

     ...industry on cutting-edge AI technology, revolutionizing performance expectations,...  ...developed a high performance RISC...  ...technologies to drive the development of next-...  ...to aid compiler engineers in understanding...  ...developer tools into ML frameworks and CI systems to improve... 
    Senior
    Performance
    Permanent employment
    Full time

    Tenstorrent

    Austin, TX
    11 hours ago
  • $135.2k - $306.4k

     ...Principal Network Development Engineer (IC5) to...  ...GPU- and accelerator-based clusters...  ...OCI’s AI superclusters...  ...requirements for RDMA performance, cluster...  ...networking, systems engineering,...  ...control, SR-IOV, queueing...  ...networking for AI/ML workloads (e...  ...analysis in high-scale... 
    Senior
    Performance
    Temporary work
    Flexible hours

    Oracle Corporation

    Austin, TX
    7 hours ago
  • $207.8k - $281.2k

     ...generation of Edge AI by combining industry...  ...that accelerate customer innovation...  ...Solution Engineering (SE), ensuring...  ...product definition, development, launch, and...  ..., system architecture,...  ...platforms, is highly desirable.Experience...  ...supports both high performance and personal wellbeing... 
    Senior
    Performance
    Work at office
    Local area
    Relocation

    ARM

    Austin, TX
    4 days ago
  • $169k - $228.6k

     ...looking for a highly motivated...  ...Architect to help accelerate our growing...  ...and AI-powered solutions...  ...business development team. This will...  ...senior engineers at both top...  ...computing, AI/ML implementation...  ...raising our performance bar as we strive...  ...computing, systems engineering,... 
    Senior
    Performance
    Flexible hours

    AmazonWebServices

    Austin, TX
    1 day ago
  • $239.6k - $324.1k

     ...software that accelerates innovation...  ...our performance bar as we...  ...integration engineers responsible...  ...deliverables in our ML...  ...expertise in high-...  ...teamCustom SoCs (System on Chip) live...  ...Learning servers. As a...  ...in leading edge nodesPreferred...  ...SOC development- Track record... 
    Senior
    Performance
    Local area
    Flexible hours

    Amazon

    Austin, TX
    2 days ago
  •  ...software engineering into the emerging...  ...agentic AI development to join...  ...that accelerate ICON's construction...  ...AI systems that automate...  ...identify high-leverage AI...  ...leading edge of the agentic...  ...AI system performance ~ Strong...  ...Protocol) server...  ...background in ML fundamentals... 
    Senior
    Performance
    Full time
    For contractors

    Icon

    Austin, TX
    11 hours ago
  • $189.2k - $372.9k

     ...intelligence (AI)...  ...scientists and engineers’ partner with...  ...demand for AI development in new...  ...strategy for AI, ML, and GenAI...  ...delivering high-quality,...  ...ensuring cutting-edge use of LLMs...  ...clients to accelerate their...  ...autonomous systems and edge AI...  ...organizational performance.... 
    Senior
    Performance
    Local area
    Visa sponsorship

    Deloitte

    Austin, TX
    1 day ago
  • $128.7k - $261.3k

     ...self-driving systems, to move us toward...  .... For the AI Kernels &...  ...turning cutting‑edge perception,...  ...export, kernel development, and performance engineering so that every...  ...cycle on our accelerators translates...  ...team builds high‑performance GPU...  ...our on‑vehicle ML inference for... 
    Senior
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Austin, TX
    2 days ago
  • Senior Systems Development Engineer - HPCJoin us as a Senior Systems Development Engineer on our High Performance Computing (HPC) Solutions Engineering team in Austin, Texas to do the best...  ...HPC Solutions team as well as the HPC & AI Innovations Lab team utilizing resources... 
    Senior
    Performance
    Work experience placement
    Start working today
    Flexible hours

    Dell Technologies

    Austin, TX
    3 days ago
  •  ...great products that accelerate next-generation...  ...experiences—from AI and data centers,...  ...and embedded systems. Grounded in a culture...  ...SoCs, and high-performance edge platforms, we develop...  ...Physical AI Software Engineer to develop next-...  ...software development.Understanding of... 
    Senior
    Performance

    AMD

    Austin, TX
    3 days ago
  • $152k - $241.5k

     ...graphics, PC gaming, and accelerated computing for more...  ...Intelligence, High-Performance Computing and...  ...critically important systems running while...  ...harness the power of AI to deliver groundbreaking...  ...operations (AIOps/ML-driven signals)...  ....Mentored other engineers and influenced... 
    Senior
    Performance
    Full time

    Nvidia

    Austin, TX
    2 days ago
  • $183k - $247.6k

     ...Learning Acceleration (MLA) team...  ...power today’s AI workloads...  ...yield & performance - it’s still...  ...a highly reliable,...  ...Verification Engineers to build the...  ...our cloud server chips. Our...  ...experience using System Verilog...  ...testbench development including:...  ..., GPU, or ML... 
    Senior
    Performance
    Local area
    Work from home
    Flexible hours

    Amazon

    Austin, TX
    4 days ago
  • $184k - $287.5k

     ...PC gaming, and accelerated computing for more...  ...potential of AI to define the...  ...talented Software Engineers to work on...  ..., and software development. The successful...  ...driven software systems, ideally applied...  ...Proficiency with modern ML frameworks (...  .... As we highly value diversity... 
    Senior
    Full time

    Nvidia

    Austin, TX
    4 days ago
  • $124k - $280k

     ...data and analytics engineering focus on...  ...implementing advanced AI and ML solutions to...  ...algorithms, models, and systems to enable...  ...develop and sustain high performing, diverse, and...  ...you will lead the development of innovative AI...  ...and how AI accelerates outcomes- Proven... 
    Senior
    Full time
    H1b

    PwC

    Austin, TX
    2 days ago
  • $151.2k - $204.6k

     ...Join Amazon's OTS Supply Chain team as a Sr. Systems Development Engineer to transform technology infrastructure through innovative Gen-AI powered solutions to critical supply...  ...decisions balancing system reliability, performance, and feature velocity while mentoring engineers... 
    Senior
    Performance
    Full time
    Temporary work
    Seasonal work
    Flexible hours
    Night shift

    Amazon

    Austin, TX
    7 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Sr. System Development Engineer, Edge & High Performance Accelerator Servers for AI/ML. Be the first to apply!