Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Hardware Systems Manager, Annapurna Labs, Trainium Machine Learning Fleet Operations

$175.1k - $236.9k

Amazon Locker

Annapurna Labs designs silicon and software that accelerates innovation. Customers choose us to create cloud solutions that solve challenges that were unimaginable a short time ago, even yesterday. Our custom chips, accelerators, and software stacks enable us to take on technical challenges that have never been seen before, and deliver results that help our customers change the world.In Annapurna Labs we are at the forefront of hardware/software co-design not just in Amazon Web Services (AWS) but across the industry. The Machine Learning Acceleration Fleet Operations Team is looking for a technical leader to manage a team of 5-10 engineers and own operations across multiple ML server platforms spanning tens of thousands of hosts globally.We are seeking a manager who combines strong technical depth in hardware systems and software development with proven people leadership. You will build and grow a high-performing team, set technical direction for fleet-scale automation and tooling, and drive operational excellence across some of the most advanced server hardware in existence. You will define your team's 6-12 month roadmap, influence org-level priorities, and represent fleet operations in VP-level reviews. You are equally comfortable debugging a complex hardware failure as you are coaching an engineer through a career development conversation.Our team has end to end ownership of some of the most advanced server hardware in the world. We drive technical debug efforts and write truly massive scale autonomous software to monitor, optimize, and remediate machine learning hardware. Come define how we operate the future of ML infrastructure.Key job responsibilities - Build, hire, mentor, and grow a team of platform development engineers responsible for ML fleet operations across multiple accelerator platforms - Define team roadmap and technical strategy for fleet health, automation, and data infrastructure — balancing near-term operational demands against long-term engineering investments - Drive operational excellence by establishing metrics, SLAs, and processes that maximize platform sellability and customer experience- Partner with hardware engineering, software engineering, and product teams to prioritize debug efforts and translate fleet learnings into permanent design fixes- Own escalation paths for critical fleet incidents and lead cross-functional war rooms to resolution- Influence org-level priorities by surfacing fleet-wide patterns and advocating for systemic improvements across the ML hardware portfolio- Raise the bar on team software practices — ensuring automation is maintainable, tested, documented, and reusable at scale- Represent fleet operations in executive reviews, providing data-driven narratives on platform health and roadmapA day in the lifeAs a Manager on the MLA Fleet Operations team, you set the direction for how your team keeps the world's most advanced ML accelerators healthy at scale.You start each day with your people — holding 1:1s, coaching engineers through ambiguous technical problems, removing blockers, and ensuring the team is focused on the highest-impact work. From there, you review fleet health with the team, understanding which issues are trending, which investigations need unblocking, and where to allocate engineering effort for maximum customer impact. You partner with hardware design teams to advocate for fleet-informed design changes and with service teams to align on deployment schedules. You balance long-term automation investments against near-term operational demands, and you represent your team's work to senior leadership with clear data and crisp narratives. When critical incidents arise, you lead the response — marshaling the right people, driving root cause, and ensuring corrective actions land.About the teamThe MLA Fleet Operations team was formed to maintain an exceptionally high quality bar for our fleet of advanced machine learning accelerators and server products. We perfect the customer experience by developing scalable software for rapid incident response times and data visualization as well as diving deep into hardware issues as they arise.Basic qualifications- Bachelor's degree in computer science, electrical engineering, or related field- 2+ years of engineering team management experience- Knowledge of and proficiency in the use of Python scripting language- Experience with general troubleshooting/debugging of hardware- Experience designing, building, operating, and managing large-scale distributed systems or web services- 7+ years of experience in systems engineering, platform engineering, SRE, or hardware operationsPreferred qualification - Experience in automating, deploying, and supporting large-scale infrastructure- Experience in server technologies such as, thermal, mechanical, power, and signal integrity- Experience working cross-functionally across several teams both technical and non-technical- Experience with GPU, ML accelerator, or high-performance computing hardware- Experience managing teams through ambiguity on new or unreleased productsAmazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at .USA, TX, Austin - 175,100.00 - 236,900.00 USD annually

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the AI Hardware Systems Manager, Annapurna Labs, Trainium Machine Learning Fleet Operations in Austin, TX vacancy
  • $136k - $184k

    Annapurna Labs designs silicon and software that accelerates...  ...the forefront of hardware/software co-design...  ...the industry. The Machine Learning Acceleration Fleet Operations Team is looking...  ...responsible for system remediation,...  ...testing story and manage tradeoffs between... 
    Operations
    Fleet
    Permanent employment
    Internship
    Flexible hours

    Amazon

    Austin, TX
    10 hours ago
  • $143.7k - $194.4k

     ...Acceleration, Trainium AI Systems, Annapurna Labs Annapurna Labs...  ...behind AWS machine learning acceleration. Trainium...  ...any of that hardware carries...  ...doing across the fleet. You will build...  ...will write and operate production code...  ...source control management, build processes... 
    Operations
    Fleet
    Internship
    Flexible hours
    Shift work
    Night shift
    Day shift

    Jobleads-US

    Austin, TX
    1 day ago
  • $183k - $247.6k

    The Annapurna AI Manufacturing, Quality...  ...Annapurna Labs focused on Machine Learning products...  ...computer systems to...  ...Inferentia, and Trainium families...  ...firmware, hardware, and silicon...  ..., Fleet Ops Systems...  ...effectively manage stress and...  ...business operations and the Company... 
    Operations
    Fleet
    Work experience placement
    Local area
    Flexible hours

    Amazon

    Austin, TX
    1 day ago
  • $159.2k - $215.3k

    The Annapurna AI Manufacturing, Quality...  ...Annapurna Labs, building...  ...between hardware design,...  ...whether each system is capable...  ...health management, including...  ...Inferentia, and Trainium families...  .... Machine Learning Annapurna...  ...Development, Fleet Ops Systems...  ...manufacturing operations or... 
    Operations
    Fleet
    Contract work
    Overseas
    Flexible hours

    Amazon

    Austin, TX
    2 days ago
  • $136k - $184k

     ...the future of AI infrastructure!...  ...AI acceleration hardware, deployed across...  ...global server fleet. Annapurna Labs (our organization...  ..., sleds and systems inside our machine learning training systems...  ...manufacturing operations throughout the...  ...Inferentia, and Trainium families of processors... 
    Operations
    Fleet
    Contract work
    Remote work
    Flexible hours

    Amazon

    Austin, TX
    3 hours ago
  • $183k - $247.6k

    Annapurna Labs (our organization within...  ...ML/AI Mechanical...  ...and ongoing fleet operations. You'll design...  ...a Cloud Hardware...  ...interconnect systems Define thermal...  ...accelerators — Trainium for...  ...precision-machined component...  ...leading, or managing more junior...  ...parental leave. Learn more about... 
    Operations
    Fleet
    Local area
    Flexible hours

    Amazon

    Austin, TX
    1 day ago
  • $137.3k - $185.7k

    Annapurna Labs designs silicon and software...  ...skills in both hardware and software....  ...solutions on machine learning products and support...  ...in AWS fleet through their...  ...reliability of products operating in the fleet....  ...of our systems, including identifying...  ...with managing ODM- 5+ years... 
    Operations
    Fleet
    Flexible hours

    Amazon

    Austin, TX
    4 days ago
  • $193.3k - $261.5k

    Annapurna Labs designs silicon and...  ...forefront of hardware/software...  ...power our Machine Learning servers and...  ...expanding fleet of...  ...seamless system integration...  ...fabric for Trainium chips. Day...  ...generation AI possible,...  ...source control management, build processes...  ..., and operations experience... 
    Operations
    Fleet
    Internship
    Local area
    Flexible hours

    Amazon

    Austin, TX
    1 day ago
  • $159.2k - $215.3k

     ...manufacturing of AI Servers and Systems based on Trainium chips across...  ...Team in AWS Annapurna Labs focused on Machine Learning products...  ...of all operational issues during...  ...power and heat management requirement-...  ..., firmware, hardware, and silicon...  ...Development, Fleet Ops Systems,... 
    Operations
    Fleet
    Flexible hours

    Amazon

    Austin, TX
    10 hours ago
  • $159.2k - $215.3k

    Within the Trainium...  ...datacenter operational performance...  ...Preparedness Lab Leader...  ...fleet performance...  ...Trainium systems to identify...  ...will manage early lifecycle...  ..., Hardware Design,...  ...and field learnings into...  ...processors.Machine Learning Annapurna functions...  ...with AI/ML acceleration... 
    Operations
    Fleet
    Flexible hours

    Amazon

    Austin, TX
    1 day ago
  • In Annapurna Labs we are at the forefront of hardware/software accelerator solutions...  ...industry. The Annapurna AI Systems team owns design,...  ...reliability and fleet operations end to end....  ...can launch on. The Trainium Manufacturing,...  ...responsible for managing server and rack manufacturing... 
    Operations
    Fleet

    Annapurna Labs (U.S.) Inc.

    Austin, TX
    1 day ago
  • $193.3k - $261.5k

    Annapurna Labs, an AWS organization with...  ..., hardware design, verification...  ..., and operations to tackle technical...  ...generation machine learning...  ...dominate in AI training and...  ...chip), inter-system connections...  ...to improve fleet healthBasic...  ...source control management, build... 
    Operations
    Fleet
    Internship
    Local area
    Flexible hours

    Amazon

    Austin, TX
    3 days ago
  • $143.7k - $194.4k

    Annapurna Labs is at the forefront of hardware/software co-design, not just...  ...highest-performing Machine Learning servers, from...  ...control systems for ML Acceleration...  ...and manage various protocols...  ...integration, and fleet metrics to deploy...  ...testing, and operations experience-... 
    Operations
    Fleet
    Internship
    Flexible hours

    Amazon

    Austin, TX
    2 days ago
  • $183k - $247.6k

     ...development and management of Compute,...  ...services.Custom machine learning chips are at...  ...heart of our Trainium machine learning...  ...the pre-silicon hardware/software co-...  ...next generation AI chips.You will...  ...software, system development or...  ...safeguard business operations and the... 
    Operations
    Local area
    Flexible hours

    Amazon

    Austin, TX
    1 day ago
  • $173.9k - $235.2k

     ...world's largest fleet of AI/ML accelerator...  ...intersection of hardware, software, and...  ...are seeking a Systems Development Engineer...  ...experts to learn from across hardware...  ...software, and operations teams.Key job...  ...manufacturing, lab, and production...  ...4. Build and manage tests covering... 
    Operations
    Fleet
    Permanent employment
    Internship
    Local area
    Worldwide
    Flexible hours
    Night shift
    Day shift

    Amazon

    Austin, TX
    1 day ago
  • $159.2k - $215.3k

    Annapurna Labs (our organization within AWS...  ...forefront of hardware/software co-design...  ...industry. Our Machine Learning Accelerator (...  ...data center operation of AWS’s next...  ...scale, debug fleet wide issues, and...  ...learning chips into Trainium servers. In...  ...Engineering, Systems Engineering,... 
    Operations
    Fleet
    Flexible hours

    Amazon

    Austin, TX
    10 hours ago
  • $183k - $247.6k

     ...the future of AI? Join the team...  ..., deliver, and operate next-generation...  ...opportunity to build the systems that define...  ...of software, hardware, and network...  ...experts, operations managers, and other...  ...to our existing fleet with the goal of...  ...our nature to learn and be curious.... 
    Operations
    Fleet
    Local area
    Flexible hours

    Amazon

    Austin, TX
    10 hours ago
  • $157.3k - $212.8k

    Annapurna Labs (our organization within AWS UC) designs...  ...We are seeking a Hardware Design Engineer...  ...a member of the Machine Learning Acceleration team...  ...to SoCs to full system testingPreferred...  ..., effectively manage stress and work safely...  ...business operations and the Company’s... 
    Operations
    Local area
    Flexible hours

    Amazon

    Austin, TX
    3 days ago
  • $148.7k - $201.2k

     ...of Generative AI cloud at AWS? Do...  ...delivering and operating AWS cloud offerings...  ...you.The AWS Hardware Engineering...  ...building intelligent systems that drive the...  ..., operations managers, and other vital...  ...decisions, and fleet health...  ...our nature to learn and be curious.... 
    Operations
    Fleet
    Internship
    Local area
    Flexible hours

    Amazon

    Austin, TX
    1 day ago
  • $173k - $245k

    Meta is seeking a Hardware Systems Engineer to...  ...next-generation AI and high-performance...  ...and data center operations, partnering with...  ...diagnostics, and remote management of AI...  ...system telemetry and fleet health data to identify...  ...intelligence and machine learning technologies in connection... 
    Operations
    Fleet
    Hourly pay
    Local area
    Remote work

    Jobleads-US

    Austin, TX
    2 days ago
  • $193.3k - $261.5k

     ...advanced custom machine learning chips. This...  ...the Trainium and Inferentia...  ...scale AI training at...  ....This role operates at the intersection...  ...of hardware and software...  ...stability of systems that support...  ...organization within Annapurna Labs (AWS). Our...  ...control management, build... 
    Operations
    Local area
    Immediate start
    Flexible hours

    Amazon

    Austin, TX
    10 hours ago
  • $168.1k - $227.4k

    Annapurna Labs designs silicon and software...  ...forefront of hardware/software co-...  ...power our Machine Learning servers and view...  ...million chip fleet of machine...  ...and existing systems experience- 5...  ...source control management, build processes...  ...testing, and operations experience-... 
    Operations
    Fleet
    Internship
    Flexible hours

    Amazon

    Austin, TX
    3 hours ago
  • $143.4k

     ...experience in system design or...  ...source control management, build processes...  ...testing, and operations. Preferred...  ...firmware for Machine Learning Accelerator (...  ...Firmware Hardware Support...  ...More: Annapurna Labs designs silicon...  ...our growing fleet of ML products... 
    Operations
    Fleet
    Full time
    Internship
    Flexible hours

    Annapurna Labs (U.S.) Inc.

    Austin, TX
    4 days ago
  •  .... ASIC RTL Design Engineer, AI Hardware The Tesla AI Hardware team...  ...chips that enhance Tesla's machine learning capabilities. This includes...  ...networks using data from Tesla's fleet, supporting technologies...  ...microarchitecture specifications. Define system-level functional... 
    Fleet

    Jobleads-US

    Austin, TX
    3 days ago
  • $148.7k - $201.2k

    AWS-Annapurna team develops the silicon...  ...most advanced machine learning accelerator...  ...provide best hardware platform for our...  ...with various system issues and data...  ...Technical Program Manager to own end-to-...  ..., silicon operations, manufacturing...  ...with CPU, GPU, AI/ML accelerator... 
    Operations
    Flexible hours

    Amazon

    Austin, TX
    3 days ago
  • $173.9k - $235.2k

     ...experienced Senior Systems Development...  ...tooling, and fleet health...  ...accelerated (AI/ML) compute fleet...  ...zero-touch operations where automation...  ...of hardware, software, system...  ...Engineers, TPMs, Managers, Principals)...  ...health across lab and production...  ...leave. Learn more about our... 
    Operations
    Fleet
    Internship
    Local area
    Worldwide
    Flexible hours

    Amazon

    Austin, TX
    10 hours ago
  • $183k - $247.6k

    Annapurna Labs designs silicon and software...  ...Custom SoCs (System on Chip) live...  ...heart of AWS Machine Learning servers. As a...  ...of hardware in our data centers...  ...Inferentia, Trainium Systems (our...  ...leading, or managing more junior engineers...  ...business operations and the Company... 
    Operations
    Local area
    Flexible hours

    Amazon

    Austin, TX
    3 days ago
  • $143.7k - $194.4k

    Annapurna Labs, an AWS organization with development...  ...engineering, hardware design,...  ...software, and operations to tackle...  ...next-generation machine learning accelerators that...  ...dominate in AI training and inference...  ..., and system-level power budgeting...  ...power management schemes, characterize... 
    Operations
    Internship
    Flexible hours

    Amazon

    Austin, TX
    1 day ago
  • $183k - $247.6k

     ...Generation of AI accelerator compute systems? Lead...  ...scale. At AWS Trainium we develop...  ...Silicon to Hardware to Software...  ...AWS Trainium Machine Learning...  ...in the AWS fleet through its...  ...of products operating in the fleet...  ...Experience using lab equipment such...  ...effectively manage stress and... 
    Operations
    Fleet
    Local area
    Flexible hours

    Amazon

    Austin, TX
    4 days ago
  • $176.6k - $247.3k

     ...validating that our hardware and software...  ...team The Annapurna Labs ML...  ...team builds the Trainium and Inferentia...  ...trains and serves machine learning models across...  ...using System Verilog and UVM...  ...effectively manage stress and work...  ...safeguard business operations and the... 
    Operations
    Local area
    Flexible hours

    Jobleads-US

    Austin, TX
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Hardware Systems Manager, Annapurna Labs, Trainium Machine Learning Fleet Operations. Be the first to apply!