Principal AI/ML HPC Specialist Technical Account Manager (STAM) , AWS Enterprise Support, NAMER-Sp
$182.8k - $247.3kAmazonWebServices
As part of the AWS Applied AI Solutions organization, we have a vision to provide business applications, leveraging Amazon’s unique experience and expertise, that are used by millions of companies worldwide to manage day-to-day operations. We will accomplish this by accelerating our customers’ businesses through delivery of intuitive and differentiated technology solutions that solve enduring business challenges. We blend vision with curiosity and Amazon’s real-world experience to build opinionated, turnkey solutions. Where customers prefer to buy over build, we become their trusted partner with solutions that are no-brainers to buy and easy to use.Are you ready to transform how businesses leverage artificial intelligence and machine learning at scale? Join our team and become a strategic partner in delivering Amazon AI/ML solutions that empower global enterprises to innovate, optimize, and achieve unprecedented operational excellence. Amazon Web Services (AWS) is seeking an experienced Principal AI/ML HPC Specialist to join our Technical Account Manager (TAM) team. You'll be at the forefront of solving complex AI HPC implementation challenges, guiding NAMER Resarch labs to enterprise customers through their most ambitious machine learning transformation journeys. By combining deep technical expertise with collaborative problem-solving, you'll help organizations unlock the full potential of artificial intelligence and machine learning technologies — from distributed model training on GPU clusters to production-grade inference at scale. AWS Support includes experts from across AWS who help our customers design, build, operate, and secure their cloud environments. Customers innovate with AWS Professional Services, upskill with AWS Training and Certification, optimize with AWS Support and Managed Services, and meet objectives with AWS Security Assurance Services. Our expertise and emerging technologies include AWS Partners, AWS Sovereign Cloud, AWS International Product, and AI/ML-native solutions. You'll join a diverse team of technical experts in dozens of countries who help customers achieve more with the AWS cloud. Key job responsibilitiesDeliver Strategic Technical Engagements — Lead comprehensive technical deep-dives and performance optimization for enterprise AI/ML workloads, including distributed training cluster architecture using AWS Parallel Computing Service (PCS) and AWS ParallelCluster, the latest GPU-accelerated computing (i.e. P6/P6e , G7/G7e instances), AWS Trainium-based training (Trn3 UltraServers), and multi-node NCCL communication tuning over EFA’s SRD protocol. Architect and Validate Innovative Solutions — Design and implement production-grade AI/ML training and inference solutions leveraging Slurm-based job scheduling, distributed training frameworks (PyTorch FSDP, DDP, DeepSpeed, Megatron-LM), SageMaker HyperPod for managed GPU clusters with automated health checks and node replacement, high-performance parallel storage (Amazon FSx for Lustre), and container runtimes on Deep Learning AMIs (DLAMIs) against reference architectures and HPC lens to ensure performance, reliability, and cost governance at scale.. Architect solutions using P6e UltraServers for multi-trillion parameter frontier models and Trn3 with the AWS Neuron SDK for cost-optimized training and inference. Enable Customer Success — Support customers in implementing business-critical HPC capabilities, including the development of large language model (LLM) (Llama, GPT-class models), physics-informed neural networks (PINNs) and surrogate models, MLOps pipelines, simulation-ML hybrid architectures orchestrated by AWS Step Functions and AWS Batch, distributed data processing, cluster observability, and governance controls for GPU/Trainium-intensive workloads. Enable Business Critical Outcomes — Partner with with service teams to enhance model training throughput, optimize NCCL collective communications, improve GPU/Trainium utilization across multi-node UltraClusters, and drive operational efficiency through proactive monitoring, automated failure recovery (HyperPod health checks), and capacity planning (EC2 Capacity Blocks for ML). Contribute to product roadmap PFR, share refrerence architecture, performance , and benchmarks with broader TAM and Technical communities Serve as Trusted Advisor and Advocate — Develop and nurture technical partnerships with enterprise stakeholders, serving as the trusted advisor for AI/ML infrastructure decisions spanning compute, networking (Elastic Fabric Adapter with SRD), storage, orchestration, and the HPC-to-AI convergence journey. A day in the lifeYour day will be dynamic and impactful, involving deep technical consultations on distributed training architectures, strategic solution design for GPU and Trainium cluster deployments, and collaborative problem-solving across multi-node ML environments. You'll engage with technical leaders, architect innovative AI/ML implementations — from Slurm-managed PCS clusters and SageMaker HyperPod to PyTorch FSDP/DeepSpeed training jobs and Neuron SDK compilation workflows — and provide expert guidance that bridges machine learning infrastructure with business objectives. You will partner with TAMs, SAs, and service teams to provide customers with AWS AI/ML best practice guidance, diving deep into machine learning infrastructure services (PCS, ParallelCluster, HyperPod, Batch), promoting customers' AI/ML workloads to production, developing regional AI/ML strategies, advising on HPC-to-AI convergence patterns (simulation-surrogate loops, physics-informed neural networks), and training field teams on distributed training patterns, GPU/Trainium cluster operations, and the use cases and benefits of artificial intelligence and machine learning at scale. About the teamWe are a collaborative group of technical innovators dedicated to pushing the boundaries of cloud computing and artificial intelligence. Our team thrives on solving complex challenges — from optimizing NCCL all-reduce operations across hundreds of GPUs to architecting elastic training clusters that scale with customer demand. We believe in continuous learning, mutual support, and driving technological advancement. Diverse ExperiencesAmazon values diverse experiences. Even if you do not meet all of the preferred qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn’t followed a traditional path, or includes alternative experiences, don’t let it stop you from applying.Why AWSAmazon Web Services (AWS) is the world’s most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating — that’s why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses.Work/Life BalanceWe value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why flexible work hours and arrangements are part of our culture. When we feel supported in the workplace and at home, there’s nothing we can’t achieve in the cloud.Inclusive Team CultureHere at AWS, it’s in our nature to learn and be curious. Our employee-led affinity groups foster a culture of inclusion that empower us to be proud of our differences. Ongoing events and learning experiences, including our Conversations on Race and Ethnicity and AmazeCon conferences, inspire us to never stop embracing our uniqueness.Mentorship and Career GrowthWe’re continuously raising our performance bar as we strive to become Earth’s Best Employer. That’s why you’ll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional.Basic qualifications- Bachelor's degree- 8+ years of experience in AI/ML, distributed computing, or GPU-accelerated infrastructure (e.g., model training, inference systems, HPC for ML)- 3+ years of hands-on experience designing, implementing, or consulting on large-scale ML training or inference architectures in a customer-facing role- 10+ years of IT development or implementation/consulting in the software, cloud computing, or AI/ML industries- Experience with at least one major deep learning framework (PyTorch, TensorFlow, JAX) in a production or research environment- Demonstrated ability to serve as a trusted technical advisor to enterprise customersPreferred qualification - Deep experience with distributed training techniques including data parallelism, model parallelism, pipeline parallelism, and Fully Sharded Data Parallel (PyTorch FSDP)- Experience with distributed training frameworks such as PyTorch DDP, DeepSpeed, and Megatron-LM for multi-node model training- Hands-on experience with GPU/accelerator cluster infrastructure: NVIDIA Blackwell (GB200, B200, B300), H100/H200 GPUs, AWS Trainium (Trn3/Trn2), NVLink/NVSwitch, InfiniBand or Elastic Fabric Adapter (EFA), and NCCL collective communications tuning- Experience with AWS Neuron SDK (torch-neuronx, neuronx-nemo-megatron) for compiling and optimizing models on Trainium and Inferentia (Inf2) instances- Familiarity with SageMaker HyperPod for managed distributed training clusters including automated health checks, node replacement, and checkpoint-based recovery- Experience with HPC job schedulers (Slurm, PBS, LSF) for orchestrating multi-node ML training workloads- Experience with high-performance parallel file systems (Amazon FSx for Lustre, GPFS/Spectrum Scale) for ML data pipelines- Familiarity with AWS Parallel Computing Service (PCS), AWS ParallelCluster, AWS Batch, or equivalent managed HPC/ML cluster services- Experience training or fine-tuning large language models (LLMs) such as Llama, GPT, or similar transformer architectures at multi-billion parameter scale- Understanding of HPC-AI convergence patterns: simulation-surrogate loops, physics-informed neural networks (PINNs), graph neural networks for molecular property prediction, and data format interoperability (HDF5, VTK, NetCDF to ML-ready tensors)- Knowledge of ML Ops tooling, container orchestration for training (Docker, Enroot, Pyxis), and Deep Learning AMIs (DLAMIs)- Experience with cluster observability and monitoring for GPU/Trainium utilization, training throughput, and job performance (CloudWatch, Prometheus, Grafana)- Experience with EC2 Capacity Blocks for ML, Capacity Reservations, or similar GPU capacity planning strategies- Experience with pipeline orchestration using AWS Step Functions for simulation-ML workflows- Experience with containers, EKS and ECS- Track record of driving operational excellence and proactive risk mitigation for mission-critical AI/ML workloads- AWS certifications (Solutions Architect Professional, Machine Learning Specialty) preferredAmazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at .USA, TX, Austin - 182,800.00 - 247,300.00 USD annuallyUSA, TX, Dallas - 182,800.00 - 247,300.00 USD annuallyUSA, VA, Herndon - 182,800.00 - 247,300.00 USD annuallyUSA, WA, Seattle - 182,800.00 - 247,300.00 USD annually
$131.3k - $177.6k
As part of the AWS Applied AI Solutions organization... ...worldwide to manage day-to-day operations... ...with AWS Support and Managed Services... ...diverse team of technical experts in dozens... ...cloud.As a Technical Account Manager (TAM) at... ...valued member of the Enterprise Support team...Amazon Web ServiceWorldwideFlexible hours$176.1k - $238.2k
...accomplished Principal, WW Storage GTM Specialist to support the growth of AWS Storage... ...you will be accountable for:Setting... ...platform with our enterprise and mid-... ..., measure, manage and coach... ..., and technical challenges,... ...with Data & AI related technologies... ...to, AI/ML, GenAI,...PrincipalAmazon Web ServiceRemote workWorldwideFlexible hours$193.7k - $262k
...with product managers to influence product... ...with other AWS Services.... ...to understand technical requirements,... ...multinational enterprises. You have domain... ...in Generative AI and key ML powered use cases... ...closely with account teams and product... ...When we feel supported in the...PrincipalAmazon Web ServiceLocal areaWorldwideFlexible hours$176.6k - $239k
...growing organizations within Amazon Web Services (AWS)? At AWS Enterprise Support we’re looking for a Sr. Technical Account Manager (TAM) to influence the largest enterprises in... ...as Compute, Storage, Database, Big Data, AI/ML, Networking, Serverless and more. You will...Amazon Web ServiceLocal areaFlexible hours$177k - $239.4k
The Applied AI... ...transforming how specialist expertise is... ...scale across AWS. We are building... ...for a Principal Technical Program Manager to own the... ...engagement, and enterprise knowledge... ...ensure joint accountability across... ...When we feel supported in the workplace... ...with AI/ML...PrincipalAmazon Web ServiceInternshipFlexible hoursDay shift$210.2k - $284.3k
...about Generative AI? In this role... ...and help enterprise customers leverage... ...along with AWS Services.The... ...role will support development of... ...Development and the AI/ML Service teams... ...across technical and non-... ...and workflow management or equivalent... ...Flexible Spending Accounts, Adoption and...PrincipalAmazon Web ServiceLocal areaWorldwideFlexible hours$208.3k - $281.8k
AWS Elemental Inference is an AI-powered video understanding... ...for a Principal Product Manager to own the product... ...of AI/ML product development... ..., and technical feasibility.... ...with strategic enterprise customers to... ...scale. We support technology that... ...Spending Accounts, Adoption and...PrincipalAmazon Web ServiceLocal areaWorldwideFlexible hours$179.9k - $243.4k
...is looking for a Principal Product Manager - Technical to own the Native AI customer and technical... ...for large enterprises or partners- Bachelor... ...limited to, AI/ML, GenAI, Analytics... ...process, including support for the interview... ...Spending Accounts, Adoption and Surrogacy...PrincipalFlexible hours$181.1k - $245k
...embrace agentic AI. We have... ...and make sure AWS is at the... ...join a deeply technical team that... ...technical product managers. You’ll have... ...for a Principal Technical Product... ...that help enterprises and system... ..., including support for the interview... ...Spending Accounts, Adoption...PrincipalAmazon Web ServiceFlexible hours$177k - $239.4k
...strategic technical program leader... ...seeking a Principal Technical Program Manager to drive... ...programs for AWS Mantle—the... ...how we scale AI inference... ...systems serving ML inference... ...secure, enterprise-grade access... ...launch to supporting models from... ...Flexible Spending Accounts, Adoption...PrincipalAmazon Web ServiceFlexible hours$153.6k - $207.8k
AWS Global Sales (AGS) drives... ..., and support for Agentic AI/ML use cases.Are... ...acumen and technical pre-sales expertise... ...the AI/ML Specialists Solution... ...America (NAMER) sales... ...strategic enterprise customers to... ...effectively manage stress and... ...Flexible Spending Accounts, Adoption...Amazon Web ServiceLocal areaWorldwideFlexible hours$235.8k - $307.85k
...control over supporting their family,... ...for a Senior Principal Technical Program Manager (P70) to lead... ...AI & Agent Platforms... ...industry leader in enterprise collaboration... ...practices and account for each candidate... ...on AI/ML platforms or... ...strategies (AWS, GCP, Azure)...PrincipalAmazon Web ServiceWork at officeLocal area$187k - $252.9k
...dynamic ProServe Account Executive (PAE... ...Web Services (AWS). In this role... ..., Engagement Managers, and Delivery... ...to translate technical concepts into... ...APN) to execute enterprise cloud... ...When we feel supported in the workplace... ...of technical specialist, design and architecture...PrincipalAmazon Web ServiceFlexible hoursDay shift$181.1k - $245k
As part of the AWS Applied AI Solutions organization... ...worldwide to manage day-to-day... ...looking for a Principal Product Manager, Technical to own and drive... .... When we feel supported in the workplace... ...for large enterprises or partners- Prior... ...Flexible Spending Accounts, Adoption and...PrincipalAmazon Web ServiceWorldwideFlexible hours$182.8k - $247.3k
...in defining technical strategies for... ...how AI agents consume... ...are seeking a Principal Solutions Architect... ...for AWS Context, the... ...customer-facing specialist Solutions... ...executives, IT management, and... ..., including support for the interview... ...Flexible Spending Accounts, Adoption...PrincipalAmazon Web ServiceLocal areaWorldwideFlexible hoursDay shift$181.1k - $245k
The AWS CX GenAIUX team is building intelligent... ...how people and AI work together on... ...We are seeking a Principal Product Manager - Technical to define the... ...development, ML/AI systems, or other... ..., including support for the interview... ...Flexible Spending Accounts, Adoption and Surrogacy...PrincipalAmazon Web ServiceImmediate startFlexible hours$181.1k - $245k
...service for the AI era. Today, thousands of enterprises and millions... ...or fully managed tools that they... ...profile: a deeply technical product... ...flows through AWS billing — you... ...pricing for ML/AI powered productsAmazon... ..., including support for the... ...Spending Accounts, Adoption and...PrincipalAmazon Web ServiceImmediate startFlexible hoursShift workDay shift$210.2k - $284.3k
AWS is seeking a Principal Customer Success Specialist - AI Driven Digital Product... ...at enterprise scale.You... ...capabilities, change management, and value... ...) and support partner-... ...with direct accountability for... ...articulate technical capabilities... ...understanding of AI/ML...PrincipalAmazon Web ServiceLocal areaFlexible hours$137.9k - $186.5k
AWS is looking for a Sr. Partner... ...Marketing Manager to own the... ...strategic AI Labs and... ...on AWS in NAMER while helping... ...enough AI/ML literacy to... ...their own in technical product... ...teams on an account-based play targeting enterprise customers evaluating... ...including support for the...Amazon Web ServiceInternshipLocal areaWorldwideFlexible hoursShift workDay shift$162.7k - $220.2k
...Generative AI background,... ...help position AWS as the cloud... ...the Worldwide Specialist Organization... ...for Managed Knowledge Bases... ...mid-market accounts to enterprise-level customers... ...specialist, and technical solutions... ...limited to, AI/ML (Artificial... ..., including support for the...Amazon Web ServiceLocal areaWorldwideFlexible hours$162.7k - $220.2k
...workflows with AWS customers... ...for AI Inference at... ...Market (GTM) Specialist to define,... ...GenAI and ML workloads with... ...of strong technical background... ...Architects, Product Managers, Marketing... ...large enterprises, ISVs, Digital... ...teaching account teams how... ...including support for the interview...Amazon Web ServiceLocal areaWorldwideFlexible hours$164k - $221.8k
...customer-obsessed Principal User... ...experiences within our AWS Agentic AI organization.... ...team of AI/ML practitioners,... ...designing for enterprise software or developer... ...or complex technical domains *Track... ..., including support for the... ...Flexible Spending Accounts, Adoption and...PrincipalAmazon Web ServiceWork experience placementWorldwideFlexible hours$164k - $221.8k
AWS AI Services is seeking a Principal UX Designer to lead some... ...across our AI/ML portfolio,... ...sophisticated technical capabilities... ..., Product Management, Customer Service... ...needs of enterprise cloud... ..., including support for the interview... ...Flexible Spending Accounts, Adoption...PrincipalAmazon Web ServiceFlexible hours$147.9k - $200.1k
The AWS Data & AI Partner GTM team supports the world's most innovative... ...Partner GTM Specialist for Data... ...• Build and manage programs for... ..., where technical excellence drives... ...helping enterprises unlock the value... ...to, AI/ML, GenAI, Analytics... ...Spending Accounts, Adoption...Amazon Web ServiceLocal areaWorldwideFlexible hours- ...Amazon Web Services, Inc. seeks a Principal Technical Product Manager for Quick's pricing and billing. You will... ...GTM. You will collaborate with AWS Billing IT, deals desks, contracting... ...extending pricing innovations to other AWS AI services. #J-18808-Ljbffr Jobleads...PrincipalAmazon Web Service
- ...Amazon Web Services (AWS) is seeking a Principal Product Manager - Technical to lead Healthcare AI solutions. You will own one or more production products, shape roadmaps, work with engineers and scientists, and drive revenue and adoption for customer outcomes in healthcare...PrincipalAmazon Web Service
$228.7k - $309.4k
...seeking a Principal Applied Scientist... ...of AI agents and... ...environments at enterprise scale.Key... ...the external technical community... ...that position AWS as the... ...administrators managing both as one... ...applying ML to real-world... ...including support for the interview... ...Spending Accounts, Adoption...PrincipalAmazon Web ServiceLocal areaFlexible hours$147.9k - $200.1k
...with Agentic AI? The Next... ...that enable AWS customers to... ...team as a Sr. Technical Business Development Specialist, Agentic AI... ...mid-market accounts to enterprise-level customers... ...accounts, managing relationships... ...When we feel supported in the... ...limited to, AI/ML (Artificial...Amazon Web ServiceImmediate startWorldwideFlexible hours$145.9k - $234.2k
...seeking a Senior Technical Program Manager (TPM) to drive project... ...management of AI-readiness... ...integrity and personal accountability, strong interpersonal... ...and lifecycle of enterprise AI systems,... ...Machine Learning (ML) heavy environments... ...are designed to support you—at work, at home...Permanent employmentFull timeWork at officeWork from home$148.7k - $201.2k
AWS Infrastructure Services (AIS) owns the design, planning... ...running. We support all AWS data... ...chain specialists, security experts... ...operations managers, and other vital... ...a Senior Technical Program... ...experience driving enterprise supply chain... ...Spending Accounts, Adoption and...Amazon Web ServiceFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal AI/ML HPC Specialist Technical Account Manager (STAM) , AWS Enterprise Support, NAMER-Sp. Be the first to apply!
- ai scientist Seattle, WA
- technical integration manager Seattle, WA
- technical supervisor Seattle, WA
- director of technical services Seattle, WA
- technical manager Seattle, WA
- senior technical director Seattle, WA
- technical superintendent Seattle, WA
- sr technical product manager Seattle, WA
- technical director Seattle, WA
- technical product manager Seattle, WA


