Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal ML Engineer - Large Scale Training Performance Optimization

AMD




WHAT YOU DO AT AMD CHANGES EVERYTHING

At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career.   

THE ROLE:

We are looking for a Principal Machine Learning Engineer to join our Models and Applications team. If you are excited by the challenge of distributed training of large models on a large number of GPUs, and if you are passionate about improving training efficiency while innovating and generating new ideas, then this role is for you. You will be part of a world class team focused on addressing the challenge of training generative AI at scale.

THE PERSON:

The ideal candidate should have experience with distributed training pipelines, be knowledgeable in distributed training algorithms (Data Parallel, Tensor Parallel, Pipeline Parallel, Expert Parallel ZeRO), and be familiar with training large models at scale.

KEY RESPONSIBILITIES:

  • Train large models to convergence on AMD GPUs at scale.
  • Improve the end-to-end training pipeline performance.
  • Optimize the distributed training pipeline and algorithm to scale out.
  • Contribute your changes to open source.
  • Stay up-to-date with the latest training algorithms.
  • Influence the direction of AMD AI platform. 
  • Collaborate across teams with various groups and stakeholders.

PREFERRED EXPERIENCE:

  • Experience with ML/DL frameworks such as PyTorch, JAX, or TensorFlow.
  • Experience with distributed training and distributed training frameworks, such as Megatron-LM, MaxText, TorchTitan.
  • Experience with LLMs or computer vision, especially large models, is a plus.
  • Experience with GPU kernel optimization is a plus.
  • Excellent Python or C++ programming skills, including debugging, profiling, and performance analysis at scale.
  • Experience with ML infra at kernel, framework, or system level
  • Strong communication and problem-solving skills.

ACADEMIC CREDENTIALS:

  • A master's degree or PhD degree in Computer Science, Artificial Intelligence, Machine Learning, or a related field.

LOCATION:

  • San Jose, CA or Bellevue, WA preferred. May consider other US markets within proximity of US AMD offices.

#LI-MV1 

#HYBRID

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

Vacancy posted 13 hours ago
Similar jobs that could be interesting for youBased on the Principal ML Engineer - Large Scale Training Performance Optimization in Bellevue, WA vacancy
  • $177k - $283.2k

     ...AI? Do you love engineering solutions enabling...  ...good?As a Principal ML Engineer at Dedrone...  ...of successfully optimizing models for the edge...  ...distributed training techniques.Implement...  ...model fairness, performance, and platform...  ...and maintaining large-scale distributed platforms... 
    Training
    Performance
    Work experience placement
    Work at office
    Remote work

    Axon

    Seattle, WA
    4 days ago
  •  ...StatesProducts - Engineering /Fulltime /...  ...team.AI Senior Principal Machine...  ...cloud computing, ML and generative...  ...building high-performance, real-time multi...  ...systems at a very large scale.Mentor and...  ...preparation, training, fine-tuning and...  ...developing ML optimization techniques in... 
    Training
    Performance
    Full time
    Shift work

    Extreme Networks, Inc.

    Seattle, WA
    4 days ago
  • $169.8k - $355.4k

    The Senior Principal AI Agent / ML Software Engineer is a Senior Staff-level,...  ...applications used in large-scale, business-critical...  ...services optimized for low latency, high...  ...including reliability, performance, security posture,...  ...GPU inference or training workloads for latency... 
    Training
    Performance
    Temporary work
    Flexible hours

    Oracle Corporation

    Seattle, WA
    3 days ago
  • $266.72k - $350.07k

     ...arelululemon is an innovative performance apparel company for yoga, running, training, and other athletic...  ...with intelligence at scale. The team leads the...  ...As a Principal AI/ML Engineer, you will define and...  ...reliability, scalability, cost optimization, privacy, and governance... 
    Training
    Performance
    Permanent employment
    Full time
    Part time
    Work visa

    Lululemon Athletica

    Seattle, WA
    20 hours ago
  • $264.3k - $322.3k

     ...Team is the engine behind all of...  ...build and scale the foundational...  ...and optimizes compute across...  ...peak workload performance, industry-leading...  ...a seasoned Principal Engineer who...  ...operated large-scale, mission...  ...GPUs for AI/ML workloads....  ...certifications and training, and... 
    Training
    Performance
    Local area
    Worldwide

    DataBricks

    Bellevue, WA
    4 days ago
  • $119.5k - $190k

     ...passionate Machine Learning Engineer II to join our growing team....  ...quality and consistency for model training and inference. Work with large datasets and apply data scaling techniques. Select and build...  ...Fine-tune models to achieve optimal performance and efficiency. Deploy... 
    Training
    Performance
    Local area
    Flexible hours

    Chewy

    Bellevue, WA
    20 hours ago
  • $117k - $152k

     ...Vector builds an offline ML platform that powers...  ...systems operate at scale across batch and...  ...our platform enables large-scale model training, feature generation,...  ...a Machine Learning Engineer to join our Offline...  ...model iterationHelp optimize performance and efficiency across... 
    Training
    Performance
    Full time
    Work at office
    Remote work
    Worldwide

    Unity Technologies

    Bellevue, WA
    2 days ago
  • $175k - $215k

     ...downstream teams on the optimization and integration...  ..., enabling engineers like you to (1) develop...  ...learning from large scale real-world data,...  ...models and model training at scale, to (3)...  ...diverse skill set of ML and geometric...  ...measuring and improving performance of pre-trained... 
    Training
    Performance
    Full time
    Remote work

    Waymo

    Kirkland, WA
    1 day ago
  • $197.3k - $313.7k

     ...Machine Learning Engineer with deep expertise in model training and finetuning to join our ML team. You'll...  ...training frameworks, optimize model...  ...strategies for large language models...  ...serve real users at scale — not just research...  ...assessment of job performance, discipline, termination... 
    Training
    Performance
    Full time

    Salesforce

    Seattle, WA
    3 days ago
  • $266k - $365.75k

     ...their missions.Our engineering teams build highly...  ..., security, and scale that is critical to...  ...data analytics and ML platform in the...  ...building systems at large scale internet companies...  ...and training, and specific work...  ...eligibility for annual performance bonus, equity, and... 
    Training
    Performance
    Local area
    Remote work
    Worldwide

    DataBricks

    Bellevue, WA
    4 days ago
  • $161.14k - $200k

     ...Machine Learning Engineer Join us in...  .... We're a high-performing, fast-moving team...  ...thinking, and scale that don't...  ...: AI and ML Research: Evaluate...  ...architecture and large foundational...  ...strategies to optimize decision-making...  ...include education, training, experience,... 
    Training
    Performance
    Work at office
    Remote work
    Flexible hours
    Shift work
    3 days per week

    Robinhood

    Bellevue, WA
    4 days ago
  • $84.9k - $209.5k

    As a Principal Core Infrastructure Engineer (AI/ML Forward Deployed Infrastructure...  ..., and at scale. Your expertise...  ...be crucial in optimizing our infrastructure for performance, reliability, and...  ...for training, testing, and deploying...  ...in working with large customers. Planning... 
    Training
    Performance
    Temporary work
    Flexible hours
    Shift work

    Oracle Corporation

    Seattle, WA
    4 days ago
  • $126.2k - $264.1k

    Implements machine learning (ML) models for production. Ensures...  ...frameworks to monitor the performance of machine learning models in...  ...readiness for deployment by scaling models, cleaning model code,...  ...alignment with design criteria of trained models and/or systems.-... 
    Training
    Performance
    Temporary work
    Flexible hours
    Shift work

    Oracle Corporation

    Seattle, WA
    3 days ago
  • $276k - $414k

     ...digital services.Snap Engineering teams build fun...  ...re looking for a Principal Machine Learning...  ...join the Content ML team at Snap! We build large-scale recommender...  ...and overall system performance.Stay up to date on...  ...Ability to design, train, deploy, and optimize state-of-the-art... 
    Performance
    Full time
    Live in
    Work at office
    Local area

    Snap

    Seattle, WA
    4 days ago
  • $114.6k - $234.6k

     ...experienced RDMA software engineer with a strong background in high-performance networking,...  ...to design, implement, optimize, and operate critical...  ...infrastructure used by large-scale AI training and inference workloads...  ...Experience supporting AI/ML infrastructure and distributed... 
    Training
    Performance
    Temporary work
    Flexible hours

    Oracle Corporation

    Seattle, WA
    1 day ago
  • $135.2k - $306.4k

     ...implementation, and performance optimization across software components...  ...infrastructure at scale.As a technical leader...  ...teams, mentor senior engineers, and help shape the...  ...networking software for large-scale AI and HPC...  ...used by distributed AI training and inference workloads... 
    Training
    Performance
    Temporary work
    Flexible hours

    Oracle Corporation

    Seattle, WA
    1 day ago
  • $206.4k - $379.1k

     ...is seeking a Principal Service Engineer to serve as the...  ...scalable, high-performance generative AI...  ...productizing and scaling a rapidly...  ...organization of ML and services engineers...  ...leading large-scale, GPU-...  ...GenAI workloads (training, inference, and/or optimization).Proven track... 
    Training
    Performance
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    Seattle, WA
    4 days ago
  • $200k - $250k

     ...Machine Learning Engineering within the Advanced...  ...annotation pipelines, ML Infrastructure and...  ...distributed training infrastructure, model...  ...and grow a high-performance team of data and ML...  ...productionize agentic AI, Large Language Models (...  ...enterprise-scale, auditable ETL pipelines... 
    Training
    Performance
    Temporary work
    Work at office
    Local area

    Metropolis Corp

    Seattle, WA
    4 days ago
  •  ...Implements machine learning (ML) models for production...  ...to monitor the performance of machine learning models...  ...for deployment by scaling models, cleaning model...  ...with design criteria of trained models and/or systems....  ...Learning, Computer Engineering, Mathematics, Physics,... 
    Training
    Performance
    Shift work

    Oracle

    Seattle, WA
    3 days ago
  • $163.9k - $270.88k

     ...Principle Software Engineer with deep...  ...deploying, and scaling generative AI...  ...contributor at the Principal level, you...  ...Engineer, AI/ML will lead the...  ...deployment of large-scale AI models...  ...ensuring they are optimized for performance, cost, and...  ...large-scale model training and deployment... 
    Training
    Performance
    Local area

    Ultimate Software

    Seattle, WA
    3 days ago
  •  ...As a Machine Learning Engineer focused on demand...  ...models and algorithms to optimize inventory management,...  ..., and preprocess large-scale data sets to ensure data...  ...preprocessing, model training, and prediction processes...  ...and optimize the performance of existing time series... 
    Training
    Performance

    TikTok

    Seattle, WA
    1 day ago
  • $202.16k - $368.22k

     ...It’s working fast, at scale, and we’re making a difference...  ..., build, and optimize the performance of large and highly scalable ML systems, algorithms, strategies...  ...preparation, feature engineering, hyperparameter tuning...  ...experience.Build and train ML models, deep learning... 
    Training
    Performance
    Full time
    Work experience placement

    TikTok

    Seattle, WA
    1 day ago
  • $148.2k - $300.96k

     ...working fast, at scale, and we’re...  ...machine learning (ML) and data mining...  ...experience and optimize consumption of...  ...initiatives from engineering viewpoints....  ...the ML model's performance and optimize performance...  ...;Optimizing, training, and deploying...  ...and deploying large-scale machine... 
    Training
    Performance
    Full time
    Work experience placement

    TikTok

    Seattle, WA
    1 day ago
  • $168.1k - $227.4k

     ...Machine Learning Engineer with expertise in...  ...system, production ML systems, and...  ...end deployment at scale of Generative AI...  ...reliability and performance• Establish scalable...  ...automated processes for large-scale data...  ...model deployment, training, and optimization• Document processes... 
    Training
    Performance
    Internship
    Worldwide
    Flexible hours

    Amazon

    Seattle, WA
    4 days ago
  • $151.8k - $265.35k

     ...Machine Learning Engineers for our GenAI...  ...scalable, high-performance generative AI systems...  ...pipelines, optimize models for...  ...and mentor other ML engineers.Job ResponsibilitiesDesign...  ...for enterprise-scale model...  ...) model training and inference—...  ...leading large-scale, GPU-intensive... 
    Training
    Performance
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    Seattle, WA
    4 days ago
  • $168.1k - $227.4k

     ...modeling problem - agent performance depends on the model,...  ...Machine Learning Engineer to build and own core...  ...reinforcement learning training systems, self-learning...  ...that runs unattended at scale.The work is concrete....  ...turn training jobs, and large GPU formations on shared... 
    Training
    Performance
    Internship
    Flexible hours
    Day shift

    Amazon

    Bellevue, WA
    2 days ago
  • $184.5k

     ...Machine Learning Engineer role is part of...  ...builds and optimizes the machine learning...  ..., deploy, and scale robust models...  ...the quality and performance of our...  ...problems into clear ML‑driven solutions...  ...including model training, evaluation, and...  ..., working with large‑scale data pipelines... 
    Training
    Performance
    Full time

    Expedia

    Seattle, WA
    20 hours ago
  • $186.1k - $300.55k

     ...Machine Learning Engineer to redefine how we...  ...in tandem with Large Language Models (...  ...infrastructure to perform safe, automated remediation...  ...on petabyte-scale datasets to training, deployment, and...  ...applying them to optimization or control...  ...or Go), CI/CD for ML, and experience deploying... 
    Training
    Performance
    Contract work
    Work at office
    Local area
    Remote work
    2 days per week

    DocuSign

    Seattle, WA
    4 days ago
  • $106.9k - $160.4k

     ...functions. As we continue to scale AI across the...  ...are seeking a skilled ML Engineer to design, build, and...  ...for developing, training, deploying, and operationalizing...  ..., including pricing optimization, industrial AI,...  ...architectures, ensuring performance, reliability, and... 
    Training
    Performance
    Full time
    Temporary work

    Weyerhaeuser

    Seattle, WA
    3 days ago
  • $156.75k - $250.8k

     ...Machine Learning Engineer to join a new...  ...multimodal data at scale: the data and...  ...pipelines, the training and evaluation...  ...pipelines for large-scale video and...  ...multimodal workloads, optimizing latency,...  ...systems across the ML lifecycle: data...  ..., design, performance metrics, code,... 
    Training
    Performance
    Work experience placement
    Work at office
    Remote work

    Axon

    Seattle, WA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal ML Engineer - Large Scale Training Performance Optimization. Be the first to apply!