Principal ML Engineer - Large Scale Training Performance Optimization
AMD
WHAT YOU DO AT AMD CHANGES EVERYTHING
At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career.
THE ROLE:
We are looking for a Principal Machine Learning Engineer to join our Models and Applications team. If you are excited by the challenge of distributed training of large models on a large number of GPUs, and if you are passionate about improving training efficiency while innovating and generating new ideas, then this role is for you. You will be part of a world class team focused on addressing the challenge of training generative AI at scale.
THE PERSON:
The ideal candidate should have experience with distributed training pipelines, be knowledgeable in distributed training algorithms (Data Parallel, Tensor Parallel, Pipeline Parallel, Expert Parallel ZeRO), and be familiar with training large models at scale.
KEY RESPONSIBILITIES:
- Train large models to convergence on AMD GPUs at scale.
- Improve the end-to-end training pipeline performance.
- Optimize the distributed training pipeline and algorithm to scale out.
- Contribute your changes to open source.
- Stay up-to-date with the latest training algorithms.
- Influence the direction of AMD AI platform.
- Collaborate across teams with various groups and stakeholders.
PREFERRED EXPERIENCE:
- Experience with ML/DL frameworks such as PyTorch, JAX, or TensorFlow.
- Experience with distributed training and distributed training frameworks, such as Megatron-LM, MaxText, TorchTitan.
- Experience with LLMs or computer vision, especially large models, is a plus.
- Experience with GPU kernel optimization is a plus.
- Excellent Python or C++ programming skills, including debugging, profiling, and performance analysis at scale.
- Experience with ML infra at kernel, framework, or system level
- Strong communication and problem-solving skills.
ACADEMIC CREDENTIALS:
- A master's degree or PhD degree in Computer Science, Artificial Intelligence, Machine Learning, or a related field.
LOCATION:
- San Jose, CA or Bellevue, WA preferred. May consider other US markets within proximity of US AMD offices.
#LI-MV1
#HYBRID
Benefits offered are described: AMD benefits at a glance.
AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.
AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.
This posting is for an existing vacancy.
$177k - $283.2k
...AI? Do you love engineering solutions enabling... ...good?As a Principal ML Engineer at Dedrone... ...of successfully optimizing models for the edge... ...distributed training techniques.Implement... ...model fairness, performance, and platform... ...and maintaining large-scale distributed platforms...TrainingPerformanceWork experience placementWork at officeRemote work- ...StatesProducts - Engineering /Fulltime /... ...team.AI Senior Principal Machine... ...cloud computing, ML and generative... ...building high-performance, real-time multi... ...systems at a very large scale.Mentor and... ...preparation, training, fine-tuning and... ...developing ML optimization techniques in...TrainingPerformanceFull timeShift work
$169.8k - $355.4k
The Senior Principal AI Agent / ML Software Engineer is a Senior Staff-level,... ...applications used in large-scale, business-critical... ...services optimized for low latency, high... ...including reliability, performance, security posture,... ...GPU inference or training workloads for latency...TrainingPerformanceTemporary workFlexible hours$266.72k - $350.07k
...arelululemon is an innovative performance apparel company for yoga, running, training, and other athletic... ...with intelligence at scale. The team leads the... ...As a Principal AI/ML Engineer, you will define and... ...reliability, scalability, cost optimization, privacy, and governance...TrainingPerformancePermanent employmentFull timePart timeWork visa$264.3k - $322.3k
...Team is the engine behind all of... ...build and scale the foundational... ...and optimizes compute across... ...peak workload performance, industry-leading... ...a seasoned Principal Engineer who... ...operated large-scale, mission... ...GPUs for AI/ML workloads.... ...certifications and training, and...TrainingPerformanceLocal areaWorldwide$119.5k - $190k
...passionate Machine Learning Engineer II to join our growing team.... ...quality and consistency for model training and inference. Work with large datasets and apply data scaling techniques. Select and build... ...Fine-tune models to achieve optimal performance and efficiency. Deploy...TrainingPerformanceLocal areaFlexible hours$117k - $152k
...Vector builds an offline ML platform that powers... ...systems operate at scale across batch and... ...our platform enables large-scale model training, feature generation,... ...a Machine Learning Engineer to join our Offline... ...model iterationHelp optimize performance and efficiency across...TrainingPerformanceFull timeWork at officeRemote workWorldwide$175k - $215k
...downstream teams on the optimization and integration... ..., enabling engineers like you to (1) develop... ...learning from large scale real-world data,... ...models and model training at scale, to (3)... ...diverse skill set of ML and geometric... ...measuring and improving performance of pre-trained...TrainingPerformanceFull timeRemote work$197.3k - $313.7k
...Machine Learning Engineer with deep expertise in model training and finetuning to join our ML team. You'll... ...training frameworks, optimize model... ...strategies for large language models... ...serve real users at scale — not just research... ...assessment of job performance, discipline, termination...TrainingPerformanceFull time$266k - $365.75k
...their missions.Our engineering teams build highly... ..., security, and scale that is critical to... ...data analytics and ML platform in the... ...building systems at large scale internet companies... ...and training, and specific work... ...eligibility for annual performance bonus, equity, and...TrainingPerformanceLocal areaRemote workWorldwide$161.14k - $200k
...Machine Learning Engineer Join us in... .... We're a high-performing, fast-moving team... ...thinking, and scale that don't... ...: AI and ML Research: Evaluate... ...architecture and large foundational... ...strategies to optimize decision-making... ...include education, training, experience,...TrainingPerformanceWork at officeRemote workFlexible hoursShift work3 days per week$84.9k - $209.5k
As a Principal Core Infrastructure Engineer (AI/ML Forward Deployed Infrastructure... ..., and at scale. Your expertise... ...be crucial in optimizing our infrastructure for performance, reliability, and... ...for training, testing, and deploying... ...in working with large customers. Planning...TrainingPerformanceTemporary workFlexible hoursShift work$126.2k - $264.1k
Implements machine learning (ML) models for production. Ensures... ...frameworks to monitor the performance of machine learning models in... ...readiness for deployment by scaling models, cleaning model code,... ...alignment with design criteria of trained models and/or systems.-...TrainingPerformanceTemporary workFlexible hoursShift work$276k - $414k
...digital services.Snap Engineering teams build fun... ...re looking for a Principal Machine Learning... ...join the Content ML team at Snap! We build large-scale recommender... ...and overall system performance.Stay up to date on... ...Ability to design, train, deploy, and optimize state-of-the-art...PerformanceFull timeLive inWork at officeLocal area$114.6k - $234.6k
...experienced RDMA software engineer with a strong background in high-performance networking,... ...to design, implement, optimize, and operate critical... ...infrastructure used by large-scale AI training and inference workloads... ...Experience supporting AI/ML infrastructure and distributed...TrainingPerformanceTemporary workFlexible hours$135.2k - $306.4k
...implementation, and performance optimization across software components... ...infrastructure at scale.As a technical leader... ...teams, mentor senior engineers, and help shape the... ...networking software for large-scale AI and HPC... ...used by distributed AI training and inference workloads...TrainingPerformanceTemporary workFlexible hours$206.4k - $379.1k
...is seeking a Principal Service Engineer to serve as the... ...scalable, high-performance generative AI... ...productizing and scaling a rapidly... ...organization of ML and services engineers... ...leading large-scale, GPU-... ...GenAI workloads (training, inference, and/or optimization).Proven track...TrainingPerformanceFull timeTemporary workLocal areaWorldwide$200k - $250k
...Machine Learning Engineering within the Advanced... ...annotation pipelines, ML Infrastructure and... ...distributed training infrastructure, model... ...and grow a high-performance team of data and ML... ...productionize agentic AI, Large Language Models (... ...enterprise-scale, auditable ETL pipelines...TrainingPerformanceTemporary workWork at officeLocal area- ...Implements machine learning (ML) models for production... ...to monitor the performance of machine learning models... ...for deployment by scaling models, cleaning model... ...with design criteria of trained models and/or systems.... ...Learning, Computer Engineering, Mathematics, Physics,...TrainingPerformanceShift work
$163.9k - $270.88k
...Principle Software Engineer with deep... ...deploying, and scaling generative AI... ...contributor at the Principal level, you... ...Engineer, AI/ML will lead the... ...deployment of large-scale AI models... ...ensuring they are optimized for performance, cost, and... ...large-scale model training and deployment...TrainingPerformanceLocal area- ...As a Machine Learning Engineer focused on demand... ...models and algorithms to optimize inventory management,... ..., and preprocess large-scale data sets to ensure data... ...preprocessing, model training, and prediction processes... ...and optimize the performance of existing time series...TrainingPerformance
$202.16k - $368.22k
...It’s working fast, at scale, and we’re making a difference... ..., build, and optimize the performance of large and highly scalable ML systems, algorithms, strategies... ...preparation, feature engineering, hyperparameter tuning... ...experience.Build and train ML models, deep learning...TrainingPerformanceFull timeWork experience placement$148.2k - $300.96k
...working fast, at scale, and we’re... ...machine learning (ML) and data mining... ...experience and optimize consumption of... ...initiatives from engineering viewpoints.... ...the ML model's performance and optimize performance... ...;Optimizing, training, and deploying... ...and deploying large-scale machine...TrainingPerformanceFull timeWork experience placement$168.1k - $227.4k
...Machine Learning Engineer with expertise in... ...system, production ML systems, and... ...end deployment at scale of Generative AI... ...reliability and performance• Establish scalable... ...automated processes for large-scale data... ...model deployment, training, and optimization• Document processes...TrainingPerformanceInternshipWorldwideFlexible hours$151.8k - $265.35k
...Machine Learning Engineers for our GenAI... ...scalable, high-performance generative AI systems... ...pipelines, optimize models for... ...and mentor other ML engineers.Job ResponsibilitiesDesign... ...for enterprise-scale model... ...) model training and inference—... ...leading large-scale, GPU-intensive...TrainingPerformanceFull timeTemporary workLocal areaWorldwide$168.1k - $227.4k
...modeling problem - agent performance depends on the model,... ...Machine Learning Engineer to build and own core... ...reinforcement learning training systems, self-learning... ...that runs unattended at scale.The work is concrete.... ...turn training jobs, and large GPU formations on shared...TrainingPerformanceInternshipFlexible hoursDay shift$184.5k
...Machine Learning Engineer role is part of... ...builds and optimizes the machine learning... ..., deploy, and scale robust models... ...the quality and performance of our... ...problems into clear ML‑driven solutions... ...including model training, evaluation, and... ..., working with large‑scale data pipelines...TrainingPerformanceFull time$186.1k - $300.55k
...Machine Learning Engineer to redefine how we... ...in tandem with Large Language Models (... ...infrastructure to perform safe, automated remediation... ...on petabyte-scale datasets to training, deployment, and... ...applying them to optimization or control... ...or Go), CI/CD for ML, and experience deploying...TrainingPerformanceContract workWork at officeLocal areaRemote work2 days per week$106.9k - $160.4k
...functions. As we continue to scale AI across the... ...are seeking a skilled ML Engineer to design, build, and... ...for developing, training, deploying, and operationalizing... ..., including pricing optimization, industrial AI,... ...architectures, ensuring performance, reliability, and...TrainingPerformanceFull timeTemporary work$156.75k - $250.8k
...Machine Learning Engineer to join a new... ...multimodal data at scale: the data and... ...pipelines, the training and evaluation... ...pipelines for large-scale video and... ...multimodal workloads, optimizing latency,... ...systems across the ML lifecycle: data... ..., design, performance metrics, code,...TrainingPerformanceWork experience placementWork at officeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal ML Engineer - Large Scale Training Performance Optimization. Be the first to apply!
- senior director engineering Bellevue, WA
- principal developer Bellevue, WA
- principal engineer Bellevue, WA
- senior civil engineer project manager Bellevue, WA
- director software engineering Bellevue, WA
- general engineer Bellevue, WA
- engineering director Bellevue, WA
- chief engineer Bellevue, WA
- hotel chief engineer Bellevue, WA
- data center chief engineer Bellevue, WA


