Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal ML Engineer - Large Scale Training Performance Optimization

Advanced Micro Devices Inc

WHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. THE ROLE:We are looking for a Principal Machine Learning Engineer to join our Models and Applications team. If you are excited by the challenge of distributed training of large models on a large number of GPUs, and if you are passionate about improving training efficiency while innovating and generating new ideas, then this role is for you. You will be part of a world class team focused on addressing the challenge of training generative AI at scale.THE PERSON:The ideal candidate should have experience with distributed training pipelines, be knowledgeable in distributed training algorithms (Data Parallel, Tensor Parallel, Pipeline Parallel, Expert Parallel ZeRO), and be familiar with training large models at scale.KEY RESPONSIBILITIES:Train large models to convergence on AMD GPUs at scale.Improve the end-to-end training pipeline performance.Optimize the distributed training pipeline and algorithm to scale out.Contribute your changes to open source.Stay up-to-date with the latest training algorithms.Influence the direction of AMD AI platform. Collaborate across teams with various groups and stakeholders.PREFERRED EXPERIENCE:Experience with ML/DL frameworks such as PyTorch, JAX, or TensorFlow.Experience with distributed training and distributed training frameworks, such as Megatron-LM, MaxText, TorchTitan.Experience with LLMs or computer vision, especially large models, is a plus.Experience with GPU kernel optimization is a plus.Excellent Python or C++ programming skills, including debugging, profiling, and performance analysis at scale.Experience with ML infra at kernel, framework, or system levelStrong communication and problem-solving skills.ACADEMIC CREDENTIALS:A master's degree or PhD degree in Computer Science, Artificial Intelligence, Machine Learning, or a related field.LOCATION:San Jose, CA or Bellevue, WA preferred. May consider other US markets within proximity of US AMD offices. #LI-MV1 #HYBRIDBenefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Principal ML Engineer - Large Scale Training Performance Optimization in San Jose, CA vacancy
  • $272k - $431.25k

     ...looking for a Machine Learning (ML) Engineer to join the GPU...  ...data centers for running large scale workloads for ETL, SQL, and ML/DL model training and inference pipelines, spanning...  ...learning solutions for performance prediction and optimization of GPU accelerated enterprise... 
    Training
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $278.1k - $347.6k

     ...runtime. As our Principal Engineer for On-Device AI...  ...from the moment a trained checkpoint leaves...  ...through export, optimization, and kernel-level...  ...Drive low-level performance work: write and tune...  ...integration between the ML runtime and the...  ...CI, and large device-farm matrices... 
    Training
    Performance
    Work at office
    Worldwide
    Relocation package

    Unity

    Mountain View, CA
    2 days ago
  • $153.2k - $234.1k

     ...transportation on a global scale.Role Overview:...  ...learning engineer working on our...  ...-of-the-art optimization, our work is at...  ...the safety and performance of the car,...  ...vehicles.As a Senior ML Infra Engineer,...  ...learning model training and evaluation...  ...building large-scale distributed... 
    Training
    Performance
    Full time
    Work at office
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    1 day ago
  • $189.3k - $290.7k

     ...transportation on a global scale. Role:Are you...  ....As a Staff ML Infra Engineer, you will drive...  ...dataset generation, training, evaluation, and...  .... From enabling large foundational driving...  ...that are performant, easy to use, and...  ...the-art training optimization techniques, including... 
    Training
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    9 hours ago
  • $296.3k

     ...Role:We are seeking a Principal AI Engineer to lead the design...  ...that powers large-scale training and cloud inference...  ...and Pytorch model optimization. This is a highly impactful...  ...optimize core AI/ML platform infrastructure...  ..., scalability, and performance across the AI/ML... 
    Training
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Flexible hours

    General Motors

    Sunnyvale, CA
    1 day ago
  • $165.2k - $223.6k

     ...forefront of maximizing performance for AWS's custom ML accelerators....  ...boundary, our engineers craft high-...  ...counts in delivering optimal performance for...  ...inference and training performance.As part...  ...that are very large, yet our teams...  ...reliability and scaling) of new and existing... 
    Training
    Performance
    Internship
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $193.3k - $261.5k

     ...forefront of maximizing performance for AWS's custom ML accelerators....  ...boundary, our engineers craft high-...  ...counts in delivering optimal performance for...  ...inference and training performance.As part...  ...that are very large, yet our teams...  ...reliability and scaling) of new and existing... 
    Training
    Performance
    Internship
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $272k - $431.25k

     ...seeking exceptional engineers to join our autonomous...  ...Be Doing:Design and train innovative large-scale models—including...  ...environments, ensuring performance, safety, and...  ...learning architectures and optimization techniques.Proven record...  ...production-grade ML models for self-... 
    Training
    Performance
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...of building, scaling, and...  ...seeking Staff and Principal level Machine Learning Engineers to solve these...  ...building, optimizing, and deploying...  ...production ML systems that...  ...infrastructure to train, evaluate,...  ...and system performance for latency,...  ...of large language and... 
    Training
    Performance
    Full time
    Work at office
    Relocation package
    Flexible hours
    Shift work

    Inworld Ai

    Mountain View, CA
    13 hours ago
  • $189k - $300k

     ...transportation on a global scale. The Data Scaling...  ...on and delivers ML models to the...  ...AV product performance through smart use...  ...uses existing very large datasets that GM has...  ...foundation model pre-training and fine-tuning with...  ...team of AI/ML engineers, data scientists and... 
    Training
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    1 day ago
  • $296.3k - $423.9k

     ...transportation on a global scale.  We are looking for a Principal Technical Lead...  ...algorithms and ML systems that...  ...will lead a high-performing team of engineers building ML-driven...  ...engineers, and lead large cross-functional initiatives...  ...scalable training pipelines using large... 
    Training
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    2 days ago
  • $272k - $431.25k

     ...seeking an exceptional Principal Perception Engineer to lead the design...  ...best practices for training and evaluation,...  ...techniques such as large-scale pretraining, distillation...  ...perception performance; analyze large-scale...  ...platforms, including optimization for latency, memory... 
    Training
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    9 hours ago
  • $231.4k - $331.8k

     ...TeamThe Cisco Hardware Engineering Team develops and...  ...innovation across the large CISCO portfolio, including...  ...finding ways to optimize cost, performance, and quality.Your ImpactWe...  ...happen on a global scale. Because our...  ...certifications, and/or training. The full salary range... 
    Training
    Performance
    Full time
    Temporary work
    For subcontractor
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    2 days ago
  •  ...experienced Deep Learning Engineer with a specialization in Large Language Models (...  ...• Design and optimize LLM architectures to improve performance, scalability, and efficiency...  ...• Fine-tune pre-trained LLMs on domain-specific...  ...LLMs on large-scale datasets • Proficient... 
    Training
    Performance

    Xforia Inc

    San Jose, CA
    13 hours ago
  • $114.6k - $234.6k

     ...interested in building large-scale distributed...  ...looking for hands-on engineers with expertise and passion...  ...for owned components; optimizes code and data paths for...  ...processing. -Design performance and load testing.System...  ...ongoing feedback and training to improve skills.... 
    Training
    Performance
    Temporary work
    Flexible hours
    Shift work

    Oracle Corporation

    Santa Clara, CA
    1 day ago
  • $240k - $320k

     ...experts for product engineering of AI-based...  ...As the Senior Principal Engineer, E2E AI Training Framework for Autonomous...  ...the development and optimization of the machinery that...  ...reliability, and performance. The ideal...  ...training AI models with large-scale fleet datasets ~... 
    Training
    Performance
    Full time
    Work experience placement
    Local area
    Flexible hours

    Bosch Group Inc

    Saratoga, CA
    2 days ago
  • $275.8k - $340.5k

     ...the team: The AV ML Infra team at GM builds...  ...productivity of ML engineers, and drive the...  ...Ensures robust model performance by running large-scale simulation workloads...  ...andoptimizeslarge-scale ML training and inference across...  ...Overview: The Principal AI/ML Engineer will... 
    Training
    Performance
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    3 days ago
  • $225k - $325.3k

     ...fully remote and can be performed from any location...  ...Platform, bridging product, engineering, and operations to...  .... Data-Driven Optimization: Monitor engineering...  ...business partners for large-scale platforms. Proficiency...  ...certifications, and/or training. The full salary range... 
    Training
    Performance
    Full time
    Temporary work
    Local area
    Remote work
    Flexible hours

    CISCO Systems

    San Jose, CA
    1 day ago
  • $165.2k - $223.6k

     ...Software Development Engineer to own the...  ...that makes large models run efficiently...  ...high-performance compute kernels...  ...operate at frontier scale with large distributed...  ...can write and optimize low-level code...  ...for a custom ML accelerator...  ...architecture, training/inference lifecycles... 
    Training
    Performance
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  •  ...an experienced AI/ML Engineer to join our Applied...  ...will develop and optimize state-of-the-art AI...  ...into production-scale solutions.KEY RESPONSIBILITIESDevelop...  ..., train, fine-tune, and...  ...pipelines leveraging high-performance GPU computing....  ...models, and large-scale machine learning... 
    Training
    Performance

    AMD

    San Jose, CA
    2 days ago
  • $184k - $287.5k

     ...Machine Learning Engineers with a background...  ...Development: Design, train, and optimize innovative...  ...coordinate entire ML workflows, covering...  ...metrics, continuous performance instrumentation,...  ...developing for large, complex systems....  ...workflows for large-scale datasets.... 
    Training
    Performance
    Full time
    Worldwide
    Night shift

    NVIDIA

    Santa Clara, CA
    5 hours ago
  •  ...THE ROLEWe are seeking a Principal GenAI Inference Optimization Engineer to join our Models and...  ...role focuses on improving performance, efficiency, and...  ...real-world deployment of large-scale models, working across the...  ...architectures.- Experience with ML frameworks (PyTorch, JAX... 
    Performance

    AMD

    San Jose, CA
    1 day ago
  • $206.4k - $379.1k

     ...is seeking a Principal Service Engineer to serve as the...  ...scalable, high-performance generative AI...  ...productizing and scaling a rapidly...  ...organization of ML and services engineers...  ...leading large-scale, GPU-...  ...GenAI workloads (training, inference, and/or optimization).Proven track... 
    Training
    Performance
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Jose, CA
    2 days ago
  • $261.5k - $353.5k

     ...and experienced Principal Machine Learning Engineer to join our Mid...  ...of end-to-end AI/ML solutions that power...  ...innovation, and scale impactful...  ...curation and model training to robust deployment...  ...such as Large Language Models...  ...strong pay for performance rewards approach... 
    Training
    Performance
    Local area

    Intuit

    Mountain View, CA
    3 days ago
  • $150k - $350k

     ...generative AI to assist engineers in RTL design,...  ...Overview We are seeking an ML Systems Engineer to optimize the performance and efficiency of large language model...  ...multi‑node clusters for training and inference that push...  ...Experience with large‑scale ML systems, GPU... 
    Training
    Performance

    ChipAgents

    San Jose, CA
    4 days ago
  • $278.1k - $417.1k

     ...runtime. As our Principal Engineer for On-Device AI...  ...from the moment a trained checkpoint leaves...  ...through export, optimization, and kernel-level...  ...Drive low-level performance work: write and tune...  ...integration between the ML runtime and the...  ...CI, and large device-farm matrices... 
    Training
    Performance
    Work at office
    Worldwide
    Relocation package

    Unity

    Mountain View, CA
    4 days ago
  • $239k - $260k

     ...Role Muon seeks a Principal Mechanical Engineer to join our...  ...industry experience to optimize and mature our...  ...improvements in performance,...  ...practices that scale with our growing...  ...onboarding, and training of mechanical engineering...  ...ADCS components. Large assembly management... 
    Training
    Performance
    Permanent employment
    Full time
    Contract work
    Temporary work
    Remote work
    Flexible hours

    Muon Space

    San Jose, CA
    3 days ago
  • $272k - $431.25k

     ...is seeking a Senior MLOps Engineering Manager to join our...  ...development, and operation of large‑scale, end‑to‑end data and ML pipelines that power NVIDIA...  ...radar—into high‑quality training, evaluation, and...  ...Doing:Lead and grow a high‑performing MLOps engineering group tasked... 
    Training
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...ROLEWe are hiring AI / ML Platform Engineers to build the platform...  ...systems that support large-scale agent execution, distributed training and inference,...  ...framework across kernel optimization, RTL/PPA optimization...  ...reproducibility, observability, performance, and developer... 
    Training
    Performance

    AMD

    Santa Clara, CA
    3 days ago
  • $145k - $200k

     ...skilled Machine Learning Engineer with deep expertise...  ..., implement, and optimize BEV-based perception...  ...perception models using large-scale datasets and well-defined...  ...a strong focus on performance, robustness, and...  ...Experience with distributed training, high-performance... 
    Training
    Performance
    Full time

    Plus.ai

    Santa Clara, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal ML Engineer - Large Scale Training Performance Optimization. Be the first to apply!