Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal ML Engineer - Large Scale Training Performance Optimization

AMD

WHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. The Role We are looking for a Principal Machine Learning Engineer to join our Models and Applications team. If you are excited by the challenge of distributed training of large models on a large number of GPUs, and if you are passionate about improving training efficiency while innovating and generating new ideas, then this role is for you. You will be part of a world class team focused on addressing the challenge of training generative AI at scale. The Person The ideal candidate should have experience with distributed training pipelines, be knowledgeable in distributed training algorithms (Data Parallel, Tensor Parallel, Pipeline Parallel, Expert Parallel ZeRO), and be familiar with training large models at scale. Key Responsibilities Train large models to convergence on AMD GPUs at scale. Improve the end-to-end training pipeline performance. Optimize the distributed training pipeline and algorithm to scale out. Contribute your changes to open source. Stay up-to-date with the latest training algorithms. Influence the direction of AMD AI platform. Collaborate across teams with various groups and stakeholders. Preferred Experience Experience with ML/DL frameworks such as PyTorch, JAX, or TensorFlow. Experience with distributed training and distributed training frameworks, such as Megatron-LM, MaxText, TorchTitan. Experience with LLMs or computer vision, especially large models, is a plus. Experience with GPU kernel optimization is a plus. Excellent Python or C++ programming skills, including debugging, profiling, and performance analysis at scale. Experience with ML infra at kernel, framework, or system level. Strong communication and problem-solving skills. Academic Credentials A master's degree or PhD degree in Computer Science, Artificial Intelligence, Machine Learning, or a related field. LOCATION San Jose, CA or Bellevue, WA preferred. May consider other US markets within proximity of US AMD offices. #HYBRID Benefits offered are described: AMD benefits at a glance. AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process. AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here. This posting is for an existing vacancy. #J-18808-Ljbffr

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Principal ML Engineer - Large Scale Training Performance Optimization in San Jose, CA vacancy
  • $291.5k - $369.1k

     ...expertise with the scale and operational...  ...and Cisco’s global engineering capabilities. Our...  ...deployment for large‑scale foundation...  .... Large‑Scale Training & Optimization – Experience optimizing...  ...monitoring of ML models. Strong...  ...sales plans earn performance-based incentive... 
    Training
    Performance
    Full time
    Temporary work
    Local area
    Flexible hours

    Cisco

    Milpitas, CA
    2 days ago
  •  ...experienced Deep Learning Engineer with a specialization in Large Language Models (...  ...• Design and optimize LLM architectures to improve performance, scalability, and efficiency...  ...• Fine-tune pre-trained LLMs on domain-specific...  ...LLMs on large-scale datasets • Proficient... 
    Training
    Performance

    Xforia Inc

    San Jose, CA
    4 days ago
  • $181.1k - $318.4k

     ...ML Research Engineer Video is at the core of nearly all Apple products...  ...lifecycle — including training infrastructure, performance optimization, data, and production...  ...colleagues to design, scale, and harden the systems...  ...distributed training, large-scale data pipelines,... 
    Training
    Performance
    Relocation

    Apple

    Cupertino, CA
    1 day ago
  • $275.8k - $340.5k

     ...the team: The AV ML Infra team at GM builds...  ...productivity of ML engineers, and drive the...  ...Ensures robust model performance by running large-scale simulation workloads...  ...andoptimizeslarge-scale ML training and inference across...  ...Overview: The Principal AI/ML Engineer will... 
    Training
    Performance
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    2 days ago
  • $261.5k - $353.5k

     ...and experienced Principal Machine Learning Engineer to join our Mid...  ...of end-to-end AI/ML solutions that power...  ...innovation, and scale impactful...  ...curation and model training to robust deployment...  ...such as Large Language Models...  ...strong pay for performance rewards approach... 
    Training
    Performance
    Local area

    Intuit

    Mountain View, CA
    2 days ago
  • $239k - $260k

     ...Muon seeks a Principal Mechanical Engineer to join our Mechanical...  ...experience to optimize and mature our...  ...improvements in performance, manufacturability...  ...practices that scale with our growing...  ...onboarding, and training of mechanical engineering...  ...components. Large assembly... 
    Training
    Performance
    Permanent employment
    Full time
    Contract work
    Temporary work
    Remote work
    Flexible hours

    Muon Space

    San Jose, CA
    2 days ago
  • $278.1k - $417.1k

     ...runtime. As our Principal Engineer for On-Device AI...  ...from the moment a trained checkpoint leaves...  ...through export, optimization, and kernel-level...  ...Drive low-level performance work: write and tune...  ...integration between the ML runtime and the...  ...CI, and large device-farm matrices... 
    Training
    Performance
    Work at office
    Worldwide
    Relocation package

    Unity

    Mountain View, CA
    3 days ago
  • $189.3k - $290.7k

     ...transportation on a global scale.  Role:...  .... As a Staff ML Infra Engineer, you will drive...  ...generation, training, evaluation, and...  ...models. From enabling large foundational...  ...pipelines that are performant, easy to use, and...  ...the-art training optimization techniques, including... 
    Training
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    5 days ago
  • $222k - $250k

     ...Principal Mechanical Engineer San Jose, CA About the Role...  ...tradeoffs and optimizations. This hybrid...  ...efforts to improve performance, reduce mass,...  ...onboarding, and training of mechanical engineering...  ...components Large assembly...  ...integration at scale. From climate monitoring... 
    Training
    Performance
    Permanent employment
    Full time
    Temporary work
    Remote work
    Flexible hours
    3 days per week

    Muon Space

    San Jose, CA
    3 days ago
  •  ..., compact chip-scale design, offering...  ...enhanced performance, efficiency, and...  ...is seeking a Principal Silicon Photonics Layout Engineer to actively lead...  ...reproducible and optimized for high-yield,...  ...Ability to parse large-scale foundry PCM...  ..., skills, training, education, market... 
    Training
    Performance

    nEye.ai

    Santa Clara, CA
    29 days ago
  • $145k - $200k

     ...skilled Machine Learning Engineer with deep expertise...  ..., implement, and optimize BEV-based perception...  ...perception models using large-scale datasets and well-defined...  ...a strong focus on performance, robustness, and...  ...Experience with distributed training, high-performance... 
    Training
    Performance

    PlusAI, Inc.

    Santa Clara, CA
    5 days ago
  • $240k - $320k

     ...experts for product engineering of AI-based...  ...As the Senior Principal Engineer, E2E AI Training Framework for Autonomous...  ...the development and optimization of the machinery that...  ...reliability, and performance. The ideal...  ...training AI models with large-scale fleet datasets ~... 
    Training
    Performance
    Full time
    Work experience placement
    Local area
    Flexible hours

    Bosch Group

    Sunnyvale, CA
    1 day ago
  • $184k - $287.5k

     ...Machine Learning Engineers with a background...  ...Development: Design, train, and optimize innovative...  ...coordinate entire ML workflows, covering...  ...metrics, continuous performance instrumentation,...  ...developing for large, complex systems....  ...workflows for large-scale datasets.... 
    Training
    Performance
    Worldwide
    Night shift

    NVIDIA

    Santa Clara, CA
    13 hours ago
  •  ...Role We are hiring AI / ML Platform Engineers to build the platform...  ...systems that support large-scale agent execution, distributed training and inference,...  ...reproducibility, observability, performance, and developer...  ...reusable across kernel optimization, RTL optimization,... 
    Training
    Performance

    AMD

    Santa Clara, CA
    3 days ago
  •  ...Machine Learning Engineer PayPal, Inc....  ...model approaches. Optimize models for performance, accuracy, and...  ...Experience with large language model (...  ...utilizing post-training optimization methods...  ...7. Production ML Systems and...  ...monitoring, and scaling using frameworks... 
    Training
    Performance
    Remote work

    PayPal

    San Jose, CA
    3 days ago
  • $313.06k

     ...brand by hundreds of large institutions seeking...  ...are looking for a Principal AI Engineer to lead the architecture...  ...deployment of large‑scale, LLM‑powered...  ...pipelines for Chatbot performance optimization. Design multi‑level...  ...routing, classifier training, and fallback strategies... 
    Training
    Performance

    United States Digital Space LLC

    San Jose, CA
    2 days ago
  • $160k - $200k

     ...fast-growing teams. As a Senior ML Infrastructure Engineer at Plus, you will design scalable...  ...of data while ensuring optimal performance for both training and inference phases. You will build...  ...will be responsible for managing large-scale GPU clusters. This role offers unparalleled... 
    Training
    Performance

    PlusAI, Inc.

    Santa Clara, CA
    4 days ago
  • $215.28k - $364.32k

     ...Staff Machine Learning Engineer Santa Clara, CA XPENG...  ...data preparation, model training, evaluation, optimization, quantization, and deployment...  ...traffic sign detection performance across diverse real-world...  .... Familiarity with large-scale data pipelines, scenario... 
    Training
    Performance
    Full time

    XPENG

    Santa Clara, CA
    1 day ago
  • $206k - $333k

     ...to build and scale AI with confidence...  ...performance with deep technical...  ...looking for a Principal Engineer to be the technical...  ...: If MLPerf (Training & Inference),...  ...and all popular ML frameworks)...  ...publication. Coordinate optimization tracks with...  ...expertise on large-scale ML... 
    Training
    Performance
    Permanent employment
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    a month ago
  •  ...results-oriented Mid-Level AI/ML Engineer to join our dynamic team....  ...development, model training, validation, and serving....  ...: Design, implement, and optimize solutions utilizing Large Language Models (LLMs) and...  ...skills to evaluate model performance, diagnose issues, and iterate... 
    Training
    Performance
    Permanent employment
    Contract work
    Local area

    Cloud Hybrid Technologies LLC

    San Jose, CA
    2 days ago
  • $164.8k - $226.6k

     ...Principal Product Engineer SiTime is the Precision Timing company...  ..., ensuring performance, resilience and scalability...  ...methods for analyzing large datasets to identify...  ...(DOE) to optimize parameters, improve...  ...experience, education, and training. In addition to base... 
    Training
    Performance

    SiTime

    Santa Clara, CA
    5 days ago
  • $182k - $192k

     ...Job Title: Principal Process Development Engineer Location : This position...  ...define, characterize, optimize, and validate...  ...experience. Ability to perform/oversee complex...  ...up, validation and scale-up in a controlled...  ...experience, education/training, key skills, and... 
    Training
    Performance
    Full time
    Contract work
    Work experience placement

    Imperative Care

    Campbell, CA
    3 days ago
  • $180k - $280k

     ...: Senior Staff SW Engineer (Systems) What you...  ...what it takes to optimize and trade-off various...  ...able to build and scale software...  ...with other software (ML, Systems) and hardware...  ...distributed, high performance software design and...  ...deployment including training, quantization,... 
    Training
    Performance
    Work experience placement

    d-Matrix

    Santa Clara, CA
    4 days ago
  •  ...businesses learn from and optimize in‑person customer...  ...deploy production‑grade ML systems with end‑to‑...  ...preprocessing, model training, deployment, inference...  ...processes for scalability and performance. Qualifications...  ...professional experience in ML engineering. Strong programming... 
    Training
    Performance
    Full time

    Catalyst Labs, LLC

    San Jose, CA
    2 days ago
  • $195k - $230k

     ...meaningful challenges at scale. Together, we...  ...Learning Engineer to help evolve our large-scale recommendation...  ...multi-objective optimization to balance...  ...systems from offline training to online...  ...drift, and system performance in production. AI...  ...-scale data and ML systems (e.g., Spark... 
    Training
    Performance
    Local area
    Work from home

    NewsBreak

    Mountain View, CA
    5 days ago
  • $219k - $351k

     ...communities. Job Title: Principal Engineer, AI System Architect...  ...bandwidth and system-scale communication . By...  ...improvements in performance, efficiency, and scalability...  ...architecture for large-scale computing...  ...DLRMs , and large-scale training and inference systems... 
    Training
    Performance
    Work at office
    Flexible hours

    Samsung Semiconductor

    San Jose, CA
    a month ago
  •  ...applications of AI & ML are bringing...  ...applied science and engineering teams to deliver...  ...scalable, high‑performance AI infrastructure...  ...foundation model training, large language model inference...  ...‑of‑the‑art LLM optimization techniques to...  ...— of large‑scale production AI systems... 
    Training
    Performance
    Full time
    Local area

    SwiftCruit

    San Jose, CA
    2 days ago
  • $214k - $289.5k

     ...Machine Learning Engineer (MLE). Senior...  ...value at Intuit scale. In this role,...  ...define and evolve ML architecture,...  ...teams for training, deployment, monitoring...  ...Deliver within large-scale strategic...  ...and system performance, ensuring continuous...  ...performance optimization. Excellent... 
    Training
    Performance

    Intuit

    Mountain View, CA
    2 days ago
  • $19 - $65 per hour

     ...Responsibilities: Identify Training Bottlenecks: Profile and analyze Bird...  ...: Design and implement high-performance custom compute kernels using CUDA,...  ...training process. Leverage LLMs for Optimization: Explore and integrate Large Language Models (LLMs) to assist... 
    Training
    Performance
    Hourly pay
    Internship

    PlusAI, Inc.

    Santa Clara, CA
    2 days ago
  •  ...About The Role As Chief Engineer , you will oversee all...  ...with asset performance. Responsibilities Operations...  ...equipment issues. Manage and optimize work order queuing...  ...development. Identify training opportunities and provide...  ...expertise in managing large, complex facilities... 
    Training
    Performance
    Ongoing contract
    Work at office
    Local area
    Flexible hours

    FOL Management

    San Jose, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal ML Engineer - Large Scale Training Performance Optimization. Be the first to apply!