Principal ML Engineer - Large Scale Training Performance Optimization
AMD
WHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. The Role We are looking for a Principal Machine Learning Engineer to join our Models and Applications team. If you are excited by the challenge of distributed training of large models on a large number of GPUs, and if you are passionate about improving training efficiency while innovating and generating new ideas, then this role is for you. You will be part of a world class team focused on addressing the challenge of training generative AI at scale. The Person The ideal candidate should have experience with distributed training pipelines, be knowledgeable in distributed training algorithms (Data Parallel, Tensor Parallel, Pipeline Parallel, Expert Parallel ZeRO), and be familiar with training large models at scale. Key Responsibilities Train large models to convergence on AMD GPUs at scale. Improve the end-to-end training pipeline performance. Optimize the distributed training pipeline and algorithm to scale out. Contribute your changes to open source. Stay up-to-date with the latest training algorithms. Influence the direction of AMD AI platform. Collaborate across teams with various groups and stakeholders. Preferred Experience Experience with ML/DL frameworks such as PyTorch, JAX, or TensorFlow. Experience with distributed training and distributed training frameworks, such as Megatron-LM, MaxText, TorchTitan. Experience with LLMs or computer vision, especially large models, is a plus. Experience with GPU kernel optimization is a plus. Excellent Python or C++ programming skills, including debugging, profiling, and performance analysis at scale. Experience with ML infra at kernel, framework, or system level. Strong communication and problem-solving skills. Academic Credentials A master's degree or PhD degree in Computer Science, Artificial Intelligence, Machine Learning, or a related field. LOCATION San Jose, CA or Bellevue, WA preferred. May consider other US markets within proximity of US AMD offices. #HYBRID Benefits offered are described: AMD benefits at a glance. AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process. AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here. This posting is for an existing vacancy. #J-18808-Ljbffr
$291.5k - $369.1k
...expertise with the scale and operational... ...and Cisco’s global engineering capabilities. Our... ...deployment for large‑scale foundation... .... Large‑Scale Training & Optimization – Experience optimizing... ...monitoring of ML models. Strong... ...sales plans earn performance-based incentive...TrainingPerformanceFull timeTemporary workLocal areaFlexible hours- ...experienced Deep Learning Engineer with a specialization in Large Language Models (... ...• Design and optimize LLM architectures to improve performance, scalability, and efficiency... ...• Fine-tune pre-trained LLMs on domain-specific... ...LLMs on large-scale datasets • Proficient...TrainingPerformance
$181.1k - $318.4k
...ML Research Engineer Video is at the core of nearly all Apple products... ...lifecycle — including training infrastructure, performance optimization, data, and production... ...colleagues to design, scale, and harden the systems... ...distributed training, large-scale data pipelines,...TrainingPerformanceRelocation$275.8k - $340.5k
...the team: The AV ML Infra team at GM builds... ...productivity of ML engineers, and drive the... ...Ensures robust model performance by running large-scale simulation workloads... ...andoptimizeslarge-scale ML training and inference across... ...Overview: The Principal AI/ML Engineer will...TrainingPerformanceLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours$261.5k - $353.5k
...and experienced Principal Machine Learning Engineer to join our Mid... ...of end-to-end AI/ML solutions that power... ...innovation, and scale impactful... ...curation and model training to robust deployment... ...such as Large Language Models... ...strong pay for performance rewards approach...TrainingPerformanceLocal area$239k - $260k
...Muon seeks a Principal Mechanical Engineer to join our Mechanical... ...experience to optimize and mature our... ...improvements in performance, manufacturability... ...practices that scale with our growing... ...onboarding, and training of mechanical engineering... ...components. Large assembly...TrainingPerformancePermanent employmentFull timeContract workTemporary workRemote workFlexible hours$278.1k - $417.1k
...runtime. As our Principal Engineer for On-Device AI... ...from the moment a trained checkpoint leaves... ...through export, optimization, and kernel-level... ...Drive low-level performance work: write and tune... ...integration between the ML runtime and the... ...CI, and large device-farm matrices...TrainingPerformanceWork at officeWorldwideRelocation package$189.3k - $290.7k
...transportation on a global scale. Role:... .... As a Staff ML Infra Engineer, you will drive... ...generation, training, evaluation, and... ...models. From enabling large foundational... ...pipelines that are performant, easy to use, and... ...the-art training optimization techniques, including...TrainingPerformanceFull timeLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours$222k - $250k
...Principal Mechanical Engineer San Jose, CA About the Role... ...tradeoffs and optimizations. This hybrid... ...efforts to improve performance, reduce mass,... ...onboarding, and training of mechanical engineering... ...components Large assembly... ...integration at scale. From climate monitoring...TrainingPerformancePermanent employmentFull timeTemporary workRemote workFlexible hours3 days per week- ..., compact chip-scale design, offering... ...enhanced performance, efficiency, and... ...is seeking a Principal Silicon Photonics Layout Engineer to actively lead... ...reproducible and optimized for high-yield,... ...Ability to parse large-scale foundry PCM... ..., skills, training, education, market...TrainingPerformance
$145k - $200k
...skilled Machine Learning Engineer with deep expertise... ..., implement, and optimize BEV-based perception... ...perception models using large-scale datasets and well-defined... ...a strong focus on performance, robustness, and... ...Experience with distributed training, high-performance...TrainingPerformance$240k - $320k
...experts for product engineering of AI-based... ...As the Senior Principal Engineer, E2E AI Training Framework for Autonomous... ...the development and optimization of the machinery that... ...reliability, and performance. The ideal... ...training AI models with large-scale fleet datasets ~...TrainingPerformanceFull timeWork experience placementLocal areaFlexible hours$184k - $287.5k
...Machine Learning Engineers with a background... ...Development: Design, train, and optimize innovative... ...coordinate entire ML workflows, covering... ...metrics, continuous performance instrumentation,... ...developing for large, complex systems.... ...workflows for large-scale datasets....TrainingPerformanceWorldwideNight shift- ...Role We are hiring AI / ML Platform Engineers to build the platform... ...systems that support large-scale agent execution, distributed training and inference,... ...reproducibility, observability, performance, and developer... ...reusable across kernel optimization, RTL optimization,...TrainingPerformance
- ...Machine Learning Engineer PayPal, Inc.... ...model approaches. Optimize models for performance, accuracy, and... ...Experience with large language model (... ...utilizing post-training optimization methods... ...7. Production ML Systems and... ...monitoring, and scaling using frameworks...TrainingPerformanceRemote work
$313.06k
...brand by hundreds of large institutions seeking... ...are looking for a Principal AI Engineer to lead the architecture... ...deployment of large‑scale, LLM‑powered... ...pipelines for Chatbot performance optimization. Design multi‑level... ...routing, classifier training, and fallback strategies...TrainingPerformance$160k - $200k
...fast-growing teams. As a Senior ML Infrastructure Engineer at Plus, you will design scalable... ...of data while ensuring optimal performance for both training and inference phases. You will build... ...will be responsible for managing large-scale GPU clusters. This role offers unparalleled...TrainingPerformance$215.28k - $364.32k
...Staff Machine Learning Engineer Santa Clara, CA XPENG... ...data preparation, model training, evaluation, optimization, quantization, and deployment... ...traffic sign detection performance across diverse real-world... .... Familiarity with large-scale data pipelines, scenario...TrainingPerformanceFull time$206k - $333k
...to build and scale AI with confidence... ...performance with deep technical... ...looking for a Principal Engineer to be the technical... ...: If MLPerf (Training & Inference),... ...and all popular ML frameworks)... ...publication. Coordinate optimization tracks with... ...expertise on large-scale ML...TrainingPerformancePermanent employmentTemporary workCasual workWork at officeFlexible hours- ...results-oriented Mid-Level AI/ML Engineer to join our dynamic team.... ...development, model training, validation, and serving.... ...: Design, implement, and optimize solutions utilizing Large Language Models (LLMs) and... ...skills to evaluate model performance, diagnose issues, and iterate...TrainingPerformancePermanent employmentContract workLocal area
$164.8k - $226.6k
...Principal Product Engineer SiTime is the Precision Timing company... ..., ensuring performance, resilience and scalability... ...methods for analyzing large datasets to identify... ...(DOE) to optimize parameters, improve... ...experience, education, and training. In addition to base...TrainingPerformance$182k - $192k
...Job Title: Principal Process Development Engineer Location : This position... ...define, characterize, optimize, and validate... ...experience. Ability to perform/oversee complex... ...up, validation and scale-up in a controlled... ...experience, education/training, key skills, and...TrainingPerformanceFull timeContract workWork experience placement$180k - $280k
...: Senior Staff SW Engineer (Systems) What you... ...what it takes to optimize and trade-off various... ...able to build and scale software... ...with other software (ML, Systems) and hardware... ...distributed, high performance software design and... ...deployment including training, quantization,...TrainingPerformanceWork experience placement- ...businesses learn from and optimize in‑person customer... ...deploy production‑grade ML systems with end‑to‑... ...preprocessing, model training, deployment, inference... ...processes for scalability and performance. Qualifications... ...professional experience in ML engineering. Strong programming...TrainingPerformanceFull time
$195k - $230k
...meaningful challenges at scale. Together, we... ...Learning Engineer to help evolve our large-scale recommendation... ...multi-objective optimization to balance... ...systems from offline training to online... ...drift, and system performance in production. AI... ...-scale data and ML systems (e.g., Spark...TrainingPerformanceLocal areaWork from home$219k - $351k
...communities. Job Title: Principal Engineer, AI System Architect... ...bandwidth and system-scale communication . By... ...improvements in performance, efficiency, and scalability... ...architecture for large-scale computing... ...DLRMs , and large-scale training and inference systems...TrainingPerformanceWork at officeFlexible hours- ...applications of AI & ML are bringing... ...applied science and engineering teams to deliver... ...scalable, high‑performance AI infrastructure... ...foundation model training, large language model inference... ...‑of‑the‑art LLM optimization techniques to... ...— of large‑scale production AI systems...TrainingPerformanceFull timeLocal area
$214k - $289.5k
...Machine Learning Engineer (MLE). Senior... ...value at Intuit scale. In this role,... ...define and evolve ML architecture,... ...teams for training, deployment, monitoring... ...Deliver within large-scale strategic... ...and system performance, ensuring continuous... ...performance optimization. Excellent...TrainingPerformance$19 - $65 per hour
...Responsibilities: Identify Training Bottlenecks: Profile and analyze Bird... ...: Design and implement high-performance custom compute kernels using CUDA,... ...training process. Leverage LLMs for Optimization: Explore and integrate Large Language Models (LLMs) to assist...TrainingPerformanceHourly payInternship- ...About The Role As Chief Engineer , you will oversee all... ...with asset performance. Responsibilities Operations... ...equipment issues. Manage and optimize work order queuing... ...development. Identify training opportunities and provide... ...expertise in managing large, complex facilities...TrainingPerformanceOngoing contractWork at officeLocal areaFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal ML Engineer - Large Scale Training Performance Optimization. Be the first to apply!
- senior principal engineer San Jose, CA
- general engineer San Jose, CA
- principal engineer San Jose, CA
- director of product engineering San Jose, CA
- data center chief engineer San Jose, CA
- chief engineer San Jose, CA
- chief design engineer San Jose, CA
- hotel chief engineer San Jose, CA
- senior civil engineer project manager San Jose, CA
- director software engineering San Jose, CA



