Senior Principal ML GPU Architect: Scale Training
$205.9k - $407.5kAdobe
At Adobe, we're driving our reinvention as an AI company and betting on Generative AI! Last year, we released many Generative AI capabilities under the Firefly umbrella - features used by more than 50% Photoshop users and generating 5B+ images, while doing AI responsibly and transparently!
We are looking to bring on a Senior Principal ML GPU Architect to lead the ML GPU optimization team in Adobe Firefly, reporting to the Head of AI/ML and Data Platforms as a member of staff. You will partner with the Director of ML Engineering who is responsible for our platform engineering resources to unlock step function changes in training and inference speed/scale for all our ML workloads.
This opportunity will not only enable you to make real world impact by optimizing ML workloads running on tens of thousands of GPUs, but also will enable you to have the opportunity to publish relevant work as either open-source or as technical publications in major conferences. The role involves hands on impact on all ML platforms powering inference, training, and data, as well as guiding the platform strategy towards higher scale and faster execution areas.
We also expect you to contribute to hiring critical talent, building and enhancing relationships with Adobe research and Adobe product teams, investing in major new initiatives in emerging technologies, and communicating goals and breakthroughs to senior leadership, to Adobe, and Adobe’s customers. The role requires experience guiding highly motivated world-class ML practitioners towards ambitious goals, generating original intellectual property, and creating real-world impact.
What you’ll do
- Help drive ML Platform technical roadmap and Strategy.
- Lead and mentor highly motivated ML GPU optimization engineers/scientists.
- Write efficient forward and backward passes in CUDA/CuTe.
- Write optimized custom layers inPytorch.
- Optimize ML training and inference code for large, distributed training/inference with FP8.
- Quality and performance analysis between data types such as BF16 and FP8 for large deep learning models.
- Understand and optimize H100 GPUs.
- Architect broader, end to end optimized training and inference code and schemes withPytorchforlarge, distributedmodels.
- Write high quality, product level code that is easy tomaintainand test following standard methodologies.
What you'll need to succeed
- Proficiencyin at least two of: Linux, Ansible, Docker, Kubernetes (7+yrs)
- Expert in Python and C++
- Expert in CUDA/CuTe, NCCL, OpenCL, Triton
- Expert inPytorch
- Experience with DDP, FSDP
- A minimum of seven years of experience in distributed computing
- A minimum of five of experience working with AWS or similar cloud infrastructure
- Experience with HW resource management for ML training and/or deployment
- S., M.S, or Ph.D. in Computer Science, ComputerEngineeringor a related area
At Adobe, you will be immersed in an exceptional work environment that is recognized around the world! You will also be surrounded by colleagues who are committed to helping each other grow through our unique Check-In approach where ongoing feedback flows freely. If you’re looking to make an impact, Adobe's the place for you. Discover what our employees are saying about their career experiences on the Adobe Life blog and explore the meaningful benefits we offer.
Adobe is an equal opportunity employer. We hire hard-working individuals, regardless of gender, race or color, ethnicity or national origin, age, disability, religion, sexual orientation, gender identity or expression, or veteran status. We know that when our employees feel appreciated and included, they can be more creative, innovative and successful. This is what it means to be Adobe For All. Learn more about our vision here.
We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us to request accommodation.
Our compensation reflects the cost of labor across several U.S. geographic markets, and we pay differently based on those defined markets. The U.S. pay range for this positionis $205,900 -- $407,500 annually. Paywithin this range varies by work locationand may also depend on job-related knowledge, skills,and experience. Your recruiter can share more about the specific salary range for the job location during the hiring process.
At Adobe, for sales roles starting salaries are expressed as total target compensation (TTC = base + commission), and short-term incentives are in the form of sales commission plans. Non-sales roles starting salaries are expressed as base salary and short-term incentives are in the form of the Annual Incentive Plan (AIP).
In addition, certain roles may be eligible for long-term incentives in the form of a new hire equity award.
Adobe is proud to be an Equal Employment Opportunity and affirmative action employer. We do not discriminate based on gender, race or color, ethnicity or national origin, age, disability, religion, sexual orientation, gender identity or expression, veteran status, or any other applicable characteristics protected by law. Learn more.
#J-18808-Ljbffr- ...Micro Devices is looking for a Principal Engineer in Santa Clara, CA to... ...infrastructure development, define GPU architecture specifications, and drive performance gains in ML systems. The role involves... ...programming, and optimizing large-scale ML systems. A Bachelor's, MS or...PrincipalTraining
- ...Senior Director, Design Engineering (Req ID: 134544) Hiring... ...position is for a Senior Principal Engineer, AI/ML System Architect. As system architect, one... ...design including AI training and inference workloads and... ..., Intel, or other modern GPU accelerators and support...PrincipalSeniorTrainingLocal areaRemote work
- ...ROLE: We are seeking a Robotics AI Architect to define and scale next‑generation Physical AI systems,... ...compute‑software co‑design across CPU, GPU, and accelerators Act as... ...Edge/accelerator subsystems Cloud (training, simulation, fleet learning) Provide...PrincipalSeniorTraining
- ...career. THE ROLE:As a Principal Engineer, you... ...by defining GPU architecture specifications... ...massive model training at scale. Your expertise... ...for distributed ML systems, you will... ...EXPERIENCE:Extensive and Senior experience... ...track record architecting distributed training...PrincipalTrainingRemote work
$272k - $431.25k
...interconnects.This Principal Architect role leads the... ...systems communicate at scale—across GPUs, DPUs,... ...systems—GPU-to-GPU, GPU-to-storage... ...bodies, and mentoring senior engineers across the... ....Understanding of ML systems concepts—... ..., or distributed training and inference patterns...PrincipalTrainingFull timeRemote work$162.7k - $284.7k
...Preferred QualificationsThe Principal HPC Architect designs, builds,... ...optimizes, and supports large scale compute environments... ...computing, AI/ML workloads, simulation,... ...applications for CPU, GPU, memory, and I/O performance... ...education level or training. We are committed to...PrincipalTrainingMinimum wageFull timeWork experience placementFlexible hours$208k - $327.75k
....We are looking for a Senior AI Architect to help define the next... ...SoCs, including GPU, CPU, DLA, memory hierarchy... ...years of experience in AI/ML systems, deep learning... ...and large-scale model systemsExperience... ...understanding of distributed training systems, scaling laws,...SeniorTrainingFull timeWorldwide- ...THE ROLE: Drive the performance of post‑training workloads on AMD Instinct™ GPUs. You’ll work... ..., and optimizer steps.Optimize multi‑GPU/multi‑node training and communication patterns... ...with SFT. LoRA and RL‑based training at scale.Strong PyTorch experience (torch....PrincipalSeniorTraining
$198.2k - $297.2k
...with us!Role and ResponsibilitiesAs a Senior Staff GPU Architect - Machine Learning, you will help lead... ...development of innovative machine learning (ML) solutions for Samsung’s premium mobile... ...to enable efficient execution of large-scale AI models. You will collaborate across...SeniorHourly payFull timeRelocation$231.1k - $358.2k
DescriptionJob Title: Sr. Principal SoC ArchitectJob... ...sensor, or modality. Edge ML applications that run completely... ...for a Chief SoC Architect to help define next... ...SiMa.ai is looking for a senior architect to lead its SoC... ...target compensation, training, company needs, and current...PrincipalSeniorTrainingFull timeWork at office- ...ROLE:We are looking for a Principal Machine Learning... ...challenge of distributed training of large models on a large... ...generative AI at scale.THE PERSON:The ideal candidate... ...:Experience with ML/DL frameworks such as PyTorch... ...a plus.Experience with GPU kernel optimization is...PrincipalTraining
- ...Oracle is seeking a seasoned software/solutions architect to mentor teams and lead the architecture of highly scalable distributed systems... ...reliability, security, and engineering excellence across large-scale data plane platforms, with opportunities for impact and...PrincipalSenior
$206.4k - $379.1k
...drives creativity at scale in design, imaging, motion... ....We're looking for a Principal Architect to build and implement... ..., merging strong ML skills with proficiency... ...infrastructure to support model training, fine-tuning,... ...intelligent systems.Mentor senior engineers and...PrincipalTrainingFull timeTemporary workLocal areaWorldwideFlexible hours$184k - $287.5k
We are now looking for a Senior GPU & Deep Learning Architect!The NVIDIA GPU Architecture group is looking for world class architects and software developers... ..., especially for deep learning workloads, both training and inference, and maintain our leadership by developing...SeniorTrainingFull time- NVIDIA Corporation is seeking a Senior Research Engineer for the Autonomous... ...Clara, CA. You will drive large-scale training pipelines for multimodal AV models, optimize GPU usage, and build simulation... ...Ideal candidates have 10+ years in ML/AI infrastructure, proficiency...SeniorTraining
$184k - $287.5k
...generation of scientific machine learning (ML) frameworks. Starting with digital... ...performant features for large scale, CUDA-backed ML training frameworks, using low level acceleration... ...scaling strategies such as kernel design, GPU porting, data structure innovations, distributed...SeniorTrainingFull time- ...Accellor is seeking a Technical Architect — AI Systems, Inference & Platform Internals to design, scale, and optimize internal AI... ...inference runtime, model serving, GPU infrastructure, and distributed... ...candidate has 10–12 years of software/ML infra experience, deep...Senior
$184k - $287.5k
...seeking an expert Solutions Architect to assist customers in building AI/ML and HPC software solutions at scale. As a member of our Solutions... ...tasks like large scale LLM training and inference.Conducting regular... ....Hands-on experience with GPU systems in general including...SeniorTrainingFull time$256k - $414k
...interactive entertainment at scale.We are looking for a Senior Manager to lead the design... ...networking for GPU-based cloud infrastructure... ...cloud gaming workloads, AI/ML training, and inference platforms by... ...specialized team of network architects focused on high-performance...SeniorTrainingFull timeLocal area- ...drive technical strategy for ML model development, training pipelines, and inference... ...modeling. Serve as the senior technical voice in design reviews... .... Mentor and develop principal and senior ML engineers... ...technical leadership on large-scale or novel ML systems. ~ Deep...PrincipalSeniorTrainingFull timeTemporary workFlexible hours
$164k - $264.5k
...looking for an experienced PCIe Subsystem Architect to help define and deliver high-... ...card and server level through full-rack scale-up and scale-out deployments. Minimum... ...mode, PAM4 signaling implications, link training, latency, error handling, and system-level...PrincipalSeniorTrainingWork experience placementWork from home- ...delivers the automation, GPU orchestration, and... ...: whether we can architect the right cluster,... ...-sales bar as we scale headcount and deal... ...genuinely stood up training and inference workloads... ...with a customer's ML infra lead, and... ...Indicative OTE: senior people-leader band...SeniorTrainingFull timeRemote work
$224k - $356.5k
...cases.Analyze and debug performance scaling bottlenecks on multi-core and multi-socket CPU and CPU/GPU systems.Work with CPU and interconnect architects to improve future CPU and system designs... ...stack, enabling faster AI model training, agentic use-cases, efficient data processing...SeniorTrainingFull time$220k - $300k
...businesses and operations at scale. SambaNova Suite™ is the first... .... Overview As a Senior Principal Machine Learning Engineer, you... ...model architectures, improving training and inference efficiency, and... ...engineer will also act as the ML expert, guiding the...PrincipalSeniorTrainingFull timeTemporary workLocal areaFlexible hours- .... THE TEAMAMD's Data Center GPU organization is transforming... ...seeking a highly accomplished Principal Modeling Architect to join the Product... ...deep analysis of emerging AI/ML, HPC, and data analytics workloads... ...neural networks), datatypes, and scaling methodologies to anticipate...PrincipalRemote work
$224k - $356.5k
...computing. An era in which our GPU acts as the brains of... ...As an AI Storage Platform Architect at NVIDIA, this position will... ...NVIDIA Dynamo), large-scale foundation model training, and agentic AI pipelines... ...storage infrastructure as a Principal Architect, Solutions Architect...SeniorTrainingFull time$184k - $287.5k
...define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving... ...NVIDIA has openings for a Deep Learning Communication Architect. We scale the DNN models and training/inference frameworks to systems with hundreds of...SeniorTrainingFull timeWork experience placement$184k - $287.5k
...next era of computing. An era in which our GPU acts as the brains of computers, robots,... ....We are seeking a world-class computer architect to contribute to the development of future... ...network on-chip design.Experience in large-scale SW development projects and strong...SeniorFull time$184k - $287.5k
...itself over two decades. Our invention of the GPU in 1999 sparked the growth of the PC... ...are looking for an outstanding hands-on architect/engineer for a Senior HPC architect role to support deployment and bringup of large-scale GPU compute clusters. Be a key player to...SeniorFull timeRemote work- ...Advanced Micro Devices, Inc. is seeking a Principal Modeling Architect in San Jose, CA, responsible for... ...advanced workload modeling for AI/ML and HPC in data center environments... ...expertise in analyzing large-scale workloads on GPU platforms. The position offers the...PrincipalRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Principal ML GPU Architect: Scale Training. Be the first to apply!
- senior lead project manager San Jose, CA
- senior robotics software engineer San Jose, CA
- senior devops engineer remote San Jose, CA
- senior sas administrator San Jose, CA
- senior IT manager San Jose, CA
- sr project manager San Jose, CA
- senior windows systems engineer San Jose, CA
- consultant senior consultant San Jose, CA
- senior manager data science San Jose, CA
- senior ui ux designer San Jose, CA



