Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Principal ML GPU Architect: Scale Training

$205.9k - $407.5k

Adobe

At Adobe, we're driving our reinvention as an AI company and betting on Generative AI! Last year, we released many Generative AI capabilities under the Firefly umbrella - features used by more than 50% Photoshop users and generating 5B+ images, while doing AI responsibly and transparently!

We are looking to bring on a Senior Principal ML GPU Architect to lead the ML GPU optimization team in Adobe Firefly, reporting to the Head of AI/ML and Data Platforms as a member of staff. You will partner with the Director of ML Engineering who is responsible for our platform engineering resources to unlock step function changes in training and inference speed/scale for all our ML workloads.

This opportunity will not only enable you to make real world impact by optimizing ML workloads running on tens of thousands of GPUs, but also will enable you to have the opportunity to publish relevant work as either open-source or as technical publications in major conferences. The role involves hands on impact on all ML platforms powering inference, training, and data, as well as guiding the platform strategy towards higher scale and faster execution areas.

We also expect you to contribute to hiring critical talent, building and enhancing relationships with Adobe research and Adobe product teams, investing in major new initiatives in emerging technologies, and communicating goals and breakthroughs to senior leadership, to Adobe, and Adobe’s customers. The role requires experience guiding highly motivated world-class ML practitioners towards ambitious goals, generating original intellectual property, and creating real-world impact.

What you’ll do

  • Help drive ML Platform technical roadmap and Strategy.
  • Lead and mentor highly motivated ML GPU optimization engineers/scientists.
  • Write efficient forward and backward passes in CUDA/CuTe.
  • Write optimized custom layers inPytorch.
  • Optimize ML training and inference code for large, distributed training/inference with FP8.
  • Quality and performance analysis between data types such as BF16 and FP8 for large deep learning models.
  • Understand and optimize H100 GPUs.
  • Architect broader, end to end optimized training and inference code and schemes withPytorchforlarge, distributedmodels.
  • Write high quality, product level code that is easy tomaintainand test following standard methodologies.

What you'll need to succeed

  • Proficiencyin at least two of: Linux, Ansible, Docker, Kubernetes (7+yrs)
  • Expert in Python and C++
  • Expert in CUDA/CuTe, NCCL, OpenCL, Triton
  • Expert inPytorch
  • Experience with DDP, FSDP
  • A minimum of seven years of experience in distributed computing
  • A minimum of five of experience working with AWS or similar cloud infrastructure
  • Experience with HW resource management for ML training and/or deployment
  • S., M.S, or Ph.D. in Computer Science, ComputerEngineeringor a related area

At Adobe, you will be immersed in an exceptional work environment that is recognized around the world! You will also be surrounded by colleagues who are committed to helping each other grow through our unique Check-In approach where ongoing feedback flows freely. If you’re looking to make an impact, Adobe's the place for you. Discover what our employees are saying about their career experiences on the Adobe Life blog and explore the meaningful benefits we offer.

Adobe is an equal opportunity employer. We hire hard-working individuals, regardless of gender, race or color, ethnicity or national origin, age, disability, religion, sexual orientation, gender identity or expression, or veteran status. We know that when our employees feel appreciated and included, they can be more creative, innovative and successful. This is what it means to be Adobe For All. Learn more about our vision here.

We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us to request accommodation.

Our compensation reflects the cost of labor across several U.S. geographic markets, and we pay differently based on those defined markets. The U.S. pay range for this positionis $205,900 -- $407,500 annually. Paywithin this range varies by work locationand may also depend on job-related knowledge, skills,and experience. Your recruiter can share more about the specific salary range for the job location during the hiring process.

At Adobe, for sales roles starting salaries are expressed as total target compensation (TTC = base + commission), and short-term incentives are in the form of sales commission plans. Non-sales roles starting salaries are expressed as base salary and short-term incentives are in the form of the Annual Incentive Plan (AIP).

In addition, certain roles may be eligible for long-term incentives in the form of a new hire equity award.

Adobe is proud to be an Equal Employment Opportunity and affirmative action employer. We do not discriminate based on gender, race or color, ethnicity or national origin, age, disability, religion, sexual orientation, gender identity or expression, veteran status, or any other applicable characteristics protected by law. Learn more.

#J-18808-Ljbffr
Vacancy posted 2 hours ago
Similar jobs that could be interesting for youBased on the Senior Principal ML GPU Architect: Scale Training in San Jose, CA vacancy
  •  ...Micro Devices is looking for a Principal Engineer in Santa Clara, CA to...  ...infrastructure development, define GPU architecture specifications, and drive performance gains in ML systems. The role involves...  ...programming, and optimizing large-scale ML systems. A Bachelor's, MS or... 
    Principal
    Training

    Advanced Micro Devices

    Santa Clara, CA
    2 hours ago
  •  ...Senior Director, Design Engineering (Req ID: 134544) Hiring...  ...position is for a Senior Principal Engineer, AI/ML System Architect. As system architect, one...  ...design including AI training and inference workloads and...  ..., Intel, or other modern GPU accelerators and support... 
    Principal
    Senior
    Training
    Local area
    Remote work

    Celestica

    San Jose, CA
    2 hours ago
  •  ...ROLE: We are seeking a Robotics AI Architect to define and scale next‑generation Physical AI systems,...  ...compute‑software co‑design across CPU, GPU, and accelerators Act as...  ...Edge/accelerator subsystems Cloud (training, simulation, fleet learning) Provide... 
    Principal
    Senior
    Training

    Advanced Micro Devices

    San Jose, CA
    4 days ago
  •  ...career. THE ROLE:As a Principal Engineer, you...  ...by defining GPU architecture specifications...  ...massive model training at scale. Your expertise...  ...for distributed ML systems, you will...  ...EXPERIENCE:Extensive and Senior experience...  ...track record architecting distributed training... 
    Principal
    Training
    Remote work

    AMD

    Santa Clara, CA
    5 days ago
  • $272k - $431.25k

     ...interconnects.This Principal Architect role leads the...  ...systems communicate at scale—across GPUs, DPUs,...  ...systems—GPU-to-GPU, GPU-to-storage...  ...bodies, and mentoring senior engineers across the...  ....Understanding of ML systems concepts—...  ..., or distributed training and inference patterns... 
    Principal
    Training
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  • $162.7k - $284.7k

     ...Preferred QualificationsThe Principal HPC Architect designs, builds,...  ...optimizes, and supports large scale compute environments...  ...computing, AI/ML workloads, simulation,...  ...applications for CPU, GPU, memory, and I/O performance...  ...education level or training. We are committed to... 
    Principal
    Training
    Minimum wage
    Full time
    Work experience placement
    Flexible hours

    KLA-Tencor

    Milpitas, CA
    4 days ago
  • $208k - $327.75k

     ....We are looking for a Senior AI Architect to help define the next...  ...SoCs, including GPU, CPU, DLA, memory hierarchy...  ...years of experience in AI/ML systems, deep learning...  ...and large-scale model systemsExperience...  ...understanding of distributed training systems, scaling laws,... 
    Senior
    Training
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...THE ROLE: Drive the performance of post‑training workloads on AMD Instinct™ GPUs. You’ll work...  ..., and optimizer steps.Optimize multi‑GPU/multi‑node training and communication patterns...  ...with SFT. LoRA and RL‑based training at scale.Strong PyTorch experience (torch.... 
    Principal
    Senior
    Training

    AMD

    San Jose, CA
    3 days ago
  • $198.2k - $297.2k

     ...with us!Role and ResponsibilitiesAs a Senior Staff GPU Architect - Machine Learning, you will help lead...  ...development of innovative machine learning (ML) solutions for Samsung’s premium mobile...  ...to enable efficient execution of large-scale AI models. You will collaborate across... 
    Senior
    Hourly pay
    Full time
    Relocation

    Samsung Semiconductor

    San Jose, CA
    4 days ago
  • $231.1k - $358.2k

    DescriptionJob Title: Sr. Principal SoC ArchitectJob...  ...sensor, or modality. Edge ML applications that run completely...  ...for a Chief SoC Architect to help define next...  ...SiMa.ai is looking for a senior architect to lead its SoC...  ...target compensation, training, company needs, and current... 
    Principal
    Senior
    Training
    Full time
    Work at office

    SiMa Technologies

    San Jose, CA
    21 hours ago
  •  ...ROLE:We are looking for a Principal Machine Learning...  ...challenge of distributed training of large models on a large...  ...generative AI at scale.THE PERSON:The ideal candidate...  ...:Experience with ML/DL frameworks such as PyTorch...  ...a plus.Experience with GPU kernel optimization is... 
    Principal
    Training

    AMD

    San Jose, CA
    3 days ago
  •  ...Oracle is seeking a seasoned software/solutions architect to mentor teams and lead the architecture of highly scalable distributed systems...  ...reliability, security, and engineering excellence across large-scale data plane platforms, with opportunities for impact and... 
    Principal
    Senior

    Oracle

    Santa Clara, CA
    2 hours ago
  • $206.4k - $379.1k

     ...drives creativity at scale in design, imaging, motion...  ....We're looking for a Principal Architect to build and implement...  ..., merging strong ML skills with proficiency...  ...infrastructure to support model training, fine-tuning,...  ...intelligent systems.Mentor senior engineers and... 
    Principal
    Training
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours

    Adobe Systems

    San Jose, CA
    1 day ago
  • $184k - $287.5k

    We are now looking for a Senior GPU & Deep Learning Architect!The NVIDIA GPU Architecture group is looking for world class architects and software developers...  ..., especially for deep learning workloads, both training and inference, and maintain our leadership by developing... 
    Senior
    Training
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • NVIDIA Corporation is seeking a Senior Research Engineer for the Autonomous...  ...Clara, CA. You will drive large-scale training pipelines for multimodal AV models, optimize GPU usage, and build simulation...  ...Ideal candidates have 10+ years in ML/AI infrastructure, proficiency... 
    Senior
    Training

    NVIDIA Corporation

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...generation of scientific machine learning (ML) frameworks. Starting with digital...  ...performant features for large scale, CUDA-backed ML training frameworks, using low level acceleration...  ...scaling strategies such as kernel design, GPU porting, data structure innovations, distributed... 
    Senior
    Training
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...Accellor is seeking a Technical Architect — AI Systems, Inference & Platform Internals to design, scale, and optimize internal AI...  ...inference runtime, model serving, GPU infrastructure, and distributed...  ...candidate has 10–12 years of software/ML infra experience, deep... 
    Senior

    Accellor

    Mountain View, CA
    2 hours ago
  • $184k - $287.5k

     ...seeking an expert Solutions Architect to assist customers in building AI/ML and HPC software solutions at scale. As a member of our Solutions...  ...tasks like large scale LLM training and inference.Conducting regular...  ....Hands-on experience with GPU systems in general including... 
    Senior
    Training
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $256k - $414k

     ...interactive entertainment at scale.We are looking for a Senior Manager to lead the design...  ...networking for GPU-based cloud infrastructure...  ...cloud gaming workloads, AI/ML training, and inference platforms by...  ...specialized team of network architects focused on high-performance... 
    Senior
    Training
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...drive technical strategy for ML model development, training pipelines, and inference...  ...modeling. Serve as the senior technical voice in design reviews...  .... Mentor and develop principal and senior ML engineers...  ...technical leadership on large-scale or novel ML systems. ~ Deep... 
    Principal
    Senior
    Training
    Full time
    Temporary work
    Flexible hours

    SambaNova Systems

    San Jose, CA
    2 days ago
  • $164k - $264.5k

     ...looking for an experienced PCIe Subsystem Architect to help define and deliver high-...  ...card and server level through full-rack scale-up and scale-out deployments. Minimum...  ...mode, PAM4 signaling implications, link training, latency, error handling, and system-level... 
    Principal
    Senior
    Training
    Work experience placement
    Work from home

    Qualcomm

    Santa Clara, CA
    3 days ago
  •  ...delivers the automation, GPU orchestration, and...  ...: whether we can architect the right cluster,...  ...-sales bar as we scale headcount and deal...  ...genuinely stood up training and inference workloads...  ...with a customer's ML infra lead, and...  ...Indicative OTE: senior people-leader band... 
    Senior
    Training
    Full time
    Remote work

    Mirantis

    San Jose, CA
    5 days ago
  • $224k - $356.5k

     ...cases.Analyze and debug performance scaling bottlenecks on multi-core and multi-socket CPU and CPU/GPU systems.Work with CPU and interconnect architects to improve future CPU and system designs...  ...stack, enabling faster AI model training, agentic use-cases, efficient data processing... 
    Senior
    Training
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $220k - $300k

     ...businesses and operations at scale. SambaNova Suite™ is the first...  .... Overview As a Senior Principal Machine Learning Engineer, you...  ...model architectures, improving training and inference efficiency, and...  ...engineer will also act as the ML expert, guiding the... 
    Principal
    Senior
    Training
    Full time
    Temporary work
    Local area
    Flexible hours

    SambaNova

    San Jose, CA
    3 days ago
  •  .... THE TEAMAMD's Data Center GPU organization is transforming...  ...seeking a highly accomplished Principal Modeling Architect to join the Product...  ...deep analysis of emerging AI/ML, HPC, and data analytics workloads...  ...neural networks), datatypes, and scaling methodologies to anticipate... 
    Principal
    Remote work

    AMD

    San Jose, CA
    3 days ago
  • $224k - $356.5k

     ...computing. An era in which our GPU acts as the brains of...  ...As an AI Storage Platform Architect at NVIDIA, this position will...  ...NVIDIA Dynamo), large-scale foundation model training, and agentic AI pipelines...  ...storage infrastructure as a Principal Architect, Solutions Architect... 
    Senior
    Training
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving...  ...NVIDIA has openings for a Deep Learning Communication Architect. We scale the DNN models and training/inference frameworks to systems with hundreds of... 
    Senior
    Training
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...next era of computing. An era in which our GPU acts as the brains of computers, robots,...  ....We are seeking a world-class computer architect to contribute to the development of future...  ...network on-chip design.Experience in large-scale SW development projects and strong... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...itself over two decades. Our invention of the GPU in 1999 sparked the growth of the PC...  ...are looking for an outstanding hands-on architect/engineer for a Senior HPC architect role to support deployment and bringup of large-scale GPU compute clusters. Be a key player to... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...Advanced Micro Devices, Inc. is seeking a Principal Modeling Architect in San Jose, CA, responsible for...  ...advanced workload modeling for AI/ML and HPC in data center environments...  ...expertise in analyzing large-scale workloads on GPU platforms. The position offers the... 
    Principal
    Remote work

    Advanced Micro Devices, Inc.

    San Jose, CA
    2 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Principal ML GPU Architect: Scale Training. Be the first to apply!