Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Principal ML GPU Architect: Scale Training

$205.9k - $407.5k

Adobe

At Adobe, we're driving our reinvention as an AI company and betting on Generative AI! Last year, we released many Generative AI capabilities under the Firefly umbrella - features used by more than 50% Photoshop users and generating 5B+ images, while doing AI responsibly and transparently!

We are looking to bring on a Senior Principal ML GPU Architect to lead the ML GPU optimization team in Adobe Firefly, reporting to the Head of AI/ML and Data Platforms as a member of staff. You will partner with the Director of ML Engineering who is responsible for our platform engineering resources to unlock step function changes in training and inference speed/scale for all our ML workloads.

This opportunity will not only enable you to make real world impact by optimizing ML workloads running on tens of thousands of GPUs, but also will enable you to have the opportunity to publish relevant work as either open-source or as technical publications in major conferences. The role involves hands on impact on all ML platforms powering inference, training, and data, as well as guiding the platform strategy towards higher scale and faster execution areas.

We also expect you to contribute to hiring critical talent, building and enhancing relationships with Adobe research and Adobe product teams, investing in major new initiatives in emerging technologies, and communicating goals and breakthroughs to senior leadership, to Adobe, and Adobe’s customers. The role requires experience guiding highly motivated world-class ML practitioners towards ambitious goals, generating original intellectual property, and creating real-world impact.

What you’ll do

  • Help drive ML Platform technical roadmap and Strategy.
  • Lead and mentor highly motivated ML GPU optimization engineers/scientists.
  • Write efficient forward and backward passes in CUDA/CuTe.
  • Write optimized custom layers inPytorch.
  • Optimize ML training and inference code for large, distributed training/inference with FP8.
  • Quality and performance analysis between data types such as BF16 and FP8 for large deep learning models.
  • Understand and optimize H100 GPUs.
  • Architect broader, end to end optimized training and inference code and schemes withPytorchforlarge, distributedmodels.
  • Write high quality, product level code that is easy tomaintainand test following standard methodologies.

What you'll need to succeed

  • Proficiencyin at least two of: Linux, Ansible, Docker, Kubernetes (7+yrs)
  • Expert in Python and C++
  • Expert in CUDA/CuTe, NCCL, OpenCL, Triton
  • Expert inPytorch
  • Experience with DDP, FSDP
  • A minimum of seven years of experience in distributed computing
  • A minimum of five of experience working with AWS or similar cloud infrastructure
  • Experience with HW resource management for ML training and/or deployment
  • S., M.S, or Ph.D. in Computer Science, ComputerEngineeringor a related area

At Adobe, you will be immersed in an exceptional work environment that is recognized around the world! You will also be surrounded by colleagues who are committed to helping each other grow through our unique Check-In approach where ongoing feedback flows freely. If you’re looking to make an impact, Adobe's the place for you. Discover what our employees are saying about their career experiences on the Adobe Life blog and explore the meaningful benefits we offer.

Adobe is an equal opportunity employer. We hire hard-working individuals, regardless of gender, race or color, ethnicity or national origin, age, disability, religion, sexual orientation, gender identity or expression, or veteran status. We know that when our employees feel appreciated and included, they can be more creative, innovative and successful. This is what it means to be Adobe For All. Learn more about our vision here.

We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us to request accommodation.

Our compensation reflects the cost of labor across several U.S. geographic markets, and we pay differently based on those defined markets. The U.S. pay range for this positionis $205,900 -- $407,500 annually. Paywithin this range varies by work locationand may also depend on job-related knowledge, skills,and experience. Your recruiter can share more about the specific salary range for the job location during the hiring process.

At Adobe, for sales roles starting salaries are expressed as total target compensation (TTC = base + commission), and short-term incentives are in the form of sales commission plans. Non-sales roles starting salaries are expressed as base salary and short-term incentives are in the form of the Annual Incentive Plan (AIP).

In addition, certain roles may be eligible for long-term incentives in the form of a new hire equity award.

Adobe is proud to be an Equal Employment Opportunity and affirmative action employer. We do not discriminate based on gender, race or color, ethnicity or national origin, age, disability, religion, sexual orientation, gender identity or expression, veteran status, or any other applicable characteristics protected by law. Learn more.

#J-18808-Ljbffr
Vacancy posted 6 hours ago
Similar jobs that could be interesting for youBased on the Senior Principal ML GPU Architect: Scale Training in San Francisco, CA vacancy
  • STN Inc in San Francisco is seeking an experienced AI Infrastructure Engineer to design, deploy, and manage large-scale GPU clusters for AI training and inference workloads. You will optimize GPU utilization, tune NCCL, CUDA, UCX, and Slurm, and work across storage, networking... 
    Senior
    Training

    STN Inc

    San Francisco, CA
    6 days ago
  • Autodesk, Inc. in San Francisco seeks a Senior Principal AI/ML Developer to shape data-driven personalization and analytic initiatives across the...  ...collaborate with product, engineering, and marketing teams to deploy robust ML solutions at scale. #J-18808-Ljbffr Autodesk
    Principal
    Senior

    Autodesk

    San Francisco, CA
    2 days ago
  • $206.4k - $379.1k

     ...drives creativity at scale in design, imaging, motion...  ....We're looking for a Principal Architect to build and implement...  ..., merging strong ML skills with proficiency...  ...infrastructure to support model training, fine-tuning,...  ...intelligent systems.Mentor senior engineers and... 
    Principal
    Training
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours

    Adobe Systems

    San Francisco, CA
    2 days ago
  • $175k - $250k

     ...Cloud ML Software Architect – Cutting-Edge Biotech + AI Startup...  ...$40M to date and are scaling rapidly. Compensation...  ...infrastructure (AWS core, GPU compute, Docker &...  ...processing and ML model training. Competitive...  ...****@*****.*** . Seniority level: Mid‑Senior level... 
    Principal
    Training
    Full time
    H1b
    Work at office
    Visa sponsorship
    3 days per week

    Strativ Group

    South San Francisco, CA
    2 days ago
  •  ...engineering team — a small, senior group focused on...  ...Science, AI/ML, or related field. PhD...  ...least 2 years in a principal engineer or lead architect role. ~ Demonstrated...  ...expertise across: LLM training and fine‑tuning,...  ...construction, large‑scale data modeling. ~ Comfort... 
    Principal
    Training
    Work at office
    Visa sponsorship
    Flexible hours
    3 days per week

    HopHR

    San Francisco, CA
    5 hours ago
  •  ...San Francisco is looking for a Senior Software Engineer to build...  ...scalable infrastructure for large‑scale training and fine-tuning of foundation...  ...training systems and optimize GPU utilization while collaborating...  ...over 5 years of experience in ML infrastructure and a strong... 
    Senior
    Training

    Baseten

    San Francisco, CA
    5 hours ago
  • A cutting-edge AI technology company based in San Francisco is seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate... 
    Senior
    Training

    Reflection AI

    San Francisco, CA
    4 days ago
  • $160k - $225k

     ...Cacheflow is seeking a Senior Software Engineer for AI Runtime at Databricks, located in San Francisco. You will be instrumental in building and scaling systems for large-scale GPU training, ensuring high throughput and resilience in training across expansive fleets of... 
    Senior
    Training

    Cacheflow

    San Francisco, CA
    5 hours ago
  • Principal or Senior Principal, Anthropic AI SolutionsAI Systems...  ...sales teams, and solution architects that leads to pre-...  ...enablement programs: designing training curricula, running...  ...compliant adoption at scale. * Enable internal and...  ...and delivering AI/ML or large language model... 
    Principal
    Senior
    Training
    Temporary work
    Work at office
    Local area

    Slalom

    San Francisco, CA
    5 days ago
  •  ..., deploy, and maintain large distributed ML training and inference clusters Develop efficient...  ...end-to-end pipelines to manage petabyte-scale datasets and model training throughout the...  ...scales Analyze, profile and debug low-level GPU operations to optimize performance Stay... 
    Senior
    Training

    Kindredventures

    San Francisco, CA
    6 days ago
  • $179k - $218k

     ...who believe in the scale of our ambition and...  ...bridged.We are seeking a Senior Staff Data Center...  ...Engineer, GPU Hardware Architecture...  ...Telemetry: Leverage AI/ML methodologies to...  ...they impact customer training runs.Technical Sparing Architecture: Architect the site-level sparing... 
    Senior
    Training
    Temporary work

    Crusoe

    San Francisco, CA
    4 days ago
  • $250k

     ...opportunities? Join a rapidly scaling AI cloud infrastructure...  ...a next-generation GPU platform designed for AI training, experimentation, and inference...  ...company is looking for a Senior / Staff Site Reliability Engineer...  ...closely with platform, ML, and infrastructure teams... 
    Senior
    Training
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  • $280k - $350k

    Principal Research Scientist - ScalingP-1227About Databricks AIAt Databricks...  ...ranging from post-training open source LLMs to developing...  ...be available to all.About the Scaling Research TeamThe Databricks AI...  ...approaches.Optimize end‑to‑end ML systems for distributed training... 
    Principal
    Training
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    4 days ago
  • $148.5k - $313.7k

     ...and overall product quality at enterprise scale Contribute to strong engineering culture...  ...Hands‑on experience applying AI/ML in product (LLMs, embeddings, or similar)...  ...assignment, compensation, promotion, benefits, training, assessment of job performance, discipline... 
    Principal
    Senior
    Training

    salesforce.com, inc.

    San Francisco, CA
    5 hours ago
  • $288k

     ...domains. We are seeking a Senior or Principal Scientist to set the...  ...formulation and architecture through training at scale, evaluation, and integration...  ...and related structure-aware ML methods Set the...  ...evaluation at scale across large GPU clusters Shape the end-... 
    Principal
    Senior
    Training
    Full time
    Work at office
    Local area
    Flexible hours

    Lila Sciences

    San Francisco, CA
    4 days ago
  • $284.32k - $355.4k

     ...See yourself at Twilio Join the team as Twilio's next Senior Principal Field Architect - AI Agents About the job The Senior Principal...  ..., and cross-functional initiatives at an enterprise scale with significant business impact. ~ Technical & Business... 
    Principal
    Senior
    Local area
    Remote work
    Worldwide
    Flexible hours

    Twilio

    San Francisco, CA
    5 days ago
  • $264.1k - $369.74k

     ...evolve rapidly as the constellation scales from first deployment through full operational...  ...for critical operations. As the Senior Principal Architect for Systems & Mission Operations...  ...hazardous materials transportation/shipping training. Required for certain Job Profiles:... 
    Principal
    Senior
    Training
    Permanent employment
    Temporary work
    Work at office
    Local area
    Worldwide
    Relocation

    Blue Origin

    San Francisco, CA
    4 days ago
  • $167.4k - $310.8k

     ...edge machine learning (ML) techniques. We are seeking a Senior or Principal Machine Learning Scientist...  ..., designing and scaling large machine learning...  ...for model architectures, training strategies, and evaluation...  ...Systems & Engineering: Architect and improve large-scale... 
    Principal
    Senior
    Training
    Full time
    Local area
    Worldwide
    Relocation package

    Genentech

    South San Francisco, CA
    2 days ago
  • Ginas Tech Jobs is seeking a Principal Machine Learning Engineer to set the technical standard for ML systems across training, inference, evaluation and deployment. This...  ...pipelines. You will own large‑scale ML systems, optimize GPU memory and latency, collaborate with... 
    Principal
    Training
    Remote job

    Ginas Tech Jobs

    San Francisco, CA
    3 days ago
  • Blue Origin seeks a Senior Principal Architect for Systems & Mission Operations Software to own end-to-end software architecture across flight and...  ...environment. The role requires deep expertise in large-scale distributed systems, C/C++, Rust, and strong leadership to... 
    Principal
    Senior

    Blue Origin

    San Francisco, CA
    5 days ago
  • $216.2k - $270.25k

    Scale GP (Scale Generative AI Platform) is an enterprise-grade Generative...  ...more. We are seeking a strong Senior Full-Stack Engineer to help us...  ...systems, data pipelines, and ML/LLM components.Integrate with...  ..., and relevant education or training. Scale employees in eligible... 
    Senior
    Training
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  •  ...Palo Alto Networks, Inc. is seeking a Senior Principal Backend Engineer in the Cortex group to lead the development and scaling of the backend for Cortex XSOAR, XDR, and XSIAM. You will help design robust data pipelines, APIs, and services, partnering with product and... 
    Principal
    Senior

    Jobleads-US

    San Francisco, CA
    6 days ago
  •  ...building production-grade ML infrastructure used by enterprise...  .... They are looking for a Senior AI/ML Engineer to own model training pipelines, evaluation...  ...and inference serving at scale. Full-time, on-site in San...  ...with distributed training, GPU optimization, or inference... 
    Senior
    Training
    Full time

    Clera

    San Francisco, CA
    1 day ago
  •  ...Software Solutions is hiring a Senior Data Engineer (Apache...  ...the design of large-scale distributed data processing...  .... Responsibilities Architect and optimize large-scale...  ...-tolerance Partner with ML engineers to deliver feature stores and training data sets at scale Drive... 
    Senior
    Training
    Flexible hours

    Appit LLC

    San Francisco, CA
    6 days ago
  • $252k - $374k

     ...Science AI (LSAI), ML engineers build and...  ...We are seeking a Principal ML Engineer to design, build, and scale the ML infrastructure...  ...systems end to end, from training pipelines and...  ...distributed training across GPU clusters Own...  ...workflows Architect ML infrastructure that... 
    Principal
    Training
    Full time
    Work at office
    Local area
    Flexible hours

    Lila Sciences

    San Francisco, CA
    1 day ago
  • $260k - $340k

     ...believe in the scale of our ambition...  ...This Role:As the Principal Systems Software...  ...push massive-scale training workloads to the...  ...(BMaaS): Architect systems that deliver raw GPU throughput via zero...  ...alongside Staff and Senior engineers to...  ...HPC projects.AI/ML Workload Expertise... 
    Principal
    Training
    Full time
    Temporary work

    Crusoe

    San Francisco, CA
    4 days ago
  • $154.38k - $193.13k

    Job DescriptionAI Architect, AI & AutomationAbout the Role The applicant...  ...and implementing AI/ML solutions across multiple domains...  ...clients, including clients at senior levels. Anticipates and proactively...  ...outcomes, faster, smarter, and at scale.Infosys Consulting is helping... 
    Principal
    Full time
    Temporary work
    Work experience placement

    Infosys Technologies

    San Francisco, CA
    2 days ago
  • $197.3k - $313.7k

     ...stage startup backed by the global scale and trust of Salesforce.What You'...  ...Actually Be DoingDesign, implement, and train novel deep learning models on large-scale GPU clusters.Prototype new...  ...highly quantitative field with an AI/ML research focusYou possess experience... 
    Principal
    Training
    Full time
    Immediate start
    Remote work

    Salesforce

    San Francisco, CA
    5 days ago
  • $127.4k - $191.1k

     ...integrated design practice. Our architects, engineers, interior...  ...the World.We are looking for a Senior Architect who shares our interest...  ...complex tasks on multiple, varying scale, highly complex projects....  ..., fringe benefits, job training, terminations or any other condition... 
    Senior
    Training
    Full time
    Contract work
    Temporary work
    Part time
    For contractors
    For subcontractor
    Casual work
    Live in
    Work at office
    Local area
    Flexible hours

    Stantec

    San Francisco, CA
    1 day ago
  •  ...RoleWe are looking for a Principal AI Engineer to join our...  ...complex problems at scale.In this role, you will...  ...sense to know why your training run is slow.You design...  ...meaningful open-source ML contributions; a research...  ...Some exposure to multi-GPU training, even at lab scale... 
    Principal
    Training
    Full time
    Temporary work
    Internship
    Immediate start

    Innovaccer

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Principal ML GPU Architect: Scale Training. Be the first to apply!