Senior Principal ML GPU Architect: Scale Training
$205.9k - $407.5kAdobe
At Adobe, we're driving our reinvention as an AI company and betting on Generative AI! Last year, we released many Generative AI capabilities under the Firefly umbrella - features used by more than 50% Photoshop users and generating 5B+ images, while doing AI responsibly and transparently!
We are looking to bring on a Senior Principal ML GPU Architect to lead the ML GPU optimization team in Adobe Firefly, reporting to the Head of AI/ML and Data Platforms as a member of staff. You will partner with the Director of ML Engineering who is responsible for our platform engineering resources to unlock step function changes in training and inference speed/scale for all our ML workloads.
This opportunity will not only enable you to make real world impact by optimizing ML workloads running on tens of thousands of GPUs, but also will enable you to have the opportunity to publish relevant work as either open-source or as technical publications in major conferences. The role involves hands on impact on all ML platforms powering inference, training, and data, as well as guiding the platform strategy towards higher scale and faster execution areas.
We also expect you to contribute to hiring critical talent, building and enhancing relationships with Adobe research and Adobe product teams, investing in major new initiatives in emerging technologies, and communicating goals and breakthroughs to senior leadership, to Adobe, and Adobe’s customers. The role requires experience guiding highly motivated world-class ML practitioners towards ambitious goals, generating original intellectual property, and creating real-world impact.
What you’ll do
- Help drive ML Platform technical roadmap and Strategy.
- Lead and mentor highly motivated ML GPU optimization engineers/scientists.
- Write efficient forward and backward passes in CUDA/CuTe.
- Write optimized custom layers inPytorch.
- Optimize ML training and inference code for large, distributed training/inference with FP8.
- Quality and performance analysis between data types such as BF16 and FP8 for large deep learning models.
- Understand and optimize H100 GPUs.
- Architect broader, end to end optimized training and inference code and schemes withPytorchforlarge, distributedmodels.
- Write high quality, product level code that is easy tomaintainand test following standard methodologies.
What you'll need to succeed
- Proficiencyin at least two of: Linux, Ansible, Docker, Kubernetes (7+yrs)
- Expert in Python and C++
- Expert in CUDA/CuTe, NCCL, OpenCL, Triton
- Expert inPytorch
- Experience with DDP, FSDP
- A minimum of seven years of experience in distributed computing
- A minimum of five of experience working with AWS or similar cloud infrastructure
- Experience with HW resource management for ML training and/or deployment
- S., M.S, or Ph.D. in Computer Science, ComputerEngineeringor a related area
At Adobe, you will be immersed in an exceptional work environment that is recognized around the world! You will also be surrounded by colleagues who are committed to helping each other grow through our unique Check-In approach where ongoing feedback flows freely. If you’re looking to make an impact, Adobe's the place for you. Discover what our employees are saying about their career experiences on the Adobe Life blog and explore the meaningful benefits we offer.
Adobe is an equal opportunity employer. We hire hard-working individuals, regardless of gender, race or color, ethnicity or national origin, age, disability, religion, sexual orientation, gender identity or expression, or veteran status. We know that when our employees feel appreciated and included, they can be more creative, innovative and successful. This is what it means to be Adobe For All. Learn more about our vision here.
We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us to request accommodation.
Our compensation reflects the cost of labor across several U.S. geographic markets, and we pay differently based on those defined markets. The U.S. pay range for this positionis $205,900 -- $407,500 annually. Paywithin this range varies by work locationand may also depend on job-related knowledge, skills,and experience. Your recruiter can share more about the specific salary range for the job location during the hiring process.
At Adobe, for sales roles starting salaries are expressed as total target compensation (TTC = base + commission), and short-term incentives are in the form of sales commission plans. Non-sales roles starting salaries are expressed as base salary and short-term incentives are in the form of the Annual Incentive Plan (AIP).
In addition, certain roles may be eligible for long-term incentives in the form of a new hire equity award.
Adobe is proud to be an Equal Employment Opportunity and affirmative action employer. We do not discriminate based on gender, race or color, ethnicity or national origin, age, disability, religion, sexual orientation, gender identity or expression, veteran status, or any other applicable characteristics protected by law. Learn more.
#J-18808-Ljbffr- STN Inc in San Francisco is seeking an experienced AI Infrastructure Engineer to design, deploy, and manage large-scale GPU clusters for AI training and inference workloads. You will optimize GPU utilization, tune NCCL, CUDA, UCX, and Slurm, and work across storage, networking...SeniorTraining
- Autodesk, Inc. in San Francisco seeks a Senior Principal AI/ML Developer to shape data-driven personalization and analytic initiatives across the... ...collaborate with product, engineering, and marketing teams to deploy robust ML solutions at scale. #J-18808-Ljbffr AutodeskPrincipalSenior
$206.4k - $379.1k
...drives creativity at scale in design, imaging, motion... ....We're looking for a Principal Architect to build and implement... ..., merging strong ML skills with proficiency... ...infrastructure to support model training, fine-tuning,... ...intelligent systems.Mentor senior engineers and...PrincipalTrainingFull timeTemporary workLocal areaWorldwideFlexible hours$175k - $250k
...Cloud ML Software Architect – Cutting-Edge Biotech + AI Startup... ...$40M to date and are scaling rapidly. Compensation... ...infrastructure (AWS core, GPU compute, Docker &... ...processing and ML model training. Competitive... ...****@*****.*** . Seniority level: Mid‑Senior level...PrincipalTrainingFull timeH1bWork at officeVisa sponsorship3 days per week- ...engineering team — a small, senior group focused on... ...Science, AI/ML, or related field. PhD... ...least 2 years in a principal engineer or lead architect role. ~ Demonstrated... ...expertise across: LLM training and fine‑tuning,... ...construction, large‑scale data modeling. ~ Comfort...PrincipalTrainingWork at officeVisa sponsorshipFlexible hours3 days per week
- ...San Francisco is looking for a Senior Software Engineer to build... ...scalable infrastructure for large‑scale training and fine-tuning of foundation... ...training systems and optimize GPU utilization while collaborating... ...over 5 years of experience in ML infrastructure and a strong...SeniorTraining
- A cutting-edge AI technology company based in San Francisco is seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate...SeniorTraining
$160k - $225k
...Cacheflow is seeking a Senior Software Engineer for AI Runtime at Databricks, located in San Francisco. You will be instrumental in building and scaling systems for large-scale GPU training, ensuring high throughput and resilience in training across expansive fleets of...SeniorTraining- Principal or Senior Principal, Anthropic AI SolutionsAI Systems... ...sales teams, and solution architects that leads to pre-... ...enablement programs: designing training curricula, running... ...compliant adoption at scale. * Enable internal and... ...and delivering AI/ML or large language model...PrincipalSeniorTrainingTemporary workWork at officeLocal area
- ..., deploy, and maintain large distributed ML training and inference clusters Develop efficient... ...end-to-end pipelines to manage petabyte-scale datasets and model training throughout the... ...scales Analyze, profile and debug low-level GPU operations to optimize performance Stay...SeniorTraining
$179k - $218k
...who believe in the scale of our ambition and... ...bridged.We are seeking a Senior Staff Data Center... ...Engineer, GPU Hardware Architecture... ...Telemetry: Leverage AI/ML methodologies to... ...they impact customer training runs.Technical Sparing Architecture: Architect the site-level sparing...SeniorTrainingTemporary work$250k
...opportunities? Join a rapidly scaling AI cloud infrastructure... ...a next-generation GPU platform designed for AI training, experimentation, and inference... ...company is looking for a Senior / Staff Site Reliability Engineer... ...closely with platform, ML, and infrastructure teams...SeniorTrainingFull timeRemote work$280k - $350k
Principal Research Scientist - ScalingP-1227About Databricks AIAt Databricks... ...ranging from post-training open source LLMs to developing... ...be available to all.About the Scaling Research TeamThe Databricks AI... ...approaches.Optimize end‑to‑end ML systems for distributed training...PrincipalTrainingLocal areaWorldwide$148.5k - $313.7k
...and overall product quality at enterprise scale Contribute to strong engineering culture... ...Hands‑on experience applying AI/ML in product (LLMs, embeddings, or similar)... ...assignment, compensation, promotion, benefits, training, assessment of job performance, discipline...PrincipalSeniorTraining$288k
...domains. We are seeking a Senior or Principal Scientist to set the... ...formulation and architecture through training at scale, evaluation, and integration... ...and related structure-aware ML methods Set the... ...evaluation at scale across large GPU clusters Shape the end-...PrincipalSeniorTrainingFull timeWork at officeLocal areaFlexible hours$284.32k - $355.4k
...See yourself at Twilio Join the team as Twilio's next Senior Principal Field Architect - AI Agents About the job The Senior Principal... ..., and cross-functional initiatives at an enterprise scale with significant business impact. ~ Technical & Business...PrincipalSeniorLocal areaRemote workWorldwideFlexible hours$264.1k - $369.74k
...evolve rapidly as the constellation scales from first deployment through full operational... ...for critical operations. As the Senior Principal Architect for Systems & Mission Operations... ...hazardous materials transportation/shipping training. Required for certain Job Profiles:...PrincipalSeniorTrainingPermanent employmentTemporary workWork at officeLocal areaWorldwideRelocation$167.4k - $310.8k
...edge machine learning (ML) techniques. We are seeking a Senior or Principal Machine Learning Scientist... ..., designing and scaling large machine learning... ...for model architectures, training strategies, and evaluation... ...Systems & Engineering: Architect and improve large-scale...PrincipalSeniorTrainingFull timeLocal areaWorldwideRelocation package- Ginas Tech Jobs is seeking a Principal Machine Learning Engineer to set the technical standard for ML systems across training, inference, evaluation and deployment. This... ...pipelines. You will own large‑scale ML systems, optimize GPU memory and latency, collaborate with...PrincipalTrainingRemote job
- Blue Origin seeks a Senior Principal Architect for Systems & Mission Operations Software to own end-to-end software architecture across flight and... ...environment. The role requires deep expertise in large-scale distributed systems, C/C++, Rust, and strong leadership to...PrincipalSenior
$216.2k - $270.25k
Scale GP (Scale Generative AI Platform) is an enterprise-grade Generative... ...more. We are seeking a strong Senior Full-Stack Engineer to help us... ...systems, data pipelines, and ML/LLM components.Integrate with... ..., and relevant education or training. Scale employees in eligible...SeniorTrainingFull time- ...Palo Alto Networks, Inc. is seeking a Senior Principal Backend Engineer in the Cortex group to lead the development and scaling of the backend for Cortex XSOAR, XDR, and XSIAM. You will help design robust data pipelines, APIs, and services, partnering with product and...PrincipalSenior
- ...building production-grade ML infrastructure used by enterprise... .... They are looking for a Senior AI/ML Engineer to own model training pipelines, evaluation... ...and inference serving at scale. Full-time, on-site in San... ...with distributed training, GPU optimization, or inference...SeniorTrainingFull time
- ...Software Solutions is hiring a Senior Data Engineer (Apache... ...the design of large-scale distributed data processing... .... Responsibilities Architect and optimize large-scale... ...-tolerance Partner with ML engineers to deliver feature stores and training data sets at scale Drive...SeniorTrainingFlexible hours
$252k - $374k
...Science AI (LSAI), ML engineers build and... ...We are seeking a Principal ML Engineer to design, build, and scale the ML infrastructure... ...systems end to end, from training pipelines and... ...distributed training across GPU clusters Own... ...workflows Architect ML infrastructure that...PrincipalTrainingFull timeWork at officeLocal areaFlexible hours$260k - $340k
...believe in the scale of our ambition... ...This Role:As the Principal Systems Software... ...push massive-scale training workloads to the... ...(BMaaS): Architect systems that deliver raw GPU throughput via zero... ...alongside Staff and Senior engineers to... ...HPC projects.AI/ML Workload Expertise...PrincipalTrainingFull timeTemporary work$154.38k - $193.13k
Job DescriptionAI Architect, AI & AutomationAbout the Role The applicant... ...and implementing AI/ML solutions across multiple domains... ...clients, including clients at senior levels. Anticipates and proactively... ...outcomes, faster, smarter, and at scale.Infosys Consulting is helping...PrincipalFull timeTemporary workWork experience placement$197.3k - $313.7k
...stage startup backed by the global scale and trust of Salesforce.What You'... ...Actually Be DoingDesign, implement, and train novel deep learning models on large-scale GPU clusters.Prototype new... ...highly quantitative field with an AI/ML research focusYou possess experience...PrincipalTrainingFull timeImmediate startRemote work$127.4k - $191.1k
...integrated design practice. Our architects, engineers, interior... ...the World.We are looking for a Senior Architect who shares our interest... ...complex tasks on multiple, varying scale, highly complex projects.... ..., fringe benefits, job training, terminations or any other condition...SeniorTrainingFull timeContract workTemporary workPart timeFor contractorsFor subcontractorCasual workLive inWork at officeLocal areaFlexible hours- ...RoleWe are looking for a Principal AI Engineer to join our... ...complex problems at scale.In this role, you will... ...sense to know why your training run is slow.You design... ...meaningful open-source ML contributions; a research... ...Some exposure to multi-GPU training, even at lab scale...PrincipalTrainingFull timeTemporary workInternshipImmediate start
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Principal ML GPU Architect: Scale Training. Be the first to apply!
- senior lead project manager San Francisco, CA
- senior robotics software engineer San Francisco, CA
- senior devops engineer remote San Francisco, CA
- senior sas administrator San Francisco, CA
- senior IT manager San Francisco, CA
- senior director of client services San Francisco, CA
- senior contracts analyst San Francisco, CA
- senior implementation consultant San Francisco, CA
- sr project manager San Francisco, CA
- senior windows systems engineer San Francisco, CA




