AI Research Scientist: Multimodal Foundation Models - Architecture, Pre-Training & Distillation
$142.3k - $263.3kApple Inc.
AI Research Scientist: Multimodal Foundation Models - Architecture, Pre-Training & Distillation Seattle, Washington, United States Machine Learning and AI The Multimodal Intelligence Team is building the next generation of foundation models for Apple experiences. We are looking for a research scientist to advance the architectures, pre‑training methods, and distillation techniques that make highly capable multimodal models practical across the Apple ecosystem. Our research spans the full foundation‑model lifecycle: model architecture, pre‑training objectives, data mixtures, optimization, scaling, distillation, and evaluation. A defining challenge of our work is to develop models that combine broad intelligence with the memory, latency, energy, and privacy requirements of on‑device deployment. You will have the opportunity to shape new research directions, conduct ambitious experiments at scale, and translate successful ideas into foundation‑model technologies that can reach Apple products. Where appropriate, this work may also lead to publications and the open sourcing of selected models, research artifacts, evaluations, or tools. Description In this role, you will investigate fundamental questions about how multimodal foundation models should be designed, trained, and distilled. You will develop and evaluate new model architectures, pre‑training objectives, data strategies, optimization methods, and teacher–student learning techniques. Your work will explore how capabilities developed in large foundation models can be effectively transferred to smaller, more efficient models without treating distillation as an isolated downstream step. A major focus of the role will be the co‑development of frontier models and efficient models for Apple silicon and on‑device intelligence. This includes designing architectures that distill effectively, studying how teacher and student models should be trained together, and developing distillation methods that preserve reasoning, multimodal understanding, instruction following, and other important capabilities under constrained model capacity. Rather than treating deployment constraints as an afterthought, you will incorporate them into the research process—from early architecture experiments and pre‑training through distillation and final model evaluation. You may thrive in this role if you: Want to invent new foundation‑model architectures rather than only adapt existing models. Enjoy combining scientific ambition with real compute, memory, latency, and energy constraints. Believe that small and efficient models can be a frontier research problem, not merely a compression exercise. Are comfortable working across model research, data, systems, and hardware boundaries. Care about translating research into private, useful, and deeply integrated intelligent experiences. Want your work to have both product impact and a presence in the broader research community. Potential research directions include: Novel dense, recurrent, state‑space, mixture‑of‑experts, and hybrid foundation‑model architectures. Multimodal pre‑training across language, images, video, audio, and sensor‑derived representations. Compute‑optimal model and data scaling, including data mixtures, curricula, tokenization, and training objectives. Architecture and algorithm co‑design for memory‑efficient and energy‑efficient inference on Apple silicon. Offline and on‑policy distillation using teacher‑generated data, logits, representations, rationales, and other supervision signals. Minimum Qualifications Hands‑on experience designing, implementing, and running large‑scale pre‑training experiments for large language models. Experience with LLM pre‑training topics such as model architecture, training objectives, data mixtures, tokenization, curricula, scaling, and optimization. Strong proficiency with modern deep learning frameworks such as PyTorch or JAX and distributed training systems. Experience evaluating pre‑trained models across language understanding, reasoning, instruction following, or multimodal capabilities. Strong understanding of transformer‑based architectures and current approaches to efficient or scalable foundation‑model training. Master’s degree, or equivalent practical experience in machine learning, computer science, or a related technical field. Preferred Qualifications Experience contributing to major foundation‑model pre‑training efforts or leading architecture experiments that influenced a large training run. Research contributions in model architecture, scaling laws, multimodal pre‑training, optimization, efficient attention, mixture‑of‑experts, state‑space models, or related areas. Experience with knowledge distillation, including offline or off‑policy distillation, on‑policy distillation, self‑distillation, sequence‑level distillation, logic matching, or representation transfer. Experience designing teacher–student training pipelines or transferring capabilities from large foundation models to smaller models. Experience with multimodal models spanning language, vision, video, audio, or other sensor modalities. Understanding of inference efficiency, memory hierarchy, hardware accelerators, or hardware–software co‑design. Strong publication record, influential open‑source contributions, or an equivalent record of applied research impact. At Apple, base pay is one part of our total compensation package and is determined within a range. This provides an opportunity to progress as you grow and develop within a role. The base pay range for this role is between $142,300 and $263,300, and your base pay will depend on your skills, qualifications, experience, and location. Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses — including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program. Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant. At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong. Learn about accessibility in Apple’s workplace Learn about reasonable accommodations for job applicants #J-18808-Ljbffr Apple Inc.
- The Multimodal Intelligence Team is building the next generation of foundation models for Apple experiences. We are looking for a research scientist to advance the architectures, pre-training methods, and distillation techniques that make highly capable multimodal models...TrainingFoundation
- Apple Inc. in Seattle, WA seeks an AI Research Scientist focusing on multimodal foundation models—architecture, pre-training, and distillation—for on-device intelligence. You will push novel architectures and efficient training methods that translate to real Apple experiences...TrainingFoundation
- Apple seeks a Research Scientist to advance multimodal foundation models for on-device experiences. You will explore architectures, pre-training objectives, and distillation methods, aiming to balance capability with memory, latency, energy, and privacy requirements for...TrainingFoundation
- ...team focuses on applied research in Generative AI, delivering intelligent... ...cutting-edge areas including multimodal foundation models, image and video... ...models through large-scale training and post-training (e.g.,... ...designing efficient model architectures and advancing...TrainingFoundation
- ...Vision-Applied Research team... ...Generative AI and CV/Multimodal Understanding... ...generative models for content... ...Engineer / Scientist who can take... ...large model distillation and... ...capabilities from foundation models into... ...enabling scalable training,... ...algorithms and architectures for large-scale...TrainingFoundation
- ...machine learning models and systems to protect... ...multilingual and multimodal content, and... ...Multimodal moderation foundation model: We study large-scale MoE architecture training and routing... ...the frontier of AI today. This topic... ...direction and produce research outcomes with...TrainingFoundationFlexible hoursShift work
$159.75k - $255.6k
...and innovative Senior AI Research Scientist to join a new team... ...on agentic video and multimodal reasoning systems.... ...functional teams to train models and develop cutting-... ...retrieval systems, foundation models, agentic systems, or RAG architectures.Experience owning and...TrainingFoundationWork experience placementWork at officeRemote work$146.6k - $183.25k
...AI Research Scientist - AI Biological Design The Allen Institute... ...teams generate foundational knowledge, tools, and... ...and large-scale AI/ML models that turn complex biological... ...Develop, train, and evaluate large-scale... ...foundation models, and multimodal learning – to address...TrainingFoundationWork at officeLocal areaRemote workVisa sponsorshipWork visaRelocation package$202.16k - $368.22k
...Traditional models still face significant... ...to build a foundational large model... ..., and multimodal product... ...large‑scale training and generation... ...through model architecture and training... ...to the research community via... ...cutting‑edge AI technologies... ...quantization, pruning, distillation, and...TrainingFoundationTemporary workLocal area$57 per hour
...machine learning models and systems to protect... ...development of multimodal moderation foundation models, focusing on training stability,... ...large‑scale MoE architecture training, synthesis... ...cutting‑edge LLM research and practical experience... ...address complex AI challenges....TrainingFoundationHourly payInternshipLocal areaFlexible hours$167.8k - $209.7k
...AI Research Scientist II – Office of the CTO The Allen Institute... ...teams generate foundational knowledge, tools, and... ...to build large AI/ML models for Biology.... ...and interpretation of multimodal data to understand biological... ...pipeline for model training and evaluation....TrainingFoundationWork experience placementWork at officeLocal areaRemote workVisa sponsorshipWork visaRelocation package$113.7k - $211.9k
Adobe Research is looking for research scientists in Generative AI to join a world-class research team. We are interested... ..., large language models, and multimodal foundation models. Job responsibilities... ...large-scale generative model training Experience of working with...TrainingFoundationTemporary work$159.75k
...and innovative Senior AI Research Scientist to join a new team... ...on agentic video and multimodal reasoning systems.... ...functional teams to train models and develop cutting-... ...retrieval systems, foundation models, agentic systems, or RAG architectures. Experience owning and...TrainingFoundationWork experience placement- ByteDance in Seattle is seeking a researcher to push the boundaries of... ...machine learning, with a focus on multimodal understanding, video/text grounding, and scalable model training. The role requires a minimum... ...contribute to product‑driven AI initiatives and collaborate across...Training
- ...is seeking a Sr. Principal Applied Scientist to lead foundational model innovation for Ads. You will drive multimodal foundation models, scalable training, and production-ready solutions while... ...high-visibility role requires deep research expertise and collaboration with...TrainingFoundation
- Research Scientist Graduate (Foundation Model, Generative AI) - 2025 Start (PhD) Join ByteDance as a Research Scientist Graduate... ...in foundation models and multimodal machine learning, especially in... ...learning frameworks and large‑scale training experience. Experience in...TrainingFoundation
- Responsibilities Develop, train, and evaluate large-scale AI/ML models across diverse biological data types. Investigate... ...to address open biological research questions. Translate complex scientific... ...(e.g., PyTorch, JAX). Strong foundation in machine learning, deep learning...TrainingFoundation
$176k - $230k
...new era, we seek AI-native thinkers across... ...are hiring an AI Research Scientist (New Grad) for... ...and curate training data pipelines —... ...research experience) Foundational expertise in... ...tuning, or reasoning model development Demonstrated... ...with agentic architectures — including tool-...TrainingFoundationFull timeFlexible hours$202.16k - $368.22k
...team oversees the distributed training, reinforcement learning... ...compilation technologies for AI foundation models. Responsibilities Design... ...AI sector, the Seed team's research spans MLLM, GenMedia, AI for... ...models and cutting‑edge multimodal capabilities. Our technology...TrainingFoundationTemporary workInternshipLocal area- ...commerce Recommendation Foundation team is dedicated... ...Foundation Model that supports multi... ...language models (LLMs), multimodal understanding,... ...encourage both research thinking and engineering... ...in model training, inference optimization... ...Familiarity with pre-training and post-...TrainingFoundation
- ...traditional models still face significant... ...to build a foundational large model... ..., and multimodal product... ...large-scale training and generation... ...through model architecture and training... ...to the research community via... ...cutting-edge AI technologies... ...quantization, pruning, distillation, and...TrainingFoundation
$142.7k - $270.95k
...add Applied Scientists in Generative AI to our world... ...preparing data, training, fine-tuning... ...large foundation models across all modalities... ..., and multimodal priors.What... ...pioneering research and development... ...AI models (pre-training and... ...or distillation.Proficiency...TrainingFoundationFull timeTemporary workLocal areaWorldwide- ...is seeking a Principal Applied Scientist to lead real-time multimodal conversational AI research. You will drive foundation models for speech and audio and post-training systems, shaping natural, human... ...ready deployment. The role spans pre-training to post-training...TrainingFoundation
- An innovative AI startup in Seattle is looking for experts to... ...develop a groundbreaking human foundation model that integrates text, speech,... ...with a strong background in training audio generation models and... ...will work alongside a top-tier research team to create lifelike...TrainingFoundation
$232.56k - $427.5k
...in Seattle is seeking candidates pursuing a PhD in a relevant field to design and build scalable infrastructure for large-scale model training. The ideal candidate will have excellent coding abilities, especially in C/C++ or Python, and experience in distributed systems...TrainingFoundation- ...LLC is seeking talented researchers to join our Trust and Safety... ...will work on building multimodal moderation foundation models and agentic moderation systems... ...with PhD in CS/AI, strong Python/Rust/C++,... ...LLM community, optimize training and inference, and help scale...TrainingFoundation
$168k - $211k
...s Applied AI Team is living... .... AI Scientists work with... ...statistical modeling to solve complex... ...strong architectural thinking,... ...Research, Bioinformatics... ...Generative AI Foundation Models &... ...management. Multimodal &... ...pruning, distillation, and efficient... ...distributed training, data parallelism...TrainingFoundationWork experience placementWork at officeImmediate startRelocation3 days per week$232.56k - $427.5k
ByteDance is seeking a Research Scientist in AI Foundation Model Infrastructure to join their Seattle team. The candidate will be responsible for designing... ...and building infrastructure for large-scale model training, optimizing distributed systems, and improving system...TrainingFoundation- Research Scientist, LLM Evaluation & Post-Training page is loaded## Research Scientist... ...is a frontier AI data foundry... ...multilingual, pre-trained datasets... ...signals drive model improvement across... ...for LLM and multimodal systems, covering... ...in LLMs or foundation models (graduate...TrainingFoundationFull timeRemote work
- ...products and research, and to the... ...models still face... ...to build a foundational large model... ...series, and multimodal content into... ...largescale training, optimizing... ...via advanced architectures, post‑training... ...cutting‑edge AI technologies... ..., pruning, distillation, and...TrainingFoundationInternship
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Research Scientist: Multimodal Foundation Models - Architecture, Pre-Training & Distillation. Be the first to apply!

