Research Scientist, Foundation Model (Video Generation)
$185k - $400kPika
About the RoleAt Pika, we are pioneering the next generation of creative infrastructure built around real-time, multimodal generation and intelligent agentic platforms. We are seeking accomplished Research Scientists in Foundation Models with expertise in pre-training and mid-training large-scale multimodal foundation models to advance our mission of making agentic, real-time generative technology accessible and transformative for millions of creators. This is a staff and lead-level opportunity.As a key member of our research team, you will design and implement core technologies, develop new methodologies for large-scale multimodal pre-training/mid-training (text, image, audio, and video), and drive innovative approaches for foundational model architecture. You will collaborate closely with engineering and product teams, shaping the future of real-time creative and agentic platforms at scale.What You’ll DoLead research and development on pre-training and mid-training of multimodal foundation models at scale.Design and prototype novel algorithms and architectures for high-fidelity, real-time multimodal synthesis and interaction across modalities.Focus on scalable data pipeline curation and model training strategies for broad, diverse, and sensory-rich datasets.Advance state-of-the-art techniques in diffusion, autoregressive, and other generative models for large-scale pre-training and fine-tuning.Identify, create, and leverage large, high-quality cross-modal datasets.Bring research advancements into production-ready systems in collaboration with engineering and product teams.Publish work in top-tier conferences and journals, and clearly communicate research both internally and externally.Stay at the forefront of foundational model and real-time multimodal AI research.What We’re Looking For5+ years of research experience in large-scale pre-training/mid-training of multimodal foundation models (LLMs, VLMs, Audio LMs, or similar), ideally at the staff or lead scientist level.Track record as a first author on major publications in top conferences or journals (e.g., NeurIPS, ICML, ICLR).Extensive hands-on experience with large-scale multimodal model design, training, and deployment.Deep understanding and implementation experience with generative architectures (diffusion, autoregressive, cross-modal, etc.).Expertise in high-throughput, scalable dataset curation and model pipeline optimization for multimodal applications.Strong programming and prototyping skills (Python, PyTorch, TensorFlow, etc.) and experience deploying research into production systems.Excellent communication and collaboration skills, and a passion for building creative enabling technology.What We OfferCompetitive salary and substantial equity in a high-growth startupFull health benefits + 401k matching and moreCollaborative, mission-driven team environment with major growth opportunitiesFlexible on-site/remote hybrid (HQ in Palo Alto, CA)About PikaPika empowers creators by building state-of-the-art agentic and multimedia platforms. Our vision is to break down technical barriers to creativity, making real-time generative and intelligent orchestration accessible to all. Join us and help shape the next evolution of creative technology!If you are a leading researcher excited to build and scale real-time multimodal foundation models, we want to hear from you.Compensation Range: $185K - $400KLocationPalo Alto HQEmployment TypeFull timeLocation TypeOn-siteDepartmentResearchCompensationUS locationBase Salary $185K – $400K • 0.1% – 0.3%
$190k - $250k
...are developing large-scale generative world models that learn to predict... ...capability serves as the foundation for scalable closed-loop training... .... We are looking for a research scientist to lead the design and development... ...realistic multi-camera video and LiDAR conditioned on...FoundationVideoTemporary workWork at officeVisa sponsorship$160.36k - $240.54k
...partner-led business model, Nuro is working toward... ...scale state-of-the-art generative models—especially... ...generative models with foundation models. Leverage large... ...in industry, or both.Research experiences in generative... ...generative model optimization, video generation, text-to-...FoundationVideoImmediate startFlexible hours- ...About the Institute of Foundation Models We are a dedicated research lab for building, understanding... ..., nurture the next generation of AI builders, and... ...class researchers, data scientists, and engineers, tackling... ...involving images, text, and video. Understanding of modern...FoundationVideo
$192k - $304.75k
...seeking an outstanding Research Scientist or Research Engineer... ...for synthetic data generation and its application to... ...training autonomous driving models of tomorrow. As part... ...-language models, foundation models, or reasoning... ...foundation models, video understanding, or large...FoundationVideoFull time$192.2k - $260k
...at Amazon's Delivery Foundation Model team, where you'll work... ...world-class scientists and engineers to pioneer... ...modalities, including image, video, and geospatial data.... ...for specific research initiatives, ensuring... ...foundation models provide generative reasoning...FoundationVideoLocal areaWorldwideFlexible hours$168k - $264.5k
...Join our groundbreaking research team as we... ...of physical AI through generative models. We are hiring Research Scientists to join our Cosmos team... ..., focusing on advanced video generative models and video... ...harness 20,000+ GPUs for foundation models Author influential...FoundationVideo$158.3k - $297k
...the Role Entails1. Engage in the research and development of large-scale video world models, including the design and construction of training datasets, foundational model algorithm design,... ...implementation, research next-generation model architectures, and push the...FoundationVideoFull timeRelocation package- The Role We are looking for a Research Scientist to join the Multi-Embodiment Generalist Agent... ...a founding member. MEGA is building foundation models for general-purpose robots beyond not... ...models, multimodal and omni models, video models, reinforcement learning, imitation...FoundationVideoFull timeWork from home
- ...company in Mountain View is seeking a World Model Research Scientist. The successful candidate will design and train generative models, requiring expertise in AI and robotics... ...include developing techniques for realistic video synthesis and maintaining sensor data consistency...Video
$185k - $400k
...are pioneering the next generation of creative... ...a staff or lead-level Research Engineer, Data to architect... ...engineering systems supporting model training for our advanced multimodal foundation models. This pivotal role... ..., image, audio, and video datasetsPartner with research...FoundationVideoRemote work$117.2k - $313.7k
...ExperienceSalesforce AI Research is a global... ...shape multiple generations of modern AI: we... ...introduced BLIP, a foundational breakthrough in multimodal... ...large language models including CodeGen... ...Research Scientists who want to build... ...language models, video understanding, and...FoundationVideoFull timeWorldwide$151.8k - $218.21k
...development, as well as research impact via open-source software... ...for a driven research scientist with a strong background... ...learning and large scale foundation models (VLMs, text-to-video models, etc). The ideal... ...how we can create the next generation of AI-powered capable...FoundationVideo- ...pioneering general world models: causal,... ...long horizons. This foundational technology... ...together a world-class research team from... ...DeepMind Gemini), video models (DeepMind... ...ambitions, we need a scientist who doesn’t just... ...that reframe what generative interactive video...FoundationVideo
$251k - $310k
...mission of the Waymo AI Foundations team is to develop... ...collaborations with other research teams in Alphabet. AI... ...from demonstration, generative modeling, Bayesian inference,... ...a Principal Research Scientist. You will :... ...world models, images, videos, 3D, using techniques...FoundationVideoFull timeTemporary workRemote work$230k - $270k
...observation alone, is an open research frontier. Our lab... ...wrong — built on multi-camera video understanding and multimodal... ...teleoperation, train open-source robot foundation models on our GPUs, and deploy them... ...real learned behaviors, generating the kind of behavioral data...FoundationVideoContract workTemporary workWorldwideFlexible hours$272k - $431.25k
...OmniDreams and FlashDreams are the foundation for a new generation of interactive world-model systems. With this release,... ...workflows, and helps customers and researchers succeed with the stack. Success... ...applied innovation in world models, video diffusion, neural rendering,...FoundationVideoFull time- Tencent’s Technology Engineering Group (TEG) seeks a research-focused engineer to advance large-scale video world models, including data set design, model pre-training, SFT, RL, and downstream applications. You will analyze R&D challenges, optimize training and inference...Video
- You'll turn Luma's industry-leading generative video models into world models: interactive, controllable, physically faithful, and useful as... ...world, and own the metrics that define success. It fits a researcher with deep generative-modeling or model-based-RL expertise...Video
$152k - $218.5k
At Toyota Research Institute (TRI), we’re on a mission... ...exploring next‑generation extended reality (XR)... ...generative AI and real ‑time video processing can enable... ...ICML, ICLR). Strong foundation in machine learning, especially... .../video generative models. Strong programming...FoundationVideoWork at officeLocal areaShift work$152k - $241.5k
...develop and adopt the next generation of Physical AI,... ...large-scale multimodal model training, robotics simulation... ...expertise in training foundation models at scale and a... ...of innovative AI research, accelerated computing... ...multimodal datasets (video, sensor data, trajectories...FoundationVideoFull time$272k - $431.25k
...building the future, we’re generating it! Our world model team is pushing the... ...robotics, and world foundation models for Physical... ...for a Senior Research Manager to lead world... ...a team of Research Scientists focused on world-model... ...foundation models, including video models, vision-...FoundationVideoFull time$174.72k - $295.68k
...time Machine Learning Engineer / Research Scientist to drive the modeling and algorithmic development of XPENG’s next-generation Vision-Language-Action (VLA) Foundation Model — the core brain that... ...and unlabeled fleet data (images, video, LiDAR, CAN bus, maps, human driving...FoundationVideoFull time$174.72k - $295.68k
...Engineers with strong expertise in generative modeling and large-scale deep... ...skills. In this role, you will research, implement, and evaluate... ...diffusion and flow-matching models, video tokenizers, and transformer-... ...Experience with multimodal foundation models and video tokenizers...FoundationVideoFull time$192k - $304.75k
We are now looking for a Senior Research Scientist focused on Multimodal Foundation Models and Robotics! NVIDIA is searching for an outstanding research scientist... ...topics: LLMs; Large vision-language models; Video generative models and diffusion algorithms; or Action-based...FoundationVideoFull time$126k - $423k
...are looking for multiple passionate Research Scientists to join the Research Group at Applied... ...cutting‑edge technology enabling next‑generation physical AI, with emphasis on the two... ...research on pretraining world‑action foundation model with various world modalities including...FoundationFull timeFor contractorsFor subcontractorCasual workWork at officeImmediate startRemote workDay shift$300k - $333k
...technically, focusing on developing model evaluations, identifying... ..., including DeepMind Research Engineers for technical delivery... ...to ensure firm engagement foundations.Analyze the enterprise AI... ...-tuning, agent harnesses), generative image, video, and audio models, and...FoundationVideo$224k - $356.5k
...seeking a Senior Applied Research Scientist with experience... ...deploying deep learning models at scale across a... ...working on the next generation of data curation and... ...extraction pipelines for foundation-model training, including... ..., image, audio and videos) used in the training...FoundationVideoFull timeWork at officeRemote workFlexible hours- Institute of Foundation Models, operating the AllWorld Team at MBZUAI, seeks researchers to develop the PAN world models that simulate... ...data pipelines for high-quality video data. The role emphasizes... ...hands-on experience with video generative models, and proficiency in...FoundationVideo
$165k - $195k
Company DescriptionThe Bosch Research and Technology Center North... ...in Silicon Valley focuses on Foundation Models, Natural Language Processing... ...modalities such as videos and images to enhance context... ...-centric AI, synthetic data generation, agentic AIProficiency with...FoundationVideoFull timeWork experience placementLocal areaWorldwide$202.35k - $303.05k
...problems in autonomous driving. You will be focusing on researching and developing state of the art generative models, with an emphasis on diffusion models, and bring... ...driving and robotics. However, experiences in video generation, text-to-image generation, image in-painting...Video
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Research Scientist, Foundation Model (Video Generation). Be the first to apply!
- molecular biology scientist Palo Alto, CA
- water quality scientist Palo Alto, CA
- machine learning scientist Palo Alto, CA
- scientist antibody discovery Palo Alto, CA
- image scientist Palo Alto, CA
- machine learning research scientist Palo Alto, CA
- materials scientist Palo Alto, CA
- health scientist Palo Alto, CA
- lead scientist Palo Alto, CA
- scientist Palo Alto, CA


