Research Scientist, Foundation Model (Video Generation)
$185k - $400kPika
About the RoleAt Pika, we are pioneering the next generation of creative infrastructure built around real-time, multimodal generation and intelligent agentic platforms. We are seeking accomplished Research Scientists in Foundation Models with expertise in post-training large-scale multimodal foundation models to advance our mission of making agentic, real-time generative technology accessible and transformative for millions of creators. This is a staff and lead-level opportunity.As a key member of our research team, you will design and implement core technologies, develop new methodologies for large-scale multimodal post-training (text, image, audio, and video), and drive innovative approaches for foundational model architecture. You will collaborate closely with engineering and product teams, shaping the future of real-time creative and agentic platforms at scale.What You’ll DoLead research and development on post-training of multimodal foundation models at scale.Design and prototype novel algorithms and architectures for high-fidelity, real-time multimodal synthesis and interaction across modalities.Focus on scalable data pipeline curation and model training strategies for broad, diverse, and sensory-rich datasets.Advance state-of-the-art techniques in diffusion, autoregressive, and other generative models for large-scale post-training and fine-tuning.Identify, create, and leverage large, high-quality cross-modal datasets.Bring research advancements into production-ready systems in collaboration with engineering and product teams.Publish work in top-tier conferences and journals, and clearly communicate research both internally and externally.Stay at the forefront of foundational model and real-time multimodal AI research.What We’re Looking For5+ years of research experience in large-scale post-training of multimodal foundation models (LLMs, VLMs, Audio LMs, or similar), ideally at the staff or lead scientist level.Track record as a first author on major publications in top conferences or journals (e.g., NeurIPS, ICML, ICLR).Extensive hands-on experience with large-scale multimodal model design, training, and deployment.Deep understanding and implementation experience with generative architectures (diffusion, autoregressive, cross-modal, etc.).Expertise in high-throughput, scalable dataset curation and model pipeline optimization for multimodal applications.Strong programming and prototyping skills (Python, PyTorch, TensorFlow, etc.) and experience deploying research into production systems.Excellent communication and collaboration skills, and a passion for building creative enabling technology.What We OfferCompetitive salary and substantial equity in a high-growth startupFull health benefits + 401k matching and moreCollaborative, mission-driven team environment with major growth opportunitiesFlexible on-site/remote hybrid (HQ in Palo Alto, CA)About PikaPika empowers creators by building state-of-the-art agentic and multimedia platforms. Our vision is to break down technical barriers to creativity, making real-time generative and intelligent orchestration accessible to all. Join us and help shape the next evolution of creative technology!If you are a leading researcher excited to build and scale real-time multimodal foundation models, we want to hear from you.Compensation Range: $185K - $400KLocationPalo Alto HQEmployment TypeFull timeLocation TypeOn-siteDepartmentResearchCompensationUS locationBase Salary $185K – $400K • 0.1% – 0.3%
$190k - $250k
...are developing large-scale generative world models that learn to predict... ...capability serves as the foundation for scalable closed-loop training... .... We are looking for a research scientist to lead the design and development... ...realistic multi-camera video and LiDAR conditioned on...FoundationVideoTemporary workWork at officeVisa sponsorshipFlexible hours- ...About the Institute of Foundation Models We are a dedicated research lab for building, understanding... ..., nurture the next generation of AI builders, and... ...class researchers, data scientists, and engineers, tackling... ...involving images, text, and video. Understanding of modern...FoundationVideo
$193.93k - $352.29k
...partner-led business model, Nuro is working toward... ...collaborate closely with researchers and engineers on the... ...teams to tackle plan generation challenges in autonomous... ...models with foundation models. Leverage large... ...generative model optimization, video generation, text-to-...FoundationVideoImmediate startFlexible hours$192.2k - $260k
...at Amazon's Delivery Foundation Model team, where you'll work... ...world-class scientists and engineers to pioneer... ...modalities, including image, video, and geospatial data.... ...for specific research initiatives, ensuring... ...foundation models provide generative reasoning...FoundationVideoLocal areaWorldwideFlexible hours$185k - $400k
...are pioneering the next generation of creative... ...a staff or lead-level Research Engineer, Data to architect... ...engineering systems supporting model training for our advanced multimodal foundation models. This pivotal role... ..., image, audio, and video datasetsPartner with research...FoundationVideoRemote work- ...intelligence into the real world. Our foundation model, Newton, understands the... ...objective sensor data and generates real-time insights into... ...electrical signals, gases, video, and other sensor modalities... ...world. We are looking for an AI Researcher to design, train, and...FoundationVideo
$117.2k - $313.7k
...ExperienceSalesforce AI Research is a global... ...shape multiple generations of modern AI: we... ...introduced BLIP, a foundational breakthrough in multimodal... ...large language models including CodeGen... ...Research Scientists who want to build... ...language models, video understanding, and...FoundationVideoFull timeWorldwide$251k - $310k
...Staff Research Scientist, Perception Waymo is an autonomous driving... ...The mission of the Waymo AI Foundations team is to develop machine... ...from demonstration, generative modeling, Bayesian inference, hierarchical... ...as world models, images, videos, 3D, using techniques such...FoundationVideoFull timeTemporary workRemote work$230k - $270k
...observation alone, is an open research frontier. Our lab... ...wrong — built on multi-camera video understanding and multimodal... ...teleoperation, train open-source robot foundation models on our GPUs, and deploy them... ...real learned behaviors, generating the kind of behavioral data...FoundationVideoContract workTemporary workWorldwideFlexible hours$152k - $218.5k
Postdoctoral Researcher At Toyota Research Institute (TRI... ...exploring next-generation extended reality (XR)... ...generative AI and real-time video processing can enable... ..., ICML, ICLR). Strong foundation in machine learning, especially... .../video generative models. Strong programming...FoundationVideoWork at officeLocal areaShift work- ...Research Engineer Dyna Robotics builds general-purpose robots powered by a proprietary embodied AI foundation model with top-in-industry generalization and real-world performance. Already... ...DOF manipulation tasks. VLA & Video Generation: Architect and scale multi-modal...FoundationVideo
$272k - $431.25k
...OmniDreams and FlashDreams are the foundation for a new generation of interactive world-model systems. With this release,... ...workflows, and helps customers and researchers succeed with the stack. Success... ...applied innovation in world models, video diffusion, neural rendering,...FoundationVideoFull time$272k - $431.25k
...building the future, we’re generating it! Our world model team is pushing the... ...robotics, and world foundation models for Physical... ...for a Senior Research Manager to lead world... ...a team of Research Scientists focused on world-model... ...foundation models, including video models, vision-...FoundationVideoFull time$152k - $241.5k
...develop and adopt the next generation of Physical AI,... ...large-scale multimodal model training, robotics simulation... ...expertise in training foundation models at scale and a... ...of innovative AI research, accelerated computing... ...multimodal datasets (video, sensor data, trajectories...FoundationVideoFull time- ...About the Institute of Foundation Models We are a dedicated research lab for building, understanding... ..., nurture the next generation of AI builders, and... ...class researchers, data scientists, and engineers, tackling... ...Experience with large-scale video or multimodal data pipelines...FoundationVideoVisa sponsorship
- ...Senior / Staff AI Research Scientist, Foundation Models RoboForce is an AI robotics company developing Physical... ...and action-conditioned generative modeling for robot learning. Decent... ...Qualifications Experience with video generation or prediction models (e....FoundationVideoWork at officeVisa sponsorship
$174.72k - $295.68k
...Engineers with strong expertise in generative modeling and large-scale deep... ...skills. In this role, you will research, implement, and evaluate... ...diffusion and flow-matching models, video tokenizers, and transformer-... ...Experience with multimodal foundation models and video tokenizers...FoundationVideoFull time$192k - $304.75k
We are now looking for a Senior Research Scientist focused on Multimodal Foundation Models and Robotics! NVIDIA is searching for an outstanding research scientist... ...topics: LLMs; Large vision-language models; Video generative models and diffusion algorithms; or Action-based...FoundationVideoFull time$300k - $333k
...technically, focusing on developing model evaluations, identifying... ..., including DeepMind Research Engineers for technical delivery... ...to ensure firm engagement foundations.Analyze the enterprise AI... ...-tuning, agent harnesses), generative image, video, and audio models, and...FoundationVideo$218.8k - $335.3k
...intelligent software, and next-generation safety and entertainment features... ...global scale. Role : As a Staff Research Scientist in the AI Research organization,... ...AI/ML techniques (e.g., foundation models, VLAs, diffusion, image/video generation, self-supervised learning...FoundationVideoFull timeLocal areaRemote workWork from homeRelocation packageFlexible hours$165k - $195k
...The Bosch Research and Technology Center North America with... ...in Silicon Valley focuses on Foundation Models, Natural Language Processing... ...auxiliary modalities such as videos and images to enhance context... ...-centric AI, synthetic data generation, agentic AI ~ Proficiency...FoundationVideoFull timeWork experience placementLocal areaWorldwide$192k - $304.75k
NVIDIA is searching for a world-class generative AI researcher to join the fundamental generative AI... ...the next generation of generative models but also has an impact on the world.... ...expected to have a strong mathematical foundation and be capable of analyzing and developing...FoundationFull time$192k - $304.75k
We are now looking for a Senior Research Scientist for Generative AI!NVIDIA is searching for a world-class researcher... ..., including image generation, video generation, 3D generation, and audio... ...great impacts with generative AI models. You will be building research prototypes...VideoFull time$171.6k - $222.2k
...help develop the next generation of advanced robotics systems... ...and large language models.We leverage advanced... ...of robotics foundation models that: - Enable... ...candidate will contribute to research that bridges the gap between... ....As an Applied Scientist, you will develop and...FoundationLocal areaWorldwideFlexible hours- ...advanced augmented dexterity for next-generation robotic platforms. As a Senior AI/ML Research Engineer, you will develop and fine-tune the foundation models—VFMs, VLMs, and VLA models—that let... ..., and context from intraoperative video, and connecting perception to reasoning...FoundationVideoLocal areaWorldwideFlexible hours
$165k - $185k
Company DescriptionThe Bosch Research and Technology Center North America with offices in Sunnyvale, California, Pittsburgh, Pennsylvania... ...research, our AI research in Silicon Valley focuses on Foundation Models, Big Data Visual Analytics, Explainable AI (XAI), Natural...FoundationWork experience placementWorldwide- ...Founded by a team of Stanford researchers and entrepreneurs with deep... ...combines deep expertise in model innovation and systems engineering... ...your mark on an ambitious, generational mission to change how the... ...'re looking for a Research Scientist who can define what "better"...
$231.5k - $405.1k
...team Our Core AI Research team develops... ...We work across LLM model post-training, agent... ...a Staff Research Scientist, you will independently... ...failure modes, generate or curate data,... ...documents, images/video, and speech/audio... ...boundaries. ~ Strong foundations in machine...FoundationVideoFull timeWork at officeImmediate startRemote workFlexible hoursShift work$174.72k - $295.68k
...time Machine Learning Engineer / Research Scientist to drive the modeling and algorithmic development of XPENG’s next-generation Vision-Language-Action (VLA) Foundation Model — the core brain that... ...and unlabeled fleet data (images, video, LiDAR, CAN bus, maps, human driving...FoundationVideoFull time$224k - $356.5k
...building cutting-edge infrastructure for large-scale foundation model training in the Generalist Embodied Agent Research (GEAR) group. Our team is leading Project GR00T,... ...tailored for multimodal datasets, including videos, text, and sensor data.Develop robust monitoring...FoundationVideoFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Research Scientist, Foundation Model (Video Generation). Be the first to apply!
- applied scientist Palo Alto, CA
- remote scientist Palo Alto, CA
- operations research scientist Palo Alto, CA
- applied sports scientist Palo Alto, CA
- health scientist Palo Alto, CA
- drug safety scientist Palo Alto, CA
- scientist biology Palo Alto, CA
- safety scientist Palo Alto, CA
- machine learning research scientist Palo Alto, CA
- application scientist Palo Alto, CA


