Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Research Scientist, Foundation Model (Video Generation)

$185k - $400k

Pika

About the RoleAt Pika, we are pioneering the next generation of creative infrastructure built around real-time, multimodal generation and intelligent agentic platforms. We are seeking accomplished Research Scientists in Foundation Models with expertise in post-training large-scale multimodal foundation models to advance our mission of making agentic, real-time generative technology accessible and transformative for millions of creators. This is a staff and lead-level opportunity.As a key member of our research team, you will design and implement core technologies, develop new methodologies for large-scale multimodal post-training (text, image, audio, and video), and drive innovative approaches for foundational model architecture. You will collaborate closely with engineering and product teams, shaping the future of real-time creative and agentic platforms at scale.What You’ll DoLead research and development on post-training of multimodal foundation models at scale.Design and prototype novel algorithms and architectures for high-fidelity, real-time multimodal synthesis and interaction across modalities.Focus on scalable data pipeline curation and model training strategies for broad, diverse, and sensory-rich datasets.Advance state-of-the-art techniques in diffusion, autoregressive, and other generative models for large-scale post-training and fine-tuning.Identify, create, and leverage large, high-quality cross-modal datasets.Bring research advancements into production-ready systems in collaboration with engineering and product teams.Publish work in top-tier conferences and journals, and clearly communicate research both internally and externally.Stay at the forefront of foundational model and real-time multimodal AI research.What We’re Looking For5+ years of research experience in large-scale post-training of multimodal foundation models (LLMs, VLMs, Audio LMs, or similar), ideally at the staff or lead scientist level.Track record as a first author on major publications in top conferences or journals (e.g., NeurIPS, ICML, ICLR).Extensive hands-on experience with large-scale multimodal model design, training, and deployment.Deep understanding and implementation experience with generative architectures (diffusion, autoregressive, cross-modal, etc.).Expertise in high-throughput, scalable dataset curation and model pipeline optimization for multimodal applications.Strong programming and prototyping skills (Python, PyTorch, TensorFlow, etc.) and experience deploying research into production systems.Excellent communication and collaboration skills, and a passion for building creative enabling technology.What We OfferCompetitive salary and substantial equity in a high-growth startupFull health benefits + 401k matching and moreCollaborative, mission-driven team environment with major growth opportunitiesFlexible on-site/remote hybrid (HQ in Palo Alto, CA)About PikaPika empowers creators by building state-of-the-art agentic and multimedia platforms. Our vision is to break down technical barriers to creativity, making real-time generative and intelligent orchestration accessible to all. Join us and help shape the next evolution of creative technology!If you are a leading researcher excited to build and scale real-time multimodal foundation models, we want to hear from you.Compensation Range: $185K - $400KLocationPalo Alto HQEmployment TypeFull timeLocation TypeOn-siteDepartmentResearchCompensationUS locationBase Salary $185K – $400K • 0.1% – 0.3%

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Research Scientist, Foundation Model (Video Generation) in Palo Alto, CA vacancy
  • $190k - $250k

     ...are developing large-scale generative world models that learn to predict...  ...capability serves as the foundation for scalable closed-loop training...  .... We are looking for a research scientist to lead the design and development...  ...realistic multi-camera video and LiDAR conditioned on... 
    Foundation
    Video
    Temporary work
    Work at office
    Visa sponsorship
    Flexible hours

    Kodiak

    Mountain View, CA
    18 days ago
  •  ...About the Institute of Foundation Models We are a dedicated research lab for building, understanding...  ..., nurture the next generation of AI builders, and...  ...class researchers, data scientists, and engineers, tackling...  ...involving images, text, and video. Understanding of modern... 
    Foundation
    Video

    Institute of Foundation Models

    Sunnyvale, CA
    24 days ago
  • $193.93k - $352.29k

     ...partner-led business model, Nuro is working toward...  ...collaborate closely with researchers and engineers on the...  ...teams to tackle plan generation challenges in autonomous...  ...models with foundation models. Leverage large...  ...generative model optimization, video generation, text-to-... 
    Foundation
    Video
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    a month ago
  • $192.2k - $260k

     ...at Amazon's Delivery Foundation Model team, where you'll work...  ...world-class scientists and engineers to pioneer...  ...modalities, including image, video, and geospatial data....  ...for specific research initiatives, ensuring...  ...foundation models provide generative reasoning... 
    Foundation
    Video
    Local area
    Worldwide
    Flexible hours

    Amazon

    Santa Clara, CA
    11 days ago
  • $185k - $400k

     ...are pioneering the next generation of creative...  ...a staff or lead-level Research Engineer, Data to architect...  ...engineering systems supporting model training for our advanced multimodal foundation models. This pivotal role...  ..., image, audio, and video datasetsPartner with research... 
    Foundation
    Video
    Remote work

    Pika

    Palo Alto, CA
    a month ago
  •  ...intelligence into the real world. Our foundation model, Newton, understands the...  ...objective sensor data and generates real-time insights into...  ...electrical signals, gases, video, and other sensor modalities...  ...world. We are looking for an AI Researcher to design, train, and... 
    Foundation
    Video

    Jobleads-US

    Palo Alto, CA
    1 day ago
  • $117.2k - $313.7k

     ...ExperienceSalesforce AI Research is a global...  ...shape multiple generations of modern AI: we...  ...introduced BLIP, a foundational breakthrough in multimodal...  ...large language models including CodeGen...  ...Research Scientists who want to build...  ...language models, video understanding, and... 
    Foundation
    Video
    Full time
    Worldwide

    Salesforce

    Palo Alto, CA
    22 days ago
  • $251k - $310k

     ...Staff Research Scientist, Perception Waymo is an autonomous driving...  ...The mission of the Waymo AI Foundations team is to develop machine...  ...from demonstration, generative modeling, Bayesian inference, hierarchical...  ...as world models, images, videos, 3D, using techniques such... 
    Foundation
    Video
    Full time
    Temporary work
    Remote work

    Waymo

    Mountain View, CA
    18 hours ago
  • $230k - $270k

     ...observation alone, is an open research frontier. Our lab...  ...wrong — built on multi-camera video understanding and multimodal...  ...teleoperation, train open-source robot foundation models on our GPUs, and deploy them...  ...real learned behaviors, generating the kind of behavioral data... 
    Foundation
    Video
    Contract work
    Temporary work
    Worldwide
    Flexible hours

    Samsung SDS America

    Mountain View, CA
    24 days ago
  • $152k - $218.5k

    Postdoctoral Researcher At Toyota Research Institute (TRI...  ...exploring next-generation extended reality (XR)...  ...generative AI and real-time video processing can enable...  ..., ICML, ICLR). Strong foundation in machine learning, especially...  .../video generative models. Strong programming... 
    Foundation
    Video
    Work at office
    Local area
    Shift work

    Toyota Research Institute

    Los Altos, CA
    1 day ago
  •  ...Research Engineer Dyna Robotics builds general-purpose robots powered by a proprietary embodied AI foundation model with top-in-industry generalization and real-world performance. Already...  ...DOF manipulation tasks. VLA & Video Generation: Architect and scale multi-modal... 
    Foundation
    Video

    DYNA Robotics Inc

    Redwood City, CA
    18 hours ago
  • $272k - $431.25k

     ...OmniDreams and FlashDreams are the foundation for a new generation of interactive world-model systems. With this release,...  ...workflows, and helps customers and researchers succeed with the stack. Success...  ...applied innovation in world models, video diffusion, neural rendering,... 
    Foundation
    Video
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  • $272k - $431.25k

     ...building the future, we’re generating it! Our world model team is pushing the...  ...robotics, and world foundation models for Physical...  ...for a Senior Research Manager to lead world...  ...a team of Research Scientists focused on world-model...  ...foundation models, including video models, vision-... 
    Foundation
    Video
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  • $152k - $241.5k

     ...develop and adopt the next generation of Physical AI,...  ...large-scale multimodal model training, robotics simulation...  ...expertise in training foundation models at scale and a...  ...of innovative AI research, accelerated computing...  ...multimodal datasets (video, sensor data, trajectories... 
    Foundation
    Video
    Full time

    Nvidia

    Santa Clara, CA
    7 days ago
  •  ...About the Institute of Foundation Models  We are a dedicated research lab for building, understanding...  ..., nurture the next generation of AI builders, and...  ...class researchers, data scientists, and engineers, tackling...  ...Experience with large-scale video or multimodal data pipelines... 
    Foundation
    Video
    Visa sponsorship

    Institute of Foundation Models

    Sunnyvale, CA
    24 days ago
  •  ...Senior / Staff AI Research Scientist, Foundation Models RoboForce is an AI robotics company developing Physical...  ...and action-conditioned generative modeling for robot learning. Decent...  ...Qualifications Experience with video generation or prediction models (e.... 
    Foundation
    Video
    Work at office
    Visa sponsorship

    Embedding VC

    Milpitas, CA
    4 days ago
  • $174.72k - $295.68k

     ...Engineers with strong expertise in generative modeling and large-scale deep...  ...skills. In this role, you will research, implement, and evaluate...  ...diffusion and flow-matching models, video tokenizers, and transformer-...  ...Experience with multimodal foundation models and video tokenizers... 
    Foundation
    Video
    Full time

    XPENG Motors

    Santa Clara, CA
    24 days ago
  • $192k - $304.75k

    We are now looking for a Senior Research Scientist focused on Multimodal Foundation Models and Robotics! NVIDIA is searching for an outstanding research scientist...  ...topics: LLMs; Large vision-language models; Video generative models and diffusion algorithms; or Action-based... 
    Foundation
    Video
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  • $300k - $333k

     ...technically, focusing on developing model evaluations, identifying...  ..., including DeepMind Research Engineers for technical delivery...  ...to ensure firm engagement foundations.Analyze the enterprise AI...  ...-tuning, agent harnesses), generative image, video, and audio models, and... 
    Foundation
    Video

    Google

    Mountain View, CA
    7 days ago
  • $218.8k - $335.3k

     ...intelligent software, and next-generation safety and entertainment features...  ...global scale. Role : As a Staff Research Scientist in the AI Research organization,...  ...AI/ML techniques (e.g., foundation models, VLAs, diffusion, image/video generation, self-supervised learning... 
    Foundation
    Video
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    1 day ago
  • $165k - $195k

     ...The Bosch Research and Technology Center North America with...  ...in Silicon Valley focuses on Foundation Models, Natural Language Processing...  ...auxiliary modalities such as videos and images to enhance context...  ...-centric AI, synthetic data generation, agentic AI ~ Proficiency... 
    Foundation
    Video
    Full time
    Work experience placement
    Local area
    Worldwide

    Bosch Group

    Sunnyvale, CA
    12 days ago
  • $192k - $304.75k

    NVIDIA is searching for a world-class generative AI researcher to join the fundamental generative AI...  ...the next generation of generative models but also has an impact on the world....  ...expected to have a strong mathematical foundation and be capable of analyzing and developing... 
    Foundation
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  • $192k - $304.75k

    We are now looking for a Senior Research Scientist for Generative AI!NVIDIA is searching for a world-class researcher...  ..., including image generation, video generation, 3D generation, and audio...  ...great impacts with generative AI models. You will be building research prototypes... 
    Video
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  • $171.6k - $222.2k

     ...help develop the next generation of advanced robotics systems...  ...and large language models.We leverage advanced...  ...of robotics foundation models that: - Enable...  ...candidate will contribute to research that bridges the gap between...  ....As an Applied Scientist, you will develop and... 
    Foundation
    Local area
    Worldwide
    Flexible hours

    Amazon

    Sunnyvale, CA
    4 days ago
  •  ...advanced augmented dexterity for next-generation robotic platforms. As a Senior AI/ML Research Engineer, you will develop and fine-tune the foundation models—VFMs, VLMs, and VLA models—that let...  ..., and context from intraoperative video, and connecting perception to reasoning... 
    Foundation
    Video
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    Sunnyvale, CA
    a month ago
  • $165k - $185k

    Company DescriptionThe Bosch Research and Technology Center North America with offices in Sunnyvale, California, Pittsburgh, Pennsylvania...  ...research, our AI research in Silicon Valley focuses on Foundation Models, Big Data Visual Analytics, Explainable AI (XAI), Natural... 
    Foundation
    Work experience placement
    Worldwide

    Robert Bosch

    Sunnyvale, CA
    a month ago
  •  ...Founded by a team of Stanford researchers and entrepreneurs with deep...  ...combines deep expertise in model innovation and systems engineering...  ...your mark on an ambitious, generational mission to change how the...  ...'re looking for a Research Scientist who can define what "better"... 

    Sanas.AI Inc.

    Palo Alto, CA
    2 days ago
  • $231.5k - $405.1k

     ...team  Our Core AI Research team develops...  ...We work across LLM model post-training, agent...  ...a Staff Research Scientist, you will independently...  ...failure modes, generate or curate data,...  ...documents, images/video, and speech/audio...  ...boundaries.  ~ Strong foundations in machine... 
    Foundation
    Video
    Full time
    Work at office
    Immediate start
    Remote work
    Flexible hours
    Shift work

    ServiceNow

    Santa Clara, CA
    20 days ago
  • $174.72k - $295.68k

     ...time Machine Learning Engineer / Research Scientist to drive the modeling and algorithmic development of XPENG’s next-generation Vision-Language-Action (VLA) Foundation Model — the core brain that...  ...and unlabeled fleet data (images, video, LiDAR, CAN bus, maps, human driving... 
    Foundation
    Video
    Full time

    XPENG Motors

    Santa Clara, CA
    3 days ago
  • $224k - $356.5k

     ...building cutting-edge infrastructure for large-scale foundation model training in the Generalist Embodied Agent Research (GEAR) group. Our team is leading Project GR00T,...  ...tailored for multimodal datasets, including videos, text, and sensor data.Develop robust monitoring... 
    Foundation
    Video
    Full time

    Nvidia

    Santa Clara, CA
    a month ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Research Scientist, Foundation Model (Video Generation). Be the first to apply!