Research Scientist, Multimodal Generative AI (Image/Video)
$188k - $262kGoogle DeepMind
: Snapshot
The role of the Research Scientist will be to develop state-of-the-art methods for multimodal generative AI models, with a primary focus on image generation and editing .
At Google DeepMind, we've built a unique culture and work environment where long-term ambitious research can flourish. Our special interdisciplinary team combines the best techniques from deep learning, reinforcement learning, and systems neuroscience to build general-purpose learning algorithms. We have already made a number of high-profile breakthroughs towards building artificial general intelligence, and we have all the ingredients in place to make further significant progress over the coming year!
About UsArtificial Intelligence could be one of humanity's most useful inventions. At Google DeepMind, we're a team of scientists, engineers, machine learning experts, and more, working together to advance the state of the art in artificial intelligence. We use our technologies for widespread public benefit and scientific discovery, and collaborate with others on critical challenges, ensuring safety and ethics are the highest priority.
The RoleResearch Scientists at Google DeepMind lead our efforts in developing novel tools, infrastructure, and algorithms towards the end goal of solving and building Artificial General Intelligence.
Having pioneered research in the world's leading academic and industrial labs, PhDs, post-docs, or professorships, Research Scientists join Google DeepMind to work collaboratively within and across Research fields. They are expected to drive independent research initiatives, work with teams on large scale AI, and develop solutions to fundamental questions in machine learning and AI.
Drawing on expertise from a variety of disciplines including deep learning, computer vision, language modeling, and advanced generative architectures, our Research Scientists are at the forefront of groundbreaking research.
Key responsibilities:- Design, rapidly implement, and rigorously evaluate cutting-edge deep learning algorithms and data curation for multimodal generative AI, with a particular emphasis on image synthesis.
- Report and present research findings and developments clearly and efficiently both internally and externally, verbally and in writing.
- Suggest and engage in team collaborations to meet ambitious research goals, while also driving significant individual contributions.
- Work in collaboration with our Ethics and Governance teams to ensure our advances in intelligence are developed ethically and provide broad benefits to humanity.
In order to set you up for success as a Research Scientist at Google DeepMind, we look for the following skills and experience:
- PhD in Computer Science, Artificial Intelligence, Machine Learning, Computer Vision, or equivalent practical experience.
- Proven experience in deep learning research and development, particularly in generative AI and related to image synthesis. This includes diffusion models and autoregressive generative models.
- Exceptional engineering skills in Python and deep learning frameworks (e.g., Jax, TensorFlow, PyTorch), with a track record of building high-quality research prototypes and systems.
- Strong publication record at top-tier machine learning, computer vision, and graphics conferences (e.g., NeurIPS, ICLR, ICML, SIGGRAPH, CVPR, ICCV).
In addition, the following would be an advantage:
- Demonstrated experience in multimodal generative modeling, especially combining large language models with visual generation (e.g., text-to-image/video systems, joint autoregressive and diffusion models).
- A keen eye for visual aesthetics and detail, coupled with a passion for creating high-quality, visually compelling generative content.
- A real passion for AI!
The US base salary range for this full-time position is between $188,000 - $262,000 + bonus + equity + benefits. Your recruiter can share more about the specific salary range for your targeted location during the hiring process.
At Google DeepMind, we value diversity of experience, knowledge, backgrounds and perspectives and harness these qualities to create extraordinary impact. We are committed to equal employment opportunity regardless of sex, race, religion or belief, ethnic or national origin, disability, age, citizenship, marital, domestic or civil partnership status, sexual orientation, gender identity, pregnancy, or related condition (including breastfeeding) or any other basis as protected by applicable law. If you have a disability or additional need that requires accommodation, please do not hesitate to let us know.
$168k - $264.5k
...forefront of the AI revolution, and our research is shaping the future... ...for a Senior Scientist to join our team and... ...in synthetic data generation for training frontier... ..., structured, and multimodal data, directly... ...data generation — image, document, video, and audio — in partnership...VideoFull timeRemote work$85k - $150k
...We're seeking a full time Research Scientist, Artificial Intelligence (... ...optimize state-of-the-art AI systems for neural decoding... ..., artifact rejection, and multimodal fusion with video, audio, and IMU data.... ...designing self-supervised or generative models (diffusion, VAEs, contrastive...VideoFull time- About Mecka AI Mecka AI is building the data infrastructure... ...dedicated to the next generation of spatial and... .... We are hiring a Research Scientist to architect and train... ...Vision, and Temporal/Video Modeling. Proven experience... ..., multi-terabyte image and video datasets for...VideoShift work
$113.7k - $211.9k
Job responsibilities Conduct cutting‑edge research and development in Generative AI Develop and transfer novel technologies to Adobe products Research... ...diffusion models Strong publication record in audio/image/video generation and audio/image/video editing Experience...VideoTemporary work$229k - $343k
...digital services.Snap’s Generative ML Platform team builds cutting-edge AI technologies that... ...Snapchatters worldwide. From multimodal LLMs and video generation to real-... ...stay up-to-date with research and are excited about... ...diffusion models for images, videos or 3DKnowledge...VideoFull timeLive inWork at officeLocal areaWorldwide$209k - $313k
...digital services.Snap’s Generative ML Platform team builds cutting-edge AI technologies that... ...Snapchatters worldwide. From multimodal LLMs and video generation to real-... ...pipelines for image, video, language or audio... ...stay up-to-date with research and are excited about...VideoFull timeLive inWork at officeLocal areaWorldwide- ...Language Data Contributor (Multimodal) - Freelance AI Trainer Project World Wide... ...pace with cutting‑edge research, and streamline communication... ...to help power the next generation of AI. We’re looking for... ...include providing audio, video, image, and written inputs—including...VideoHourly payContract workFor contractorsFreelanceRemote work
- ...Language Data Contributor (Multimodal) - Freelance AI Trainer Project World Wide... ...pace with cutting‑edge research, and streamline communication... ...to help power the next generation of AI. We’re looking for... ...include providing audio, video, image, and written inputs—including...VideoHourly payContract workFor contractorsFreelanceRemote work
- ...security‑first enterprise AI company. We build... ...customers. Cohere is a team of researchers, engineers, designers,... ...in the power of multimodal AI to revolutionise the... ...multimodal tasks such as image or video captioning, speech‑to‑text generation. Bonus: Publications in...VideoFull timeWork at officeLocal areaRemote workHome office
- Scale AI, Inc. is seeking Research Scientists and Research Engineers with expertise in LLM post-training (SFT... ...LLM capabilities in text and multimodal modalities. You will develop novel... ...leading labs to influence the next generation of generative AI. The role emphasizes...
- Snap Inc. is seeking a Machine Learning Engineer to join the Generative ML Platform team. You will develop innovative ML technology and... ...models. You will work with state-of-the-art pipelines for image, video, language and audio generation, collaborating with cross-functional...Video
- Responsibilities: Conduct research and develop models to advance the state-of-the-art in generative computer vision technologies, with... ...customization of video content, including the creation... ...deploying neural networks for multimodal tasks. Strong understanding of...VideoLocal area
- TELUS Digital is seeking a Multimodal AI Content Specialist to join our Global Community. In this role, you'll evaluate the relationship... ...culturally resonant. Your responsibilities include auditing image and video datasets, conducting safety and bias detection, and helping...VideoRemote job
$150k - $200k
...is a human-centered AI lab focused on creative and multimodal outputs, where... ...taste defines the next generation of AI capabilities.... ...like across design, video, imagery, and... ...Labs is expanding its research work with frontier... ...domains such as design, image, video, and agentic...VideoWork at office$15 per hour
Welo Global in the United States is seeking a detail-oriented Generative AI Analyst to support AI training, data annotation, evaluation... ...diverse content types—from text and translations to audio, images, and video—with pay at 15 USD per hour. You will follow project...VideoRemote jobHourly payFreelance$147k - $211k
...Research Scientist, Multimodal Alignment, Safety, and Fairness Kirkland, Washington, US; Mountain View... ...Research Scientists with expertise in AI research and experience in... ...deep learning, computer vision, and generative architectures. This role requires independent...Full time$251k - $310k
...Staff Research Scientist, Perception Waymo is an autonomous... ...mission of the Waymo AI Foundations team is... ...from demonstration, generative modeling, Bayesian inference... ...state-of-the-art Multimodal LLMs and World models... ...as world models, images, videos, 3D, using techniques...VideoFull timeTemporary workRemote work$251k - $310k
...mission of the Waymo AI Foundations team is... ...with other research teams in Alphabet.... ...from demonstration, generative modeling, Bayesian... ...Principal Research Scientist . You Will Use Multimodal LLMs to perform 3D... ...such as world models, images, videos, 3D, human animation...VideoFull timeTemporary workRemote work- ...Prolific isn't just enabling AI innovation - we're... ...to train the next generation of AI models. Through... ...platform, we empower researchers and companies to access... ...deliver physical-world multimodal data collection from the... ...of large-scale video, audio and multimodal...VideoRemote work
$113.7k - $211.9k
Adobe is seeking a skilled researcher specializing in Generative AI to conduct transformative research and development. This role involves developing new technologies and collaborating with top-tier engineers and researchers. Candidates should possess a Ph.D. in related...$15 per hour
Generative AI Associate - Flexible Hours Remote - Georgia Innodata (Nasdaq... ...writer, linguist, educator, researcher, or just deeply passionate... ...original data, such as modifying images (rotation, flipping, cropping... ...), or altering audio/video signals (speed modification,...VideoHourly payFull timePart timeRemote workFlexible hoursShift work- ...Stack Developer at our rapidly scaling AI-driven video production studio, you will play a pivotal... ...the core technology that powers next-generation content creation. You'll architect and... ...Integrate AI/ML models (e.g., fine-tuned image/video models, generation APIs,...VideoContract workRemote work
- ...scalable ML systems that understand content across audio, video, text, and images. You will enable automation, improve quality, and unlock new... ...proficiency with PyTorch or TensorFlow, and a passion for multimodal data, system design, and cross-team collaboration. #J-188...Video
- ...made it possible for AI to impact clinical... ...time.Translational Scientist, Applied Machine... ...capable of hypothesis generation, experimental design, and multimodal ML modeling... ...transcriptomic and pathology imaging data, along with... ...collaboration: Work with Research, Engineering & Data...Full timeRemote workShift work
$55 - $90 per hour
...premier project with one of the world's top AI labs. This role pays between $55-90/hour.... ..., Statistics and Data Science, AI / ML Research, Computer Science, Game Development or... ...time including this and a few onboarding videos if you are hired. Currently, we are only...Video$320k - $400k
As a Research Scientist on our team, you will partner with Research Engineers... ...on our track record of AI-powered solutions (e.g., Bits... ...for Observability -- Training multimodal foundation models that learn... ...You'll Do:Conduct research in generative AI and machine learning,...- ...seeks a professional to build perception models for their AI companions, blending ML research with practical product needs. You will work in the... ...should have strong fundamentals in machine learning, a multimodal thinking approach, and the ability to execute research...Work at office
$148.1k - $282.1k
The Research and AI team builds foundational generative AI models and applications for Adobe products, enabling customers to ideate,... ...generative capabilities. This position supports multimodal generative capabilities across images, video, and more. It is a technical role with...VideoFull timeTemporary workLocal areaWorldwide$70k - $80k
...range $70,000.00/yr - $80,000.00/yr The AI Creative Video Producer sits at the intersection of... ...turn a simple concept into a stunning AI-generated video that inspires action. Responsibilities... .... Experiment with new AI tools (video, image, and text-to-motion) to push creative...VideoFull timeFlexible hours$227.5k - $324.99k
...from licensed catalog to user-generated content is trusted, safe,... ...grow—driven by advances in AI and new creation tools—we’... ...understanding of content across audio, video, text, and images—enabling automation,... ...in or have worked with multimodal machine learning You...VideoWork from homeWorldwideFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Research Scientist, Multimodal Generative AI (Image/Video). Be the first to apply!
- materials scientist New York, NY
- cosmetic scientist New York, NY
- scientist assay development New York, NY
- entry level research scientist New York, NY
- health scientist New York, NY
- quality control scientist New York, NY
- bioanalytical scientist New York, NY
- deep learning scientist New York, NY
- research associate scientist New York, NY
- application scientist New York, NY


