Multimodal Audio-Video Enhancement Scientist
Velvet
A data research company in San Francisco seeks a Research Scientist to develop and enhance models for audiovisual data processing. This role involves researching novel methods for audio and video enhancement, running large-scale experiments, and collaborating with engineers for workflow integration. Ideal candidates have a strong background in deep learning, proficiency in PyTorch, and familiarity with signal processing. The environment is fast-paced, rewarding impactful applied research efforts. #J-18808-Ljbffr Velvet
- About Luma AI Luma’s mission is to build multimodal AGI. Through our research on video, 3D, and now multimodal models at Luma, we believe that AI needs to... ...jointly trained over all signal modalities - text, video, audio, images - analogous to the human brain. To advance our...VideoAudioWorldwide
- ...datasets that power the next generation of multimodal AI. Founded by Lucas Mantovani (ex Meta... ...frontier labs. We’re hiring a Research Scientist to develop and fine‑tune models for video and audio data processing and enhancement, as well as to conduct data‑oriented research...VideoAudioImmediate startShift work
$117.2k - $313.7k
...looking for outstanding AI Research Scientists and Research Engineers. Our... ...agents, and coding agents.Multimodal and Computer Vision: Vision-language models (VLM), video understanding, visual grounding... ...turn-taking, and low-latency audio processing.Core Modeling and Post...VideoAudioFull time- ...We train large‑scale robot foundation models from massive multimodal datasets spanning video, proprioception, action traces, language, and more. You... ...at scale (for generative models e.g. including language, audio, or video models) Have led or significantly contributed...VideoAudio
- Overview Research Scientist - VidGen (Relocation offered). This range is provided by... ...diffusion, rectified flow, and GANs) for video. Optimise proprietary infrastructure... ..., ICLR, etc.). Bonus: exposure to multimodal data (video + audio) or real-time systems. Work at the forefront...VideoAudioFull timeRelocationVisa sponsorshipRelocation package
$206.3k - $388k
...We’re looking for a Principal ML Engineer to architect and scale the multimodal data processing pipelines and infrastructure behind Adobe Firefly’s multimodal foundation models (image, video, audio). In this role, you’ll sit at the intersection of data engineering and...VideoAudioFull timeTemporary workLocal areaWorldwide- ...many users. You will own ML projects end-to-end, from research through deployment and iteration, and develop multimodal models across video, text, images, and audio. You’ll work with LLMs and external AI APIs to create content understanding models and semantic search...VideoAudio
$114.2k - $306.6k
...looking for outstanding AI Research Scientists / Research Engineers.**Our... ...and autonomous workflows.* **Multimodal & Computer Vision:** Vision-language models (VLM), video understanding, and visual... ...turn-taking, and low-latency audio processing.* **Efficient Systems...VideoAudioFull time- ...Team OpenAI is at the center of some of the highest-impact multimodal work in AI. ChatGPT serves a massive global audience, and enables... ...post‑training or evaluations, or advanced safety for image, video, or audio systems. This role is based in San Francisco, CA. We use a...VideoAudioWork at officeRelocation package
- ...'ll work directly with frontier AI labs to tackle challenging multimodal data problems, fine-tune models, build evaluation systems, and deliver measurable improvements in dataset quality across video, audio, images, and text. This is an onsite role requiring full-time...VideoAudioFull time
- .... Ltd. is seeking an Applied Research Engineer to design scalable pipelines for large-scale video understanding. You will work on multimodal AI applications, including CV, audio, and NLP tasks, building production-ready systems. You will optimize inference performance,...VideoAudio
$171.2k - $214k
...systems for the world's most important decisions.We are looking for an AI Product Manager to support Multimodal & Coding AI data verticals, including audio, image, video, and world models. This role is designed for someone early in their product career who is excited...VideoAudioFull time$141.1k - $281.7k
...technologies across all modalities, including text, image, and videos, and creates an industry-leading technical platform to improve... ...content and ranking based on a query to asset type affinity. Process Multimodal (Text, Image, Video) intent and contextually assist the user by...VideoTemporary work- ...achieve this through pioneering research in multimodal AI for modeling human-to-human communication (language, audio, and video), as well as generating audio-visual avatar... ...Role We’re looking for an experienced Research Scientist/Engineer with a focus on model optimization...VideoAudioRemote workRelocation packageFlexible hours
- ...hiring builders to join our Multimodal AI group, an industry-leading... ...include through voice, images, video, or new modalities we have... ...spans immersive UIs, realtime audio processing, evaluation systems... ...product managers, designers, data scientists, and go‑to‑market teams....VideoAudio
$204k - $300k
...algorithms, digital signal processing, audio engineering, image processing, computer... ...AccomplishAs a senior research leader in the Multimodal Experiences Lab, you will shape the... ...sensory technologies.• Drive projects that enhance multimodal content creation, delivery,...AudioFull timeLocal areaWorldwideFlexible hours$170k - $225k
...commerce ecosystems including video shopping platforms, social... ...continuous improvement \• Develop multimodal ML solutions processing video... ..., textual data, imagery, and audio at enterprise scale \•... ...cutting -edge AI tooling to enhance development velocity \• Partner...VideoAudioFull timeImmediate start$230k - $400k
...just on text but also other modes of data, including images, video and audio. Such models have potential to augment human creativity and... ...we are very concerned about the risks introduced by powerful multimodal AIs. The Multimodal team builds and studies multimodal...VideoAudioWork experience placementWork at officeHome officeVisa sponsorshipRelocation packageFlexible hours$157.9k - $284.4k
...creation. To help everyone create compelling videos, podcasts, and more, we are re-imagining music... ...process dramatically. We are looking for a research scientist (junior or senior) who is passionate about music generation, audio sequence modeling, and/or generative AI/ML, to...VideoAudioTemporary work- ...approaches across large language models (LLMs), speech models, and multimodal AI to build human-like voice agents capable of operating in... ...and human feedback. Build feedback loops that continuously enhance AI performance. Stay current with cutting-edge ML research. Evaluate...Audio
- ...the development of cutting‑edge multimodal foundation models that can comprehend videos just like humans do. Our models... ...External Partner Collaboration: Enhance dataset and process quality through... ...working with research scientists and engineers. Expertise or interest...VideoWork at officeWorldwideFlexible hours
- ...shape our real-time voice and video AI capabilities, building the foundation for intelligent, multimodal patient interactions. What... ...information. Continuously enhance retrieval accuracy and data lineage... ...models that combine text, audio, and visual inputs Knowledge...VideoAudio
$206.9k - $279.9k
...(CXI) team delivers the best audio and visual experiences for music... ...in one endpoint (e.g., video on Fire TV, voice on Alexa, personalization... ...technologies (generative AI, multimodal interfaces, ambient computing... ...with Engineers on product enhancements- Experience in project...VideoAudioLocal areaFlexible hoursDay shift- ...Position) III: A Declarative Data Management System for Semantic Multimodal Workflows National Science Foundation (NSF) - University of... ...Last verified Jul 27, 2026 The amount of data stored as videos, audio recordings, and documents is growing rapidly and now far exceeds...VideoAudio
- ...our startup journey. We create engaging videos covering the job market, stock analysis,... ...text overlays, transitions, and B-roll to enhance storytelling Create engaging thumbnails and... ...align with channel branding Optimize audio quality, color correction, and pacing for...VideoAudioPart timeInternship
- Thinking Machines Lab in San Francisco is actively researching audio capabilities, blending rigorous theory with practical engineering to build multimodal AI systems. You will work across pre-training, post-training, and product to develop models that understand and generate...Audio
- ...machine learning through large-scale video understanding and multimodal data. The company develops advanced... ...address challenges in computer vision, audio processing, and natural language... ...batching, caching, and other performance enhancements. Leverage foundation models and...VideoAudioFull timeH1bVisa sponsorship
- OpenAI is at the center of high-impact multimodal AI. The Chat and Multimodal Safety team builds safe, scalable models and evaluations for text, vision, and audio tasks. As a Researcher on the Chat and Multimodal Safety team in San Francisco, you will shape model perception...AudioWork at officeRelocation package
- ...building open, state-of-the-art models for video generation towards unlocking the right... ...overview:We are seeking an exceptional Research Scientist to join our team, focusing on alignment... ...teams to ensure alignment methods enhance rather than inhibit model capabilitiesQualifications...VideoRelocation
$228.7k - $309.4k
...podcast catalogs.As a Principal Applied Scientist, you'll drive innovation in music content... ...systems, leveraging foundational models to enhance content intelligence and understanding-... ...our customers discover the most relevant audio content including music, podcasts and audiobooks...AudioTemporary workLocal areaWorldwideFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Multimodal Audio-Video Enhancement Scientist. Be the first to apply!
- scientist ii San Francisco, CA
- scientist 1 San Francisco, CA
- image scientist San Francisco, CA
- downstream processing scientist San Francisco, CA
- entry level research scientist San Francisco, CA
- qc scientist San Francisco, CA
- research scientist San Francisco, CA
- analytical scientist San Francisco, CA
- research scientist - biology San Francisco, CA
- genomics scientist San Francisco, CA


