AI Researcher (Multimodal Audio/Video Generation) at Series B multimodal AI lab
Jack & Jill
This is a job that Jill, our AI Recruiter, is recruiting for on behalf of one of our customers. She will pick the best candidates from Jack's network. The next step is to speak to Jack. Job Title AI Researcher (Multimodal Audio/Video Generation) Salary Not Disclosed Company Description Series B multimodal AI lab Job Description Lead research on audio-visual avatar generation at a cutting-edge lab building real-time conversational humans. You will design diffusion-based models for high-fidelity talking heads and neural avatars, bridging the gap between verbal and non-verbal communication. Partner with engineering teams to ship groundbreaking research into production-ready applications across healthcare and education. Location San Francisco, USA Why this role is remarkable Shape the future of human-computer interaction by building AI avatars that see, hear, and respond with human-like emotional intelligence. Join a well-funded Series B startup backed by top-tier global venture capital firms with a culture of shipping research directly to production. Work directly with founders and a high-caliber research team to define the architecture for real-time, multimodal conversational timing and rendering. What You Will Do Design and develop state-of-the-art diffusion models for long-video generation and audio-visual modeling to create life-like talking heads. Drive innovation in multimodal perception and neural rendering, capturing intricate verbal and non-verbal signals in perfectly synchronized flows. Mentor junior researchers and set strategic research directions while publishing impactful work at top-tier venues like CVPR and NeurIPS. The ideal candidate Holds a PhD in Computer Science or a related field with 2-3+ years of experience applying large-scale generative models in industry. Possesses deep expertise in diffusion models, PyTorch, and GPU-optimized workflows with a proven track record of publication at premier AI conferences. Demonstrates mastery of multimodal generation spanning video and audio, ideally with experience in 3D graphics or Gaussian splatting techniques. Who are Jack & Jill? Ok, I'll go first. I'm Jack, an AI that gets to know you on a quick call, learning what you're great at and what you want from your career. Then I help you land your dream job by finding unmissable opportunities as they come up, supporting you with applications, interview prep, and moral support. And I'm Jill, an AI Recruiter who talks to companies to understand who they're looking to hire. Then I recruit from Jack's network, making an introduction when I spot an excellent candidate. How does this work? Jack's an AI agent for job searching and career coaching. He works for you. Jill is the AI recruiter working for the company. She recruits from Jack's network. If it's a match and the company wants to meet you, they'll make the intro. In the meantime, if you'd like, Jack will send you excellent alternatives. We never post fake jobs This isn't a trick. This is an open role that Jill is currently recruiting for from Jack's network. Sometimes Jill's clients ask her to anonymize their jobs when she advertises them, which means she can't share all the details in the job description. We appreciate this can make them look a bit suspect, but there isn't much we can do about it. Give Jack a spin! You could land this role. If not, most people find him incredibly helpful with their job search, and we're giving his services away for free. #J-18808-Ljbffr Jack & Jill
$160k - $250k
...This is a job that Jill, our AI Recruiter, is recruiting... ...Company Description: Series B backed multimodal AI lab Job Description: You... ...a real-time conversational video interface, bridging the gap... ...You will collaborate with research teams to integrate state-of...VideoFull time$204k - $300k
...Technology Group (ATG) is the research division of the... ...engineering, such as AI/ML, algorithms,... ...digital signal processing, audio engineering, image... ...research leader in the Multimodal Experiences Lab, you will shape the... ...Invent and advance next-generation interactive and...AudioFull timeLocal areaWorldwideFlexible hours- Thinking Machines Lab in San Francisco is actively researching audio capabilities, blending rigorous theory with practical engineering to build multimodal AI systems. You will work across pre-training... ...develop models that understand and generate audio with high fidelity and...Audio
- AI Researcher (Computer Vision/Multimodal/Generative AI) About the Role We are hiring ML Researchers to develop novel approaches that advance the frontier of... ...‑on human‑centric visual representation learning video‑based modeling and temporal consistency multimodal...Video
- ...some of the highest-impact multimodal work in AI. ChatGPT serves a massive... ...experiences. We develop the research, training methods, and... ...advanced safety for image, video, or audio systems. In this role, you... ...video understanding, image generation, audio, or multimodal reasoning...VideoAudioWork at officeRelocation package
- ...improve dataset quality. You'll work directly with frontier AI labs to tackle challenging multimodal data problems, fine-tune models, build evaluation... ...measurable improvements in dataset quality across video, audio, images, and text. This is an onsite role requiring full...VideoAudioFull time
$250k
Research Scientist / Engineer - Multimodal Agent SF Bay Area, CA • Remote,... ...time About Luma AI Luma’s mission is... ...our research on video, 3D, and now... ...- text, video, audio, images - analogous... ...a recent $900M Series C and our partnership... ...that can generate, understand, and...VideoAudioFull timeRemote workWorldwide- ...adventure game where an AI companion is the... ...world's leading generative AI studio—we're the... ...define a new kind of video game experience.... ...Astra, and top-tier AI researchers. As an early... ...text-to-motion or audio-to-motion models.... ...worked at an industrial lab on this and related...VideoAudioWork at officeVisa sponsorship
- ...is hiring builders to join our Multimodal AI group, an industry-leading team defining the next generation of human-AI interaction. Our... ...include through voice, images, video, or new modalities we have yet... ...spans immersive UIs, realtime audio processing, evaluation systems...VideoAudio
- OpenAI is at the center of high-impact multimodal AI. The Chat and Multimodal Safety team builds safe, scalable models and evaluations for text, vision, and audio tasks. As a Researcher on the Chat and Multimodal Safety team in San Francisco, you will shape model perception...AudioWork at officeRelocation package
- Guidant Solutions Pvt. Ltd. is seeking an Applied Research Engineer to design scalable pipelines for large-scale video understanding. You will work on multimodal AI applications, including CV, audio, and NLP tasks, building production-ready systems. You will optimize inference...VideoAudio
$171.2k - $214k
Scale is the leading AI data foundry, helping fuel the most exciting... ...AI Product Manager to support Multimodal & Coding AI data verticals, including audio, image, video, and world models. This role is... ...operations, engineering, research, and go-to-market teams to ensure...VideoAudioFull time- About Us Tavus is a research lab pioneering human computing... .... We’re building AI Humans: a new... ...at scale. We’re a Series B company backed by world... ...Tavus’s Conversational Video Interface, our multimodal real‑time... ...systems , whether video, audio, or model inference,...VideoAudio
- ...users a day. We're hiring an AI Researcher to build the next generation of real-time, interactive... ...language processing, multimodal AI, human-computer interaction... ...modeling Knowledge of audio tokenization, neural... ...at an industrial research lab, major AI organization, technology...Audio
$117.2k - $313.7k
...AI Research Scientist And Research Engineer... ...coding agents. Multimodal and Computer Vision... ...language models (VLM), video understanding,... ..., and low-latency audio processing. Core... ...inference, and time series modeling. What... ...leave the lab and reach users at...VideoAudio- An innovative AI company based in California is seeking an experienced AI Researcher focusing on computer vision and multimodal AI. The candidate will develop novel architectures that improve various aspects of generative models, with direct implications for real-world...
- ...About Us Sieve is an AI research lab building the world's highest-quality multimodal datasets — spanning video, audio, images, text, and 3D. We combine exabyte-scale data infrastructure... ...of just ~25 people. We also raised our Series A from Tier 1 firms such as Matrix...VideoAudio
- ...team in San Francisco, focusing on video-language data collection, labeling,... ...repetitive tasks, and collaborate with research and product teams to drive... ...Python automation skills, experience with data pipelines, and a passion for multimodal AI. #J-18808-Ljbffr Twelve-LabsVideo
- ...Sieve is a multi-modal lab curating the world's highest... ...datasets — spanning video, audio, images, text, and 3D.... ...and novel multimodal understanding techniques... ...data. We partner with top AI labs and did $XXM last... ...people. We also raised our Series A from Tier 1 firms such...VideoAudio
- ...Intelligence is the product lab company behind... ...of state-of-the-art multimodal models across design,... ...dev, game dev, image, video, audio, slide generation, and more, through the... ...benchmark for AI-generated visuals, and... ...capabilities Publish research, technical reports, and...VideoAudioRelocationVisa sponsorship
- Postdoctoral Researcher, Computer Vision (PhD... ...role at Jobright.ai Postdoctoral Researcher... ...for the next generation of AI • Work with... ...(images, video, text, audio, speech and other... ...Professional Research Series Postdoctoral Fellow... ..., Orozco lab Geospatial Data Scientist...VideoAudioFull timePart timeWork experience placementInternship
- ...Intelligence is the product lab company behind... ...of state-of-the-art multimodal models across design... ..., game dev, image, video, audio, slide generation, and more, through... ...referenced benchmark for AI-generated visuals,... ...directly with the researchers at frontier labs to...VideoAudioRelocationVisa sponsorship
- Human Computer Lab in San Francisco is looking for a machine learning engineer to develop intelligent robotic systems. The candidate will work on multimodal models integrating computer vision, audio processing, and behavior. Key qualifications include 3+ years in machine...Audio
- ...About the Team API Multimodal builds the developer-facing... ...bring OpenAI’s image, audio, and real-time model capabilities... ...-scale APIs for image generation, speech transcription,... ...partner closely with Research and Inference to bring... ...help make multimodal AI useful at scale. Model...AudioFull timeInternship
- AI Researcher Location: San Francisco About Hum.ai is building planetary... ...edge, where we’re scaling generative transformer diffusion... ...do we do? We’re building multimodal foundation models for the... ...transformer models, especially on video or time series data. Preference for San...VideoRemote work
- ...Technical Staff, Lead Researcher San Francisco, CA;... ...DoorDash is building an AI Research org from... ..., and Dashers generating data that no academic lab and few companies can... ...-the-wild image and video capture to operational... ...operational settings Multimodal understanding — vision...VideoLocal area
$200.9k - $257.5k
...You Will Contribute To Altos Altos Labs is building a world-class AI ecosystem to solve the most complex... ...data. By implementing large-scale multimodal data fusion, you will move beyond simple... ..., and accessible to our global research community. Responsibilities Model...Local area$150k - $350k
...Our client is the only AI research lab exclusively focused on video data, combining exabyte-scale... ...with top AI labs and generated meaningful revenue last quarter... ...2022 · ~12 people (Series A) · Industry: AI Tools... ...across computer vision, audio processing, and text processing...VideoAudioFull timeH1bVisa sponsorship$350k
Thinking Machines Lab in San Francisco is looking for new team members to advance multimodal learning and visual perception science. This role involves developing AI architectures, datasets, and evaluation tools to fuse text and images effectively. The ideal candidate has...- ...Description Job Description The Research Role: We're looking for an experienced AI Researcher to join our team and... ...other modalities like 3d, text, video and audio. You may be a good fit if:... .... This one of the only labs in the world where you can combine...VideoAudioWork at officeVisa sponsorship
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Researcher (Multimodal Audio/Video Generation) at Series B multimodal AI lab. Be the first to apply!
- senior researcher San Francisco, CA
- machine learning researcher San Francisco, CA
- researcher San Francisco, CA
- senior design researcher San Francisco, CA
- design researcher San Francisco, CA
- qualitative researcher San Francisco, CA
- data collection researcher San Francisco, CA
- product researcher San Francisco, CA
- survey researcher San Francisco, CA
- legal researcher San Francisco, CA



