Audio AI Researcher: Multimodal, Low-Latency Modeling
Mosaic.tech
Thinking Machines Lab in San Francisco is actively researching audio capabilities, blending rigorous theory with practical engineering to build multimodal AI systems. You will work across pre-training, post-training, and product to develop models that understand and generate audio with high fidelity and low latency. The role requires deep experimentation, code writing, and collaboration with researchers, engineers, and designers to push the foundations of how AI learns and communicates. #J-18808-Ljbffr Mosaic.tech
- ...Lab in San Francisco is looking for researchers to advance audio capabilities in AI systems. This role blends... ...communication and collaboration through audio models. The ideal candidate will have a... ..., and experience in audio or multimodal models. We foster a collaborative...Audio
- ...building the human layer of AI. Our mission is to... ...through pioneering research in multimodal AI for modeling human-to-human communication (language, audio, and video), as well... ...trade-offs across latency, cost, and quality Partner... ...such as low‑rank adapters Strong...AudioRemote workRelocation packageFlexible hours
- ...architectures and training strategies for photorealistic virtual try‑on. You will work on multimodal learning, integrating vision, language, and video modalities into scalable models. A Ph.D. or Master’s in a relevant field is expected, along with proven experience in...Suggested
- ...veteran behind Project Astra, and top-tier AI researchers. As an early member of this team, you... ...of the best image and video generation model teams in the world on data collection.... ...experience in training text-to-motion or audio-to-motion models. You might have...AudioWork at officeVisa sponsorship
- ...and implementing infrastructure for large-scale multimodal models, focusing on high-performance delivery of audio and image inputs. You'll collaborate closely with researchers and product teams to push the boundaries of AI technology, ensuring reliable production...Audio
- A cutting-edge AI startup is searching for an experienced AI Researcher eager to advance generative AI. This role requires a PhD and 5+ years of research experience, focusing on developing innovative models that harness earth observation. The ideal candidate will demonstrate...Remote job
- ...based in San Francisco is seeking a Machine Learning Researcher to enhance K-12 education through AI. This role combines advanced technical skills with a... ...Candidates should possess expertise in generative AI models, Python programming, and have a graduate degree in a...Flexible hours
$180k
...s mission is to create AI systems that can accurately... .... All engineers and researchers are expected to have... ...About the Role The multimodal team at xAI creates magical... ...image, video, and audio. As a multimodal researcher... ..., you will drive the model’s multimodal capability...AudioLocal areaRelocation- We are seeking a talented Multimodal AI Systems Architect to develop and optimize AI systems that seamlessly integrate vision and audio models. This role focuses on enhancing our voice-to-voice... ...loops. Optimize streaming latency for voice-to-voice AI interactions....Audio
$180k - $260k
Machine Learning Researcher / Engineer, Multimodal LLMs Location: San Francisco... ...enterprises to build AI phone agents at scale... ...we are building the models and infrastructure... ...integrating streaming audio, tool execution, and... ...models You obsess over latency, correctness, and...AudioRemote jobWork at office- ...technology firm is seeking a Senior Applied Researcher to join their team in San Francisco.... ...focuses on building vision-language models that comprehend complex real-world environments... ..., tackling foundational problems in multimodal AI. Candidates should have experience in...
- ...photorealistic virtual try‑on and human‑centric visual representation. Advance multimodal learning by integrating vision, language, and video modalities into scalable generative models. Translate research prototypes into production‑ready systems, optimizing for realism,...
- AI Researcher (Computer Vision/Multimodal/Generative AI) About the Role We are hiring ML Researchers to develop novel approaches that advance the frontier... ...This role exists because current generative and vision models are not designed for photorealistic human...
- Cartesia, an AI company based in San Francisco, is hiring a Researcher for the Post‑Training team. You will design new techniques... ...for preference optimization, model evaluation, and feedback‑driven... ..., you’ll shape how Cartesia’s multimodal foundation models learn, align...
- An innovative AI company based in California is seeking an experienced AI Researcher focusing on computer vision and multimodal AI. The candidate will develop novel architectures that improve various aspects of generative models, with direct implications for real-world...
- ...D world where humanlike AI agents are able to interact... ...an exceptional AI researcher/engineer to join our team... ...high-level reasoning using multimodal LLMs with fast, low-level action models. The ideal candidate... ...architectures with 300-500ms latency Implementing multi-modal...
$350k
A leading AI research company in San Francisco is hiring for a position on their Audio team. The role involves developing and training advanced audio models, optimizing performance, and working collaboratively across teams. Ideal candidates will have strong expertise in...AudioFlexible hours$250k
Research Scientist / Engineer - Multimodal Agent SF Bay Area, CA • Remote, International • London... ...Full-time About Luma AI Luma’s mission is to build... ...video, 3D, and now multimodal models, we believe that AI needs... ...modalities - text, video, audio, images - analogous to the...AudioFull timeRemote workWorldwide- DeepRec.ai is hiring a Machine Learning Researcher, Audio to advance fully speech-to-speech conversational AI. You will contribute across speech-to-text, text... ...-audio reasoning, training at scale and delivering models into production fast. You will handle large audio datasets...AudioRemote work
- ...technology firm in San Francisco is seeking a Senior Applied Researcher in Audio Understanding to tackle complex audio perception tasks. You will... ...of traditional speech recognition, emphasizing large-scale model development and innovative research. The ideal candidate will...AudioRelocation package
- The Research Role: We're looking for an experienced AI Researcher to join our team and help us push the boundaries of AI... ...and develop large-scale diffusion models with a primary focus on images,... ...modalities like 3d, text, video and audio. You may be a good fit if: You love...AudioWork at officeVisa sponsorship
- ...dataset quality. You'll work directly with frontier AI labs to tackle challenging multimodal data problems, fine-tune models, build evaluation systems, and deliver... ...improvements in dataset quality across video, audio, images, and text. This is an onsite role requiring...AudioFull time
$191k - $273.6k
...Technology Group (ATG) ATG is the research division of Dolby. Its mission is... ...electrical engineering, such as AI/ML, algorithms, digital signal processing, audio engineering, image processing, computer... ...research at the intersection of multimodal AI and immersive sensory...AudioFull timeLocal areaWorldwideFlexible hours$310k
...our most advanced models - including our GPT... ...partner closely with Research to bring the next... ...boundaries of what AI can do. We're expanding into multimodal inference, building... ...that handle image, audio, and other non-text... ...for high-throughput, low-latency delivery of image...Audio- About Liquid AI Spun out of MIT CSAIL, we... ...hardware, ensuring low latency, minimal memory usage... ...Liquid Foundation Models across pre-training, vision, audio, and emerging... ...We treat data as a research problem, not an infrastructure... ..., audio, or multimodal. What We're Looking...Audio
$191k - $273.6k
Dolby Laboratories in San Francisco is seeking a leader in multimodal AI to drive innovative research. In this role, you will lead a team to develop cutting-edge technologies at the intersection of AI and sensory experiences. The ideal candidate has 8+ years of experience...Flexible hours- ...Technical Staff @ Lotus AI Lotus AI is a... ...may work across model training and fine-... ...for intelligent, multimodal patient... ...speech. Develop low-latency, streaming voice agents... ...that combine text, audio, and visual inputs... ...clinicians, and AI researchers to build something...Audio
- ...inference and RL‑driven training: making models significantly faster and more cost‑... ...engine and/or training stack. Have a solid research foundation in your area(s) of depth: Track... ..., and scheduling strategies for low‑latency, high‑throughput inference. Implement and...Full time
- ...institution based in San Francisco is looking for an Applied Researcher to work on AI-powered products. This role involves delivering innovative... ...Responsibilities include partnering with teams to build AI models and conducting impactful research. #J-18808-Ljbffr Capital...
- ...financial services firm in San Francisco is seeking an Applied Researcher II to develop innovative AI systems. In this role, you will collaborate with a cross-functional team to build AI foundation models, engage in impactful research, and translate complex work into business...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Audio AI Researcher: Multimodal, Low-Latency Modeling. Be the first to apply!
- remote researcher San Francisco, CA
- trend researcher San Francisco, CA
- survey researcher San Francisco, CA
- design researcher San Francisco, CA
- security researcher San Francisco, CA
- academic researcher San Francisco, CA
- content researcher San Francisco, CA
- criminal researcher San Francisco, CA
- senior design researcher San Francisco, CA
- court researcher San Francisco, CA

