Audio AI Researcher: Multimodal, Low-Latency Modeling
Mosaic.tech
Thinking Machines Lab in San Francisco is actively researching audio capabilities, blending rigorous theory with practical engineering to build multimodal AI systems. You will work across pre-training, post-training, and product to develop models that understand and generate audio with high fidelity and low latency. The role requires deep experimentation, code writing, and collaboration with researchers, engineers, and designers to push the foundations of how AI learns and communicates. #J-18808-Ljbffr Mosaic.tech
- Kotoba’s speech models are licensed to Fortune... .... We’re hiring an AI Researcher to build the next... ..., prosody, and latency. Your work runs the... ...At our core is a low-latency, high-accuracy... ...processing, multimodal AI, human‑computer... ...modeling Knowledge of audio tokenization,...Audio
$204k - $300k
...Advanced Technology Group (ATG) is the research division of the company. ATG’s... ...electrical engineering, such as AI/ML, algorithms, digital signal processing, audio engineering, image processing,... ...a senior research leader in the Multimodal Experiences Lab, you will shape the...AudioFull timeLocal areaWorldwideFlexible hours- ...and implementing infrastructure for large-scale multimodal models, focusing on high-performance delivery of audio and image inputs. You'll collaborate closely with researchers and product teams to push the boundaries of AI technology, ensuring reliable production...Audio
- ...veteran behind Project Astra, and top-tier AI researchers. As an early member of this team, you... ...of the best image and video generation model teams in the world on data collection.... ...experience in training text-to-motion or audio-to-motion models. You might have...AudioWork at officeVisa sponsorship
- ...About Us [Tavus]( is a research lab pioneering human computing... .... We’re building AI Humans: a new interface... ...-time human simulation models let machines see, hear,... ...to lead research in audio-visual avatar generation... ...techniques. Experience in multimodal generation — spanning...AudioFull timeRemote work
- ...About the Team API Multimodal builds the developer... ...bring OpenAI’s image, audio, and real-time model capabilities into... ...speech generation, and low-latency voice interactions.... ...closely with Research and Inference to bring... ...help make multimodal AI useful at scale. Model...AudioFull timeInternship
$117.2k - $313.7k
...SalesforceSalesforce is the #1 AI CRM, where humans with... ...AI Research is a global leader in... ...foundational breakthrough in multimodal AI; trained state-of-the-art large language models including CodeGen; and... ...speech recognition, and low-latency audio processing.Core Modeling...AudioFull timeWorldwide$84.13 - $91.34 per hour
AI Researcher - Efficient AI (Contractor) Step into the innovative world... ...make modern LLMs, VLMs, multimodal models, and AI agents faster, smaller... ...methods (PTQ, QAT, pruning, low-rank approximation, etc) for... ..., constrained decoding, low-latency generation, and kernel-level...Full timeContract workTemporary workFor contractorsLocal areaImmediate start$117.2k - $313.7k
Salesforce AI Research is looking for outstanding AI Research Scientists... ...problems, develop novel models, and bridge the gap between... ..., and coding agents. Multimodal and Computer Vision: Vision... ...human‑like turn‑taking, and low‑latency audio processing. Core Modeling...Audio$114.2k - $306.6k
...duplicating efforts.*Salesforce Research advances state-of-the-art AI techniques, developing models and prototypes that pave the... ...autonomous workflows.* **Multimodal & Computer Vision:** Vision-... ...human-like turn-taking, and low-latency audio processing.* **Efficient...AudioFull time- ...our most advanced models - including our GPT... ...partner closely with Research to bring the next... ...boundaries of what AI can do. We’re expanding into multimodal inference, building... ...that handle image, audio, and other non-text... ...for high-throughput, low-latency delivery of image...AudioFull time
$206.3k - $388k
...architect and scale the multimodal data processing... ...multimodal foundation models (image, video, audio). In this role, you’ll... ...serving, throughput/latency tradeoffs) Experience... ...the full stack, from low-level systems and GPU... ...into impact, powered by AI and driven by human...AudioFull timeTemporary workLocal areaWorldwide- ...the OpenAI API, powering models including GPT-5 and a growing set of multimodal capabilities across text, image, audio, and video. Our team also... ...systems are a combination of low-latency, high scale, high... ...About OpenAI OpenAI is an AI research and deployment company dedicated...AudioFull time
- ...Tavus is a research lab pioneering... ...We’re building AI Humans: a new interface... ...simulation models let machines see... ...Interface, our multimodal real-time... ...concurrent programs. Low-level concepts... ...whether video, audio, or model... ...a millisecond latency budget feels like...AudioFull time
- ...is to architect AI that learns... ...pioneering the model architectures that... ...builds real-time multimodal intelligence:... ...We started as a research lab, pioneered... ...directly shape the latency, reliability,... ...that carries audio between users and... ..., SIP, RTP, or low-latency audio....AudioFull timeWork at officeVisa sponsorshipFlexible hours
- LG Electronics is seeking a Contract AI Researcher focusing on Efficient AI in Santa Clara, CA, hybrid work arrangement. You will explore model compression, quantization, efficient inference, and architectures to make LLMs/VLMs faster and more deployable on devices. You...Contract work
- ...What you’ll do Own a multimodal ML work-stream from problem definition... ...objectives, data strategies, model approaches, and success... ...workflow disruption. Communicate research findings and technical decisions... ...team as we build and deploy AI solutions that empower physicians...Full timeRemote workShift work
- ...institution based in San Francisco is looking for an Applied Researcher to work on AI-powered products. This role involves delivering innovative... ...Responsibilities include partnering with teams to build AI models and conducting impactful research. #J-18808-Ljbffr Capital...
$171.2k - $214k
Scale is the leading AI data foundry, helping fuel... ...AI, including frontier model training, enterprise adoption... ...Manager to support Multimodal & Coding AI data verticals, including audio, image, video, and world... ...operations, engineering, research, and go-to-market teams...AudioFull time- ...As a Data Engineer - Multimodal Systems , you will be... ...variety of modalities (text, audio, image) Designing and... ...others in a high-paced research setting Can rapidly... ...the bar to impact as low as possible We all... ...do and love discussing AI Benefits and Perks...AudioWork at officeRelocation package
- ...Technical Staff @ Lotus AI Who we are... ...may work across model training and fine-... ...for intelligent, multimodal patient interactions... ...speech. Develop low-latency, streaming voice... ...that combine text, audio, and visual inputs... ...clinicians, and AI researchers to build something...Audio
- ...Francisco, California. The Role: As a Research Engineer - Audio & Speech Models , you will be a core contributor on... ...we aim to minimize the bar to impact as low as possible We all enjoy what we do and love discussing AI Benefits and Perks: Comprehensive...AudioWork at officeRelocation package
- ...Bracket Bot is building low-cost, general-purpose robots... ...Kilpatrick (Google AI), Mohith Mothukuri (Physical... ...to own the full audio pipeline — from microphones... ...— under strict latency constraints. What you... ...stack: Voice-to-voice models (e.g. real-time speech systems...AudioFull time
- ...build, train, and serve AI models tailored to their own... ..., image, embedding, audio, and multimodal workloads. Today,... ...understand what inference latency, throughput, and cost... ...with AI research or open-model communities... ...infrastructure, from low-latency inference to...AudioPermanent employmentFlexible hours
$190.2k - $345.65k
...multimedia generative AI — deep-tuned image, video, and 3D models built on each... ...image, video, 3D, audio) and model-derived... ...into structured, low-latency, searchable intelligence... ...model-training or research role — you won't... ..., the hybrid and multimodal retrieval stack...AudioFull timeTemporary workLocal areaWorldwide$150k - $250k
About Distyl AI Distyl is an applied AI technology company partnering... ...social organizations.We research and deploy technologies that power... ...agents, evaluation harnesses, multimodal integrations, or workflow... ...effectiveExperience Building with Models, Not Just Building Models: We...Work at office3 days per week- ...on site in the specified location(s).As an AI Researcher within Schwab’s AI Strategy & Transformation... ...impact in real environments—where latency, reliability, cost, and regulatory considerations matter as much as model performance.In this role, you’ll take ideas...Full timeWork at office
$216.3k - $280.8k
Meet the TeamAt Foundation AI, we are leading frontier AI research across Cisco. Our mission is to advance the state of artificial intelligence... ...of enterprise AI.Our research spans foundation models, agentic AI, multimodal learning, reasoning systems, scalable training...Full timeTemporary workLocal areaFlexible hours- ...Description Job Description The Research Role: We're looking for an experienced AI Researcher to join our team and... ...and develop large-scale diffusion models with a primary focus on images,... ...modalities like 3d, text, video and audio. You may be a good fit if:...AudioWork at officeVisa sponsorship
$180k - $270k
...the world's most trusted AI work companion for... ...high-throughput, ultra-low-latency inference engines for large language models or foundational speech models... ...Token (or Time-To-First-Audio) in real-time streaming... ...collaborative, fast-paced research. Gear & Perks: Choice...AudioFull timeWork at officeWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Audio AI Researcher: Multimodal, Low-Latency Modeling. Be the first to apply!
- field researcher San Francisco, CA
- product researcher San Francisco, CA
- security researcher San Francisco, CA
- data collection researcher San Francisco, CA
- machine learning researcher San Francisco, CA
- court researcher San Francisco, CA
- researcher San Francisco, CA
- senior researcher San Francisco, CA
- independent researcher San Francisco, CA
- remote researcher San Francisco, CA




