Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Audio AI Researcher: Multimodal, Low-Latency Modeling

Mosaic.tech

Thinking Machines Lab in San Francisco is actively researching audio capabilities, blending rigorous theory with practical engineering to build multimodal AI systems. You will work across pre-training, post-training, and product to develop models that understand and generate audio with high fidelity and low latency. The role requires deep experimentation, code writing, and collaboration with researchers, engineers, and designers to push the foundations of how AI learns and communicates. #J-18808-Ljbffr Mosaic.tech

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Audio AI Researcher: Multimodal, Low-Latency Modeling in San Francisco, CA vacancy
  • Kotoba’s speech models are licensed to Fortune...  .... We’re hiring an AI Researcher to build the next...  ..., prosody, and latency. Your work runs the...  ...At our core is a low-latency, high-accuracy...  ...processing, multimodal AI, human‑computer...  ...modeling Knowledge of audio tokenization,... 
    Audio

    Kotoba

    San Francisco, CA
    5 days ago
  • $204k - $300k

     ...Advanced Technology Group (ATG) is the research division of the company. ATG’s...  ...electrical engineering, such as AI/ML, algorithms, digital signal processing, audio engineering, image processing,...  ...a senior research leader in the Multimodal Experiences Lab, you will shape the... 
    Audio
    Full time
    Local area
    Worldwide
    Flexible hours

    Dolby

    San Francisco, CA
    5 days ago
  •  ...and implementing infrastructure for large-scale multimodal models, focusing on high-performance delivery of audio and image inputs. You'll collaborate closely with researchers and product teams to push the boundaries of AI technology, ensuring reliable production... 
    Audio

    Jobleads-US

    San Francisco, CA
    5 days ago
  •  ...veteran behind Project Astra, and top-tier AI researchers. As an early member of this team, you...  ...of the best image and video generation model teams in the world on data collection....  ...experience in training text-to-motion or audio-to-motion models. You might have... 
    Audio
    Work at office
    Visa sponsorship

    Spellbrush

    San Francisco, CA
    2 days ago
  •  ...About Us [Tavus]( is a research lab pioneering human computing...  .... We’re building AI Humans: a new interface...  ...-time human simulation models let machines see, hear,...  ...to lead research in audio-visual avatar generation...  ...techniques. Experience in multimodal generation — spanning... 
    Audio
    Full time
    Remote work

    Tavus

    San Francisco, CA
    11 days ago
  •  ...About the Team API Multimodal builds the developer...  ...bring OpenAI’s image, audio, and real-time model capabilities into...  ...speech generation, and low-latency voice interactions....  ...closely with Research and Inference to bring...  ...help make multimodal AI useful at scale. Model... 
    Audio
    Full time
    Internship

    OpenAI

    San Francisco, CA
    1 day ago
  • $117.2k - $313.7k

     ...SalesforceSalesforce is the #1 AI CRM, where humans with...  ...AI Research is a global leader in...  ...foundational breakthrough in multimodal AI; trained state-of-the-art large language models including CodeGen; and...  ...speech recognition, and low-latency audio processing.Core Modeling... 
    Audio
    Full time
    Worldwide

    Salesforce

    San Francisco, CA
    3 days ago
  • $84.13 - $91.34 per hour

    AI Researcher - Efficient AI (Contractor) Step into the innovative world...  ...make modern LLMs, VLMs, multimodal models, and AI agents faster, smaller...  ...methods (PTQ, QAT, pruning, low-rank approximation, etc) for...  ..., constrained decoding, low-latency generation, and kernel-level... 
    Full time
    Contract work
    Temporary work
    For contractors
    Local area
    Immediate start

    LG Electronics

    San Francisco, CA
    3 days ago
  • $117.2k - $313.7k

    Salesforce AI Research is looking for outstanding AI Research Scientists...  ...problems, develop novel models, and bridge the gap between...  ..., and coding agents. Multimodal and Computer Vision: Vision...  ...human‑like turn‑taking, and low‑latency audio processing. Core Modeling... 
    Audio

    salesforce.com, inc.

    San Francisco, CA
    3 days ago
  • $114.2k - $306.6k

     ...duplicating efforts.*Salesforce Research advances state-of-the-art AI techniques, developing models and prototypes that pave the...  ...autonomous workflows.* **Multimodal & Computer Vision:** Vision-...  ...human-like turn-taking, and low-latency audio processing.* **Efficient... 
    Audio
    Full time

    Niebles

    San Francisco, CA
    4 days ago
  •  ...our most advanced models - including our GPT...  ...partner closely with Research to bring the next...  ...boundaries of what AI can do. We’re expanding into multimodal inference, building...  ...that handle image, audio, and other non-text...  ...for high-throughput, low-latency delivery of image... 
    Audio
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  • $206.3k - $388k

     ...architect and scale the multimodal data processing...  ...multimodal foundation models (image, video, audio). In this role, you’ll...  ...serving, throughput/latency tradeoffs) Experience...  ...the full stack, from low-level systems and GPU...  ...into impact, powered by AI and driven by human... 
    Audio
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Francisco, CA
    5 days ago
  •  ...the OpenAI API, powering models including GPT-5 and a growing set of multimodal capabilities across text, image, audio, and video. Our team also...  ...systems are a combination of low-latency, high scale, high...  ...About OpenAI OpenAI is an AI research and deployment company dedicated... 
    Audio
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...Tavus is a research lab pioneering...  ...We’re building AI Humans: a new interface...  ...simulation models let machines see...  ...Interface, our multimodal real-time...  ...concurrent programs. Low-level concepts...  ...whether video, audio, or model...  ...a millisecond latency budget feels like... 
    Audio
    Full time

    Tavus

    San Francisco, CA
    1 day ago
  •  ...is to architect AI that learns...  ...pioneering the model architectures that...  ...builds real-time multimodal intelligence:...  ...We started as a research lab, pioneered...  ...directly shape the latency, reliability,...  ...that carries audio between users and...  ..., SIP, RTP, or low-latency audio.... 
    Audio
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Cartesia

    San Francisco, CA
    20 hours ago
  • LG Electronics is seeking a Contract AI Researcher focusing on Efficient AI in Santa Clara, CA, hybrid work arrangement. You will explore model compression, quantization, efficient inference, and architectures to make LLMs/VLMs faster and more deployable on devices. You... 
    Contract work

    LG Electronics

    San Francisco, CA
    5 days ago
  •  ...What you’ll do Own a multimodal ML work-stream from problem definition...  ...objectives, data strategies, model approaches, and success...  ...workflow disruption. Communicate research findings and technical decisions...  ...team as we build and deploy AI solutions that empower physicians... 
    Full time
    Remote work
    Shift work

    Radai

    San Francisco, CA
    9 days ago
  •  ...institution based in San Francisco is looking for an Applied Researcher to work on AI-powered products. This role involves delivering innovative...  ...Responsibilities include partnering with teams to build AI models and conducting impactful research. #J-18808-Ljbffr Capital... 

    Capital One

    San Francisco, CA
    3 days ago
  • $171.2k - $214k

    Scale is the leading AI data foundry, helping fuel...  ...AI, including frontier model training, enterprise adoption...  ...Manager to support Multimodal & Coding AI data verticals, including audio, image, video, and world...  ...operations, engineering, research, and go-to-market teams... 
    Audio
    Full time

    Scale AI

    San Francisco, CA
    5 days ago
  •  ...As a Data Engineer - Multimodal Systems , you will be...  ...variety of modalities (text, audio, image) Designing and...  ...others in a high-paced research setting Can rapidly...  ...the bar to impact as low as possible We all...  ...do and love discussing AI Benefits and Perks... 
    Audio
    Work at office
    Relocation package

    Zyphra

    San Francisco, CA
    11 days ago
  •  ...Technical Staff @ Lotus AI Who we are...  ...may work across model training and fine-...  ...for intelligent, multimodal patient interactions...  ...speech. Develop low-latency, streaming voice...  ...that combine text, audio, and visual inputs...  ...clinicians, and AI researchers to build something... 
    Audio

    Lotus Health AI, Inc

    San Francisco, CA
    5 days ago
  •  ...Francisco, California. The Role: As a Research Engineer - Audio & Speech Models , you will be a core contributor on...  ...we aim to minimize the bar to impact as low as possible We all enjoy what we do and love discussing AI Benefits and Perks: Comprehensive... 
    Audio
    Work at office
    Relocation package

    Zyphra

    San Francisco, CA
    11 days ago
  •  ...Bracket Bot is building low-cost, general-purpose robots...  ...Kilpatrick (Google AI), Mohith Mothukuri (Physical...  ...to own the full audio pipeline — from microphones...  ...— under strict latency constraints. What you...  ...stack: Voice-to-voice models (e.g. real-time speech systems... 
    Audio
    Full time

    Bracket Bot Inc.

    San Francisco, CA
    1 day ago
  •  ...build, train, and serve AI models tailored to their own...  ..., image, embedding, audio, and multimodal workloads. Today,...  ...understand what inference latency, throughput, and cost...  ...with AI research or open-model communities...  ...infrastructure, from low-latency inference to... 
    Audio
    Permanent employment
    Flexible hours

    Fireworks AI

    San Francisco, CA
    2 days ago
  • $190.2k - $345.65k

     ...multimedia generative AI — deep-tuned image, video, and 3D models built on each...  ...image, video, 3D, audio) and model-derived...  ...into structured, low-latency, searchable intelligence...  ...model-training or research role — you won't...  ..., the hybrid and multimodal retrieval stack... 
    Audio
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Francisco, CA
    3 days ago
  • $150k - $250k

    About Distyl AI Distyl is an applied AI technology company partnering...  ...social organizations.We research and deploy technologies that power...  ...agents, evaluation harnesses, multimodal integrations, or workflow...  ...effectiveExperience Building with Models, Not Just Building Models: We... 
    Work at office
    3 days per week

    Distyl AI

    San Francisco, CA
    3 days ago
  •  ...on site in the specified location(s).As an AI Researcher within Schwab’s AI Strategy & Transformation...  ...impact in real environments—where latency, reliability, cost, and regulatory considerations matter as much as model performance.In this role, you’ll take ideas... 
    Full time
    Work at office

    The Charles Schwab Corporation

    San Francisco, CA
    7 days ago
  • $216.3k - $280.8k

    Meet the TeamAt Foundation AI, we are leading frontier AI research across Cisco. Our mission is to advance the state of artificial intelligence...  ...of enterprise AI.Our research spans foundation models, agentic AI, multimodal learning, reasoning systems, scalable training... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Francisco, CA
    6 days ago
  •  ...Description Job Description The Research Role: We're looking for an experienced AI Researcher to join our team and...  ...and develop large-scale diffusion models with a primary focus on images,...  ...modalities like 3d, text, video and audio. You may be a good fit if:... 
    Audio
    Work at office
    Visa sponsorship

    Spellbrush

    San Francisco, CA
    13 days ago
  • $180k - $270k

     ...the world's most trusted AI work companion for...  ...high-throughput, ultra-low-latency inference engines for large language models or foundational speech models...  ...Token (or Time-To-First-Audio) in real-time streaming...  ...collaborative, fast-paced research. Gear & Perks: Choice... 
    Audio
    Full time
    Work at office
    Worldwide

    Plaud

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Audio AI Researcher: Multimodal, Low-Latency Modeling. Be the first to apply!