Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Research Scientist-Voice and Audio Ai

RST Recruitment

Job Description

Job Description

This is a research-driven, high-impact role for ML researchers who want to push the boundaries of real-time AI. As a Founding Machine Learning Research Engineer at Retell, you'll focus on advancing model capabilities for human-like voice agents operating in complex, real-world environments.

You'll explore new approaches across LLMs and audio models, design novel evaluation methods, and prototype systems that improve reasoning, latency, and conversational quality. Your work will directly influence production systems, bridging cutting-edge research with real-world deployment.

If you're excited about solving open-ended ML problems, experimenting rapidly, and shaping how voice AI systems think and perform, this is a unique opportunity to do so at scale.

KEY RESPONSIBILITIES

  • Research & Experimentation – Explore and develop new techniques across LLMs and audio models to improve reasoning, latency, and conversational quality in real-time systems.
  • Model Prototyping – Rapidly build and iterate on experimental models and pipelines, turning research ideas into working prototypes.
  • Evaluation & Benchmarking – Design novel evaluation frameworks, datasets, and metrics to measure performance on complex, real-world voice tasks.
  • Bridge Research to Production – Collaborate closely with engineering to translate research insights into deployable systems.
  • Human Feedback Loops – Develop methods to incorporate human evaluation into model improvement, especially for subjective conversational quality.
  • Advance the Frontier – Stay at the cutting edge of ML research and bring new ideas into Retell's product and infrastructure.

HOW TO THRIVE

  • Strong ML Research Background – You've worked on advanced ML problems (for example: LLM pre-training and post training, transcription model training, text to speech model training, or multimodal systems), either in industry or academia.
  • Deep Technical Foundation – Comfortable with PyTorch, model architectures, and the math behind modern machine learning.
  • Experimental Mindset – You enjoy exploring open-ended problems and iterating quickly on ideas.
  • Bridging Theory & Practice – You can translate research into systems that work in real-world environments.
  • Startup-Ready – You thrive in fast-paced environments with high ownership and ambiguity.
  • Collaborative & Clear Communicator – You can explain complex ideas and work cross-functionally to drive impact.

Tech stack:
PyTorch, LLMs, Audio/Speech Models, Text-to-Speech (TTS), Automatic Speech Recognition (ASR), Multimodal Systems, Python

Seniority:

1 - 5 years of experience in audio/multimodal AI research or engineering

Work experience:

Working at a high bar company with AI products (MAANG, vc backed startup, etc.)

Experience with LLM pre-training or post-training, evals, and translating research to production

Coming from another top voice ai or audio startup (Cartesia, Eleven Labs, Descript, etc.)

Education:

Degree in CS, ML, or closely related field (PhD prefered)

Recent publications in voice, audio, or speech AI

Hard skills:

Hands-on PyTorch and audio/speech model development

Experience with TTS, ASR, or multimodal audio systems

Pre-training experience at scale

Miscellaneous:

Comfortable with intense startup pace including weekend work

Vacancy posted 25 days ago
Similar jobs that could be interesting for youBased on the Research Scientist-Voice and Audio Ai in Redwood City, CA vacancy
  • $117.2k - $313.7k

     ...About the Role Salesforce AI Research is seeking outstanding AI Research Scientists / Research Engineers to build and deploy high...  ...GUI agents Speech Intelligence – Voice intelligence, TTS/ASR, human‑like turn‑taking, low‑latency audio processing Efficient Systems – Scalable... 
    Audio
    Full time

    100 Salesforce, Inc.

    Palo Alto, CA
    2 days ago
  • $117.2k - $313.7k

     ...Salesforce AI Research is looking for outstanding AI Research Scientists and Research Engineers to discover new research problems...  ...AI agents. Speech Intelligence: Voice intelligence, text‑to‑speech/...  ...like turn‑taking, and low‑latency audio processing. Core Modeling and... 
    Audio

    Salesforce.Com Inc

    Palo Alto, CA
    4 days ago
  • $34 per hour

     ...join our team as Data Labeling Analysts, supporting speech and voice AI systems. This is a high-impact production role focused on building...  ...that power real-world AI systems. You'll be working with audio, speech, and language data — helping ensure models are trained... 
    Audio
    Work experience placement
    Remote work

    Welocalize

    Burlingame, CA
    6 days ago
  •  ...Description Salesforce AI Research is looking for outstanding AI Research Scientists / Research Engineers. Do you want to...  ...agents. Speech Intelligence: Voice intelligence, TTS/ASR, human-like turn-taking, and low-latency audio processing. Efficient Systems:... 
    Audio
    Full time

    Salesforce

    Palo Alto, CA
    a month ago
  • $26 - $28 per hour

     ...join our team as Data Labeling Analysts, supporting speech and voice AI systems. This is a high-impact production role focused on building...  ...that power real-world AI systems. You'll be working with audio, speech, and language data — helping ensure models are trained... 
    Audio
    Full time
    Work experience placement
    Remote work
    Visa sponsorship

    Welocalize

    Burlingame, CA
    4 days ago
  • $200k - $280k

     ...ABOUT RETELL AI Retell AI is using the first principles to reimagine the call center with cutting edge voice AI. We believe voice is still the most natural way humans communicate...  ...’ll fine‑tune large language models and audio models, evaluate them with rigorous... 
    Audio
    H1b
    Work at office

    Retell AI

    Redwood City, CA
    7 hours ago
  • $180k

     ...SpaceXAI's mission is to create AI systems that can accurately...  ...ABOUT THE ROLE: The Grok Voice Product team builds seamless,...  ...code. Deep interest in voice/audio AI. Strong generalist engineer...  ...collaborating across engineering, research, and product teams to ship... 
    Audio
    Temporary work

    SpaceXAI

    Palo Alto, CA
    10 days ago
  • $26 - $28 per hour

     ...What if your language expertise could help improve the speech and voice AI systems used by millions of people worldwide? WHAT YOU’LL DO...  ...tasks across speech and voice datasets. • Work with audio and language data, including transcription, categorization, and... 
    Audio
    Hourly pay
    Full time
    Worldwide
    Visa sponsorship

    Welo Data

    Burlingame, CA
    a month ago
  • $150k

     ...Description SpaceXAI's mission is to create AI systems that can accurately understand the...  ...ABOUT THE ROLE: You will join the Grok Voice Model team to help build the world's best...  ...pipeline: massive data curation, premium audio processing, frontier speech-language pre-training... 
    Audio
    Temporary work

    SpaceXAI

    Palo Alto, CA
    9 days ago
  • $35 - $45 per hour

    xAI is seeking an AI Tutor specialized in multilingual audio to enhance Grok's voice interactions. The role involves curating and annotating audio data to improve speech recognition globally. Candidates should have native proficiency in Italian and be proficient in English... 
    Audio
    Hourly pay
    Remote work
    Flexible hours

    Xai

    Palo Alto, CA
    5 days ago
  • $26 - $28 per hour

     ...join our team as Data Labeling Analysts, supporting speech and voice AI systems. This is a high-impact production role focused on building...  ...that power real-world AI systems. You’ll be working with audio, speech, and language data — helping ensure models are trained... 
    Audio
    Full time
    Work experience placement
    Remote work
    Visa sponsorship

    Welo Data

    Burlingame, CA
    5 days ago
  •  ...date investors like a16z, General Catalyst, GV, and Accel and enjoy multi-year runway. About The Role We’re looking for an AI Research Scientist to advance the methodological frontier of AI in healthcare. This role is ideal for someone who has demonstrated strong research... 
    Temporary work
    Work at office
    Monday to Friday
    Monday to Thursday
    Flexible hours

    Sprinter Health

    Menlo Park, CA
    1 day ago
  •  ...generative models e.g. including language models, audio models, or video models) Have led or significantly...  ...physical world. The company embraces both large-scale AI and robotics as core to its DNA. Our team of researchers, roboticists, and company builders come from OpenAI... 
    Audio

    Generalist

    San Mateo, CA
    2 days ago
  •  ...About the Role Progress in speech AI is only as meaningful as our ability to measure...  ...-world disfluency. We’re looking for a Research Scientist who can define what "better" actually means...  ...applied research experience in speech, audio, or NLP, with a demonstrated focus on... 
    Audio

    Sanas

    Palo Alto, CA
    3 days ago
  •  ...Femtosense—was founded in 2018 by researchers from the Brains in Silicon Lab...  ...pioneered a high-performance AI accelerator integrated with an...  ...-based and classical DSP audio algorithms for our SPU platform...  ...detection, sound localization, voice identification, voice interfaces... 
    Audio
    Work experience placement

    Femtosense

    San Bruno, CA
    15 hours ago
  • $45 - $100 per hour

     ...Role Description Mercor is partnering with a leading AI research group to engage mathematics professionals in a...  ...enhance model performance. Contribute data in text, voice, and video formats, including annotations, audio recordings, or video sessions. Use proprietary... 
    Audio
    Hourly pay
    Contract work
    Work at office
    Local area
    Remote work

    CloudDevs

    Palo Alto, CA
    3 days ago
  • $25 per hour

     ...to leveraging different perspectives, voices, and backgrounds to promote a thoughtful...  ...approach to technology. Through research-driven and explorative AI, we aim to create real-time interactions that seamlessly blend text, audio, visuals and beyond. About the Role... 
    Audio
    Hourly pay
    Flexible hours

    Breakout Tools

    Mountain View, CA
    3 days ago
  • $170k - $190k

     ..., our mission is simple: deliver the best AI-powered customer experience—faster than anyone...  ...the design and delivery of end-to-end voice AI solutions, combining large language models...  ..., text-to-speech, and real-time streaming audio pipelines. This role requires a hands-on... 
    Audio
    Full time

    Asapp

    Mountain View, CA
    7 hours ago
  •  ...outside sales and service teams work. Their AI technology captures and analyzes real-...  ...natural speech interaction and real-time audio understanding. Develop and optimize ML...  ...critical insights from previously unstructured voice data. Build agents capable of... 
    Audio
    Full time

    Catalyst Labs, LLC

    Palo Alto, CA
    3 days ago
  • $25 per hour

     ...to leveraging different perspectives, voices, and backgrounds to promote a thoughtful...  ...approach to technology. Through research‑driven and explorative AI, we aim to create real‑time interactions that seamlessly blend text, audio, visuals, and beyond. About the Role:... 
    Audio
    Hourly pay
    Flexible hours

    Breakout Tools

    Mountain View, CA
    2 days ago
  •  ...electronics systems and semiconductors where AI can design and create beyond human...  ...team of previous Stanford professors, SAIL researchers, Olympiad medalists (IPhO, IOI, etc.), CTOs...  ...models (e.g., combining text, image, or audio inputs). Bonus Points Background... 
    Audio
    Full time

    Voltai

    Palo Alto, CA
    7 hours ago
  •  ...across the entire robotics stack. We're training state-of-the-art AI models that leverage our large-scale, high-quality, real-world...  ...more time on the things they value most. As a Machine Learning Research Engineer, you will work on the software and algorithms that... 

    Sunday

    Redwood City, CA
    4 days ago
  • $205k - $235k

     ...than ever before. Generative AI platforms represent a major breakthrough...  ...need to include their unique voice and style and ensure...  ...types including text, image, audio and video enabling us to generate...  ...company   ~ Publications and research experiences in ML venues and/... 
    Audio
    Full time
    Work at office
    Local area
    Flexible hours
    3 days per week

    Typeface

    Palo Alto, CA
    7 hours ago
  • $160k - $195k

     ...through integration, launch, and production across Syntiant’s edge AI, processor, sensor, and software solutions. The ideal...  ...background, with experience across processors, microcontrollers, sensors, audio, connectivity, or related semiconductor technologies. ~ Working... 
    Audio
    Temporary work
    Flexible hours

    Syntiant

    Redwood City, CA
    5 days ago
  • $197k - $291k

    # Software Engineer Manager II, Audio and Video, YouTubeGoogle • onsite • Mountain View, CA, USA • full\_timePay: USD 197000.00 - USD 29...  ...achieve as a team with sellers, shape the future of advertising in the AI-era, and make a real impact on the millions of companies and... 
    Audio
    Temporary work

    Epic Games (Portuguese)

    Mountain View, CA
    3 days ago
  •  ...AI Research Intern We're looking for an AI Research Intern to join our AI team and explore cutting-edge research across...  ...enhancement, super-resolution, restoration) Speech & audio (e.g. speech enhancement, voice cloning, voice generation) Multimodal understanding... 
    Audio
    Internship
    Local area
    Remote work
    Worldwide
    Flexible hours
    3 days per week

    OpusClip

    Mountain View, CA
    3 days ago
  •  ...A growing AI technology startup is seeking an ML Engineer to design and deploy production-grade systems. The role involves using Python and collaborating with teams to optimize customer interactions through advanced AI applications. Candidates should have a degree in... 
    Audio

    Catalyst Labs, LLC

    Mountain View, CA
    3 days ago
  •  ...Machine Learning Research Scientist At Autoscience Institute, we create AI systems that autonomously conduct AI research. Recently, we announced the first AI agent to autonomously create peer-reviewed literature (ICLR 2025 Workshops). We are passionate about pushing... 
    Full time
    Flexible hours

    Autoscience Institute

    Menlo Park, CA
    5 days ago
  • $250k - $350k

    About the Lab The 1X World Model Lab is an embodied AI research organization focused on pretraining the foundation models to accelerate the emergence of embodied intelligence. As the lab grows, researchers contribute where they have the most leverage, and the problems worth... 
    Local area

    1X

    San Carlos, CA
    5 days ago
  • $150k - $250k

     ...Body Language Engineer, AI Companion Location: Palo Alto, CA (on‑site) About 1X We build humanoid robots that work alongside...  ...years of experience training multi‑modal models combining language, audio, and vision ~ Experience in robotics — such as path planning... 
    Audio

    1x.tech

    Palo Alto, CA
    7 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Research Scientist-Voice and Audio Ai. Be the first to apply!