Research Scientist-Voice and Audio Ai
RST Recruitment
Job Description
Job Description
This is a research-driven, high-impact role for ML researchers who want to push the boundaries of real-time AI. As a Founding Machine Learning Research Engineer at Retell, you'll focus on advancing model capabilities for human-like voice agents operating in complex, real-world environments.
You'll explore new approaches across LLMs and audio models, design novel evaluation methods, and prototype systems that improve reasoning, latency, and conversational quality. Your work will directly influence production systems, bridging cutting-edge research with real-world deployment.
If you're excited about solving open-ended ML problems, experimenting rapidly, and shaping how voice AI systems think and perform, this is a unique opportunity to do so at scale.
KEY RESPONSIBILITIES
- Research & Experimentation – Explore and develop new techniques across LLMs and audio models to improve reasoning, latency, and conversational quality in real-time systems.
- Model Prototyping – Rapidly build and iterate on experimental models and pipelines, turning research ideas into working prototypes.
- Evaluation & Benchmarking – Design novel evaluation frameworks, datasets, and metrics to measure performance on complex, real-world voice tasks.
- Bridge Research to Production – Collaborate closely with engineering to translate research insights into deployable systems.
- Human Feedback Loops – Develop methods to incorporate human evaluation into model improvement, especially for subjective conversational quality.
- Advance the Frontier – Stay at the cutting edge of ML research and bring new ideas into Retell's product and infrastructure.
HOW TO THRIVE
- Strong ML Research Background – You've worked on advanced ML problems (for example: LLM pre-training and post training, transcription model training, text to speech model training, or multimodal systems), either in industry or academia.
- Deep Technical Foundation – Comfortable with PyTorch, model architectures, and the math behind modern machine learning.
- Experimental Mindset – You enjoy exploring open-ended problems and iterating quickly on ideas.
- Bridging Theory & Practice – You can translate research into systems that work in real-world environments.
- Startup-Ready – You thrive in fast-paced environments with high ownership and ambiguity.
- Collaborative & Clear Communicator – You can explain complex ideas and work cross-functionally to drive impact.
Tech stack:
PyTorch, LLMs, Audio/Speech Models, Text-to-Speech (TTS), Automatic Speech Recognition (ASR), Multimodal Systems, Python
Seniority:
1 - 5 years of experience in audio/multimodal AI research or engineering
Work experience:
Working at a high bar company with AI products (MAANG, vc backed startup, etc.)
Experience with LLM pre-training or post-training, evals, and translating research to production
Coming from another top voice ai or audio startup (Cartesia, Eleven Labs, Descript, etc.)
Education:
Degree in CS, ML, or closely related field (PhD prefered)
Recent publications in voice, audio, or speech AI
Hard skills:
Hands-on PyTorch and audio/speech model development
Experience with TTS, ASR, or multimodal audio systems
Pre-training experience at scale
Miscellaneous:
Comfortable with intense startup pace including weekend work
$117.2k - $313.7k
...About the Role Salesforce AI Research is seeking outstanding AI Research Scientists / Research Engineers to build and deploy high... ...GUI agents Speech Intelligence – Voice intelligence, TTS/ASR, human‑like turn‑taking, low‑latency audio processing Efficient Systems – Scalable...AudioFull time$117.2k - $313.7k
...Salesforce AI Research is looking for outstanding AI Research Scientists and Research Engineers to discover new research problems... ...AI agents. Speech Intelligence: Voice intelligence, text‑to‑speech/... ...like turn‑taking, and low‑latency audio processing. Core Modeling and...Audio$34 per hour
...join our team as Data Labeling Analysts, supporting speech and voice AI systems. This is a high-impact production role focused on building... ...that power real-world AI systems. You'll be working with audio, speech, and language data — helping ensure models are trained...AudioWork experience placementRemote work- ...Description Salesforce AI Research is looking for outstanding AI Research Scientists / Research Engineers. Do you want to... ...agents. Speech Intelligence: Voice intelligence, TTS/ASR, human-like turn-taking, and low-latency audio processing. Efficient Systems:...AudioFull time
$26 - $28 per hour
...join our team as Data Labeling Analysts, supporting speech and voice AI systems. This is a high-impact production role focused on building... ...that power real-world AI systems. You'll be working with audio, speech, and language data — helping ensure models are trained...AudioFull timeWork experience placementRemote workVisa sponsorship$200k - $280k
...ABOUT RETELL AI Retell AI is using the first principles to reimagine the call center with cutting edge voice AI. We believe voice is still the most natural way humans communicate... ...’ll fine‑tune large language models and audio models, evaluate them with rigorous...AudioH1bWork at office$180k
...SpaceXAI's mission is to create AI systems that can accurately... ...ABOUT THE ROLE: The Grok Voice Product team builds seamless,... ...code. Deep interest in voice/audio AI. Strong generalist engineer... ...collaborating across engineering, research, and product teams to ship...AudioTemporary work$26 - $28 per hour
...What if your language expertise could help improve the speech and voice AI systems used by millions of people worldwide? WHAT YOU’LL DO... ...tasks across speech and voice datasets. • Work with audio and language data, including transcription, categorization, and...AudioHourly payFull timeWorldwideVisa sponsorship$150k
...Description SpaceXAI's mission is to create AI systems that can accurately understand the... ...ABOUT THE ROLE: You will join the Grok Voice Model team to help build the world's best... ...pipeline: massive data curation, premium audio processing, frontier speech-language pre-training...AudioTemporary work$35 - $45 per hour
xAI is seeking an AI Tutor specialized in multilingual audio to enhance Grok's voice interactions. The role involves curating and annotating audio data to improve speech recognition globally. Candidates should have native proficiency in Italian and be proficient in English...AudioHourly payRemote workFlexible hours$26 - $28 per hour
...join our team as Data Labeling Analysts, supporting speech and voice AI systems. This is a high-impact production role focused on building... ...that power real-world AI systems. You’ll be working with audio, speech, and language data — helping ensure models are trained...AudioFull timeWork experience placementRemote workVisa sponsorship- ...date investors like a16z, General Catalyst, GV, and Accel and enjoy multi-year runway. About The Role We’re looking for an AI Research Scientist to advance the methodological frontier of AI in healthcare. This role is ideal for someone who has demonstrated strong research...Temporary workWork at officeMonday to FridayMonday to ThursdayFlexible hours
- ...generative models e.g. including language models, audio models, or video models) Have led or significantly... ...physical world. The company embraces both large-scale AI and robotics as core to its DNA. Our team of researchers, roboticists, and company builders come from OpenAI...Audio
- ...About the Role Progress in speech AI is only as meaningful as our ability to measure... ...-world disfluency. We’re looking for a Research Scientist who can define what "better" actually means... ...applied research experience in speech, audio, or NLP, with a demonstrated focus on...Audio
- ...Femtosense—was founded in 2018 by researchers from the Brains in Silicon Lab... ...pioneered a high-performance AI accelerator integrated with an... ...-based and classical DSP audio algorithms for our SPU platform... ...detection, sound localization, voice identification, voice interfaces...AudioWork experience placement
$45 - $100 per hour
...Role Description Mercor is partnering with a leading AI research group to engage mathematics professionals in a... ...enhance model performance. Contribute data in text, voice, and video formats, including annotations, audio recordings, or video sessions. Use proprietary...AudioHourly payContract workWork at officeLocal areaRemote work$25 per hour
...to leveraging different perspectives, voices, and backgrounds to promote a thoughtful... ...approach to technology. Through research-driven and explorative AI, we aim to create real-time interactions that seamlessly blend text, audio, visuals and beyond. About the Role...AudioHourly payFlexible hours$170k - $190k
..., our mission is simple: deliver the best AI-powered customer experience—faster than anyone... ...the design and delivery of end-to-end voice AI solutions, combining large language models... ..., text-to-speech, and real-time streaming audio pipelines. This role requires a hands-on...AudioFull time- ...outside sales and service teams work. Their AI technology captures and analyzes real-... ...natural speech interaction and real-time audio understanding. Develop and optimize ML... ...critical insights from previously unstructured voice data. Build agents capable of...AudioFull time
$25 per hour
...to leveraging different perspectives, voices, and backgrounds to promote a thoughtful... ...approach to technology. Through research‑driven and explorative AI, we aim to create real‑time interactions that seamlessly blend text, audio, visuals, and beyond. About the Role:...AudioHourly payFlexible hours- ...electronics systems and semiconductors where AI can design and create beyond human... ...team of previous Stanford professors, SAIL researchers, Olympiad medalists (IPhO, IOI, etc.), CTOs... ...models (e.g., combining text, image, or audio inputs). Bonus Points Background...AudioFull time
- ...across the entire robotics stack. We're training state-of-the-art AI models that leverage our large-scale, high-quality, real-world... ...more time on the things they value most. As a Machine Learning Research Engineer, you will work on the software and algorithms that...
$205k - $235k
...than ever before. Generative AI platforms represent a major breakthrough... ...need to include their unique voice and style and ensure... ...types including text, image, audio and video enabling us to generate... ...company ~ Publications and research experiences in ML venues and/...AudioFull timeWork at officeLocal areaFlexible hours3 days per week$160k - $195k
...through integration, launch, and production across Syntiant’s edge AI, processor, sensor, and software solutions. The ideal... ...background, with experience across processors, microcontrollers, sensors, audio, connectivity, or related semiconductor technologies. ~ Working...AudioTemporary workFlexible hours$197k - $291k
# Software Engineer Manager II, Audio and Video, YouTubeGoogle • onsite • Mountain View, CA, USA • full\_timePay: USD 197000.00 - USD 29... ...achieve as a team with sellers, shape the future of advertising in the AI-era, and make a real impact on the millions of companies and...AudioTemporary work- ...AI Research Intern We're looking for an AI Research Intern to join our AI team and explore cutting-edge research across... ...enhancement, super-resolution, restoration) Speech & audio (e.g. speech enhancement, voice cloning, voice generation) Multimodal understanding...AudioInternshipLocal areaRemote workWorldwideFlexible hours3 days per week
- ...A growing AI technology startup is seeking an ML Engineer to design and deploy production-grade systems. The role involves using Python and collaborating with teams to optimize customer interactions through advanced AI applications. Candidates should have a degree in...Audio
- ...Machine Learning Research Scientist At Autoscience Institute, we create AI systems that autonomously conduct AI research. Recently, we announced the first AI agent to autonomously create peer-reviewed literature (ICLR 2025 Workshops). We are passionate about pushing...Full timeFlexible hours
$250k - $350k
About the Lab The 1X World Model Lab is an embodied AI research organization focused on pretraining the foundation models to accelerate the emergence of embodied intelligence. As the lab grows, researchers contribute where they have the most leverage, and the problems worth...Local area$150k - $250k
...Body Language Engineer, AI Companion Location: Palo Alto, CA (on‑site) About 1X We build humanoid robots that work alongside... ...years of experience training multi‑modal models combining language, audio, and vision ~ Experience in robotics — such as path planning...Audio
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Research Scientist-Voice and Audio Ai. Be the first to apply!
- drug safety scientist Redwood City, CA
- molecular biology scientist Redwood City, CA
- safety scientist Redwood City, CA
- validation scientist Redwood City, CA
- support scientist Redwood City, CA
- water quality scientist Redwood City, CA
- machine learning research scientist Redwood City, CA
- manufacturing scientist Redwood City, CA
- senior principal scientist Redwood City, CA
- research scientist - biology Redwood City, CA



