Research Engineer, Audio and Speech
$200k - $400kDecagon
About Decagon
Decagon is the leading conversational AI platform empowering every brand to deliver concierge customer experiences.
Our technology enables industry-defining enterprises like Avis Budget Group, Block’s Cash App and Square, Chime, Oura Health, and Hunter Douglas to deploy AI agents that power personalized, deeply satisfying interactions across voice, chat, email, SMS, and every other channel.
We’re building a future where customer experiences are being redefined from support tickets and hold music to faster resolutions, richer conversations, and deeper relationships. We’re proud to be backed by world-class investors who share that vision, including a16z, Accel, Bain Capital Ventures, Coatue, and Index Ventures, along with many others.
We’re an in-office company, driven by a shared commitment to excellence and velocity. Our values — Just Get It Done, Invent What Customers Want, Winner’s Mindset, and The Polymath Principle — shape how we work and grow as a team.
About the Team
Read more about the Speech Research Team's work:
The Research team develops the model and decision-making stack that powers Decagon’s conversational agents for enterprise support. We research, adapt, and implement state-of-the-art techniques in model training, prompting, orchestration, and evaluation in order to make our agents more accurate, robust, and efficient in real-world deployments.
Our goal is to push the frontier of applied conversational AI: agents that reliably understand nuanced intent, track long context, and take the right actions under uncertainty. We measure success the way customers feel it: higher resolution rates, better user satisfaction, and consistent behavior at scale.
About the Role
As a Research Engineer focused on Audio and Speech, you’ll be responsible for building the models and agent harnesses that power Decagon’s real-time voice agents and taking them all the way from idea to production. Your work will advance multimodal and full-duplex systems that can listen, reason, speak, and respond naturally in real time.
We’re looking for strong engineers who want to build the next generation of AI voice agents. People here own their work end-to-end, ship real improvements, and are trusted to make high-impact technical decisions.
In this role, you will
Design and build next-generation agent harnesses optimized for streaming speech, turn-taking, interruptions, overlapping speech, and continuous interaction
Research and train multimodal and full-duplex models that jointly understand audio, reason, and generate speech
Improve speech recognition, voice activity detection, endpointing, and speech generation across diverse speakers, environments, domains, and languages
Build evaluations and use production calls to ship measurable improvements in accuracy, latency, naturalness, and task outcomes
Optimize end-to-end inference for responsiveness, throughput, stability, and cost, partnering with Voice Platform and Infrastructure teams to deploy at scale
Your background looks something like this
2+ years of experience in speech, audio ML, multimodal ML, or production machine learning
Experience developing or adapting autoregressive, diffusion, flow-matching, or codec-based speech models
Hands-on experience with streaming agent systems, low-latency inference, production model serving, and evaluation on real-world audio
Fluency in Python and a modern deep-learning framework such as PyTorch, with strong foundations in machine learning and signal processing
A track record of taking research ideas from prototype to reliable, measurable production impact
Even better if you have
Familiarity with speech-to-speech or full-duplex models
Experience with telephony, multilingual speech, noisy-channel robustness, speaker adaptation, or expressive speech generation
Compensation
$200K – $400K + Offers Equity
Benefits
We proudly offer the following benefits for our full-time employees:
Medical, Dental, and Vision benefits for you and your family
Life Insurance and Disability Benefits
Retirement Plan (e.g., 401K, pension)
Parental Leave
Fertility and family building benefits through Carrot
Monthly stipend to support your wellness, lifestyle, and work-life balance
Daily lunches and snacks in the office to keep you at your best
Take what you need vacation policy (subject to local requirements; UK employees receive 25 days of statutory leave)
These benefits are described in more detail in Decagon’s policies, may vary by location, and can change at any time according to applicable compensation and benefits plans.
- ...Job Description Zyphra is an artificial intelligence company based in San Francisco, California. The Role: As a Research Engineer - Audio & Speech Models , you will be a core contributor on Zyphra’s Audio Team, building the next generation of open-source...AudioWork at officeRelocation package
- AI Talent Now, LLC in San Francisco, CA, is seeking a Senior Machine Learning Research Engineer to own the full lifecycle of ML research and development for speech and audio models. You will design, train, deploy, and fine-tune state-of-the-art systems with end-to-end ownership...AudioWork at office
- ...Research EngineerWe are looking for a Research Engineer to join the research team at ElevenLabs. You will thrive in the role... ...system specialized for text-to-speech projects. This includes establishing... ...particularly within the realm of audio and text-to-speech domains....AudioRemote work
- ...About PhonicPhonic is a product and research lab focused on powering the most realistic... ...tier 1 VCs.About The RoleAs a Research Engineer at Phonic, you'll sit at the intersection... ...on models across the voice stack (speech, audio, language, and beyond), while also building...AudioWork at office
- ...ElevenlabsElevenLabs is an AI research and product company... ...marketers to generate and edit speech, music, image, and video across... ...developers access to our leading AI audio foundational models.... ...their lives. We are researchers, engineers, and operators. IOI medalists...AudioImmediate startRemote work
- Senior Machine Learning Research Engineer San Francisco, United States | Posted on 09/03/2026 AI... ...research and development at the frontier of audio AI. You'll join a fast-growing team... ...-tier investors, building cutting-edge speech and audio models that power the data...AudioWork at officeRelocation
$164.6k - $313.3k
...Design AI group (SODA) is looking for a driven Data/ML engineer to push the boundaries of audio GenAI. Join the team behind Firefly Generate Sound Effects... ...products.We’re a small, collaborative and efficient research team looking for highly motivated candidates of all levels...AudioFull timeTemporary workLocal areaWorldwide- ...AI Research Engineer OpportunityPoly is building a better file storage platform for everyone. We've raised $8M from YC, Bloomberg Beta, Felicis... ...to give an answer.We're maximally multimodal. We support any audio, video, document, office file, slide deck, text file, PDF,...AudioWork at office
- A technology company in San Francisco is looking for an innovative Applied Research Engineer to develop high-performance solutions in computer vision and audio processing. You will collaborate with cross-functional teams, conduct applied research, and lead experimental...Audio
- Huxley in San Francisco is seeking an Edge AI Research Scientist to advance speech and audio AI for devices with limited resources. You will design compact... ..., combining research rigor with practical system engineering to deliver scalable, production-ready solutions. #J...Audio
- Applied Research Engineer - San Francisco As an applied research engineer at Confidential, you’ll build high performance building blocks and... ...techniques to solve them. You will be working in the computer vision, audio processing, and text processing domains. You’re likely a good...Audio
- Plaud Inc. is seeking senior AI researchers to join our SpeechLLM lab in San Francisco. You will help build and train large-scale audio/speech models and push the boundaries of human-AI interaction. We value hands-on experience with PyTorch or JAX, distributed training...Audio
$180k - $270k
...throughput, ultra-low-latency inference engines for large language models or foundational speech models. Understand the... ...To-First-Token (or Time-To-First-Audio) in real-time streaming environments... ...highly collaborative, fast-paced research. Gear & Perks: Choice of top-...AudioFull timeWork at officeWorldwide- ...an inflection point where advances in speech, language models, and clinical AI can... ...Position Overview We are hiring two ML Engineers / Researchers to help build the next generation of... ...in one of two areas: Speech & Audio: Build state-of-the-art medical speech...AudioFull time
- .... This full-time role focuses on building end-to-end speech technologies across TTS, STT, and neural codecs, with... ...You will push theory to production, train on massive audio datasets, and collaborate with engineering and product teams to ship capabilities to customers quickly...AudioFull time
- We are seeking an Edge AI Research Scientist to develop next-generation speech and audio AI systems that run efficiently on smartphones, wearables, and other resource... ...research, systems optimization, and production engineering. You will design compact model architectures,...Audio
$110.7k - $379.2k
Position Summary Research Engineer — Post-Training & Small Language Models (SLMs), Healthcare AI Three hundred fifty million Americans... ...-LLM, TGI, or Ollama. • Experience with multimodal models, speech models, or domain-specific foundation models; experience using...Local areaVisa sponsorship- ...immediately. No speculative research track here. If you want your... ...single, ambitious goal: a fully speech-to-speech conversational AI... ...text, text-to-speech, neural audio codecs, and getting LLMs to understand... ...a day, working closely with engineering and product teams to get your...AudioPermanent employmentFull timeImmediate start
- ...hiring for a Product Manager focused on AI speech (text-to-speech, speech-to-text, voice... ...development, collaborating closely with engineering and analysis colleagues, as well as our... ...Experience working in a product role at a voice/audio software company (e.g. ElevenLabs,...Audio
$34 per hour
...individuals to join our team as Data Labeling Analysts, supporting speech and voice AI systems. This is a high-impact production role... ...that power real-world AI systems. You'll be working with audio, speech, and language data — helping ensure models are trained on...AudioWork experience placementRemote work$26 - $28 per hour
...individuals to join our team as Data Labeling Analysts, supporting speech and voice AI systems. This is a high-impact production role... ...that power real-world AI systems. You'll be working with audio, speech, and language data — helping ensure models are trained on...AudioWork experience placementRemote work- ...Research EngineerLotus Health is a groundbreaking primary care app that integrates your... ...prescriptions.Our team includes ex-founders and engineers who have built and scaled consumer apps... ...structured tool useExperience building speech or multimodal pipelines for medical...
$195k - $365k
...proven track record of building and training large-scale audio or speech models from the ground up, whether that involves unified... ...audio architectures. Love living at the intersection of research and engineering, eager to design novel sequence modeling architectures one...AudioFull timeWork at officeWorldwide- ...automate offline and live evals that keep our speech and multimodal models honest in... ...models, not slide decks — partner with research and infra to prototype, train, and deploy... ...Qualifications:Expert-level PyTorch.Proven software engineer who loves ML; comfortable writing...Full timeContract workShift work
$224.5k - $256k
...Empathetic . Your Role As a Sr. AI Engineer: Speech, you'll be a senior technical leader on... ...speech recognition, enhancement, audio intelligence, and real-time inference.... ...engineering: whether you are strongest in research, systems, or both, you'll help turn advances...AudioWork at officeRemote work$180k - $240k
...inference calls monthly, process 1M+ hours of audio daily, and power 2 billion+ end-user... ...role: We're looking for a Senior Design Engineer to own the craft and feel of AssemblyAI's... ...before applying. Keep Exploring AssemblyAI: Speech-to-text Streaming speech-to-text Speech...AudioLive in$200k - $225k
About AlembicAlembic is where top engineers are solving marketing's hardest problem: proving what actually works. If you're looking for... ...teams and customers to leverage advanced analyticsDocument research and implementation decisions for reproducibility and knowledge...$197.3k - $313.7k
...TeamSalesforce AI is looking for talented software and platform engineers to embed in our AI team to bridge the gap between frontier AI... ...where your engineering skills directly enable world-class research and products used by millions?At Salesforce, we are driving the...Full time- Factory is seeking innovative Research Engineers to design and integrate advanced AI and ML capabilities that revolutionize productivity and accelerate innovation within software organizations.What you will do and achieve:Design, develop, and deploy AI-driven agentic systems...Work at office
$150k - $250k
...manufacturing, consumer goods, and global social organizations.We research and deploy technologies that power AI-native operations — both... ...of 20+ F500s. What We Are Looking ForAt Distyl, Research Engineers build the bridge between frontier AI research and production systems...Work at office3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Research Engineer, Audio and Speech. Be the first to apply!
- research programmer San Francisco, CA
- senior research engineer San Francisco, CA
- junior machine learning research engineer San Francisco, CA
- deep learning research engineer San Francisco, CA
- research engineer San Francisco, CA
- ai research engineer San Francisco, CA
- research assistant engineering San Francisco, CA
- research software engineer San Francisco, CA
- audio video San Francisco, CA
- live audio San Francisco, CA



