Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Research Scientist, Speech & Audio

$160k - $185k
Full-time

Innodata Inc.

Innodata (Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale. We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.

Scope of the Role:

Where models actually differ now is robustness across accents, noise, and code-switching; speaker diarization; the naturalness of generated speech; and latency under streaming. Measuring those honestly, and building the data that trains for them, is gated as much by data and evaluation design as by architecture. Innodata builds that data and those evaluations for the customers and frontier labs advancing speech and audio models, and we are hiring a Research Scientist to own the science behind it.

You will partner directly with the customers and frontier labs building ASR, text-to-speech, speech-to-speech and conversational voice, diarization, and audio-language models, as interested in the data behind them as in the models themselves. Your work is judgment: which conditions and languages a benchmark must cover to be honest, what a transcription convention should be for a given objective, and when an automated metric can be trusted versus when a human ear is required. You will also partner closely with our transcription and linguistics lead, whose standards directly shape what the models learn.

What You’ll Own:

  • You will define how Innodata designs, structures, and evaluates audio data for speech and audio models, and you will validate those choices experimentally. Concretely, you will:
  • Translate the requirements of speech and audio models — ASR, text-to-speech and speech generation, speech-to-speech and conversational voice, speaker diarization and verification, audio-language models, and streaming systems — into concrete data specifications: modalities, transcription and annotation schemas, sampling, and evaluation criteria.
  • Build evaluation methodology that goes past word error rate — semantic accuracy, robustness to noise and accent, code-switching, diarization error rate (DER), naturalness and intelligibility of generated speech, and streaming latency — and know when automated metrics hold and when they don't.
  • Decide how existing and incoming audio should be structured, enriched, and sampled for coverage that fits the model objective — across languages, accents, and acoustic conditions (studio, real-world, telephonic), speaker demographics, emotional and paralinguistic range, scripted versus spontaneous speech, and single- versus multi-speaker settings, including low-resource and code-switched speech.
  • Partner with the transcription and linguistics lead to turn model objectives into transcription specifications, and to quantify how transcription conventions and quality move ASR and speech-model results.
  • Partner with the audio solutions and engineering team so the audio we collect is built for the model objective: you specify what good data and evaluation require, and they scope programs with customers and capture audio to spec.
  • Run experiments that prove data decisions matter: fine-tune and evaluate models on Innodata data, with ablations tying specific data choices to measurable improvement.
  • Design adversarial and stumping evaluations — noisy, accented, and adversarial audio — that surface where speech systems fail, and turn those failures into better data.
  • Publish. Turn what you learn into benchmarks, methodology, and papers that advance the field and earn the trust of the customers and frontier labs we partner with.
  • Work with annotation teams, subject-matter experts, and the synthetic- and augmented-audio pipeline to turn specifications into operational plans.

You’ll Thrive in This Role If You Have:

  • Roughly 5+ years of hands-on industry experience in speech or audio ML. We weight practical experience over formal credentials; a PhD with a compelling, current research agenda can offset the lower end.
  • A Bachelor's degree in computer science, electrical engineering, or a related technical or quantitative field is required; an advanced degree (MS or PhD) in a relevant field is preferred.
  • Trained and evaluated speech or audio models yourself — ASR, TTS, speech-to-speech, speaker, or audio-language models — with strong PyTorch fundamentals.
  • Fluency in the toolchains and metrics speech work runs on: ESPnet, NeMo, SpeechBrain, or Kaldi, HuggingFace, forced alignment, and WER/CER and the metrics that go beyond them.
  • Hands-on experience with multilingual, accented, dialectal, low-resource, or code-switched speech, and with synthetic or augmented audio (TTS pipelines, noise and room-response simulation).
  • A way of thinking in datasets: you have built evaluation sets, reasoned about coverage across conditions, and argued about what makes speech data good for a given objective.
  • A track record the field recognizes: first-author publications or strong open-source contributions at venues such as Interspeech, ICASSP, ASRU, SLT, or NeurIPS.
  • The ability to work directly with the research scientists at the customers and frontier labs we partner with, and to explain data and modeling decisions clearly to both expert and non-expert audiences, backed by a rigorous, reproducible approach to experiments and documentation.
  • Bonus: interest or hands-on experience in responsible-AI evaluation and red-teaming — for example spoofing and voice-cloning robustness, or bias across accents and languages.

The expected salary range for this position is $160,000 - $185,000 p/year, based on experience, skills, and qualifications.

Please be aware of recruitment scams involving individuals or organizations falsely claiming to represent employers. Innodata will never ask for payment, banking details, or sensitive personal information during the application process. To learn more on how to recognize job scams, please visit the Federal Trade Commission’s guide at

If you believe you’ve been targeted by a recruitment scam, please report it to Innodata at View email address on aiapply.co and consider reporting it to the FTC at ReportFraud.ftc.gov .

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Research Scientist, Speech & Audio in Remote vacancy
  •  ...built. We're looking for an experienced researcher to be a driving force behind this work: shaping...  ...multimodal models (vision, video, audio, language) — SFT, RLHF, DPO or GRPO, and...  ...tooling, and quality/coverage measurement Speech and audio models — streaming ASR and diarization... 
    Audio
    Work visa

    Noösphere

    Seattle, WA
    2 days ago
  •  ...Descript's Research team builds the models behind the product's most distinctive features: Video Regenerate...  ...work with. Some recent work from the team: - Audio editing by latent inpainting : regenerating a masked span of speech - Video Regenerate : regenerating a speaker's... 
    Audio
    Full time

    Descript

    Remote
    18 days ago
  •  ...The ideal candidate will have hands-on experience with the Web Speech API and familiarity with major speech frameworks. This remote contract...  ...for individuals with over 3 years of experience in voice UI or audio processing, focused on optimizing performance and user... 
    Audio
    Contract work
    Remote work

    New York Technology Partners

    United States
    3 days ago
  • $35 - $45 per hour

     ...technology firm is seeking an AI Tutor specialized in multilingual audio capabilities. This position focuses on training Grok to excel in...  ...based on experience. Remote work is possible, aiming to bridge language barriers and improve AI's speech processing. #J-18808-Ljbffr... 
    Audio
    Hourly pay
    Full time
    Part time
    Remote work

    Pantera Capital

    New York, NY
    4 days ago
  •  ...Machine Learning Scientist Rime builds voice AI for enterprises...  ...experiences at scale. Our text-to-speech models are purpose-built for...  ...the intersection of product, research, and craft. Building voice...  ...experience in speech, audio, ML, or computational linguistics... 
    Audio
    Remote work
    Visa sponsorship

    Rime Labs

    United States
    2 days ago
  • $45 per hour

     ...Commitment: 2–10 hours per week for 1–2 weeks Responsibilities Record high-quality Hebrew audio recordings to support AI model training Deliver clear, natural speech with a neutral-to-upbeat “customer service” tone Follow transcripts precisely or adjust them... 
    Audio
    Hourly pay
    Contract work
    Remote work
    10 hours per week

    Crossing Hurdles

    United States
    4 days ago
  • SpaceXAI is seeking an AI Tutor specialized in multilingual audio capabilities to train Grok for voice interactions and speech processing across languages and accents. You will curate and annotate high-quality audio data to improve global accessibility and natural spoken... 
    Audio
    Remote job
    Worldwide

    aitrainer

    Brooklyn, NY
    5 days ago
  • $39.5 per hour

     ...Audit Spanish-language AI training data by carefully comparing speech recordings with transcripts and word-level timing. This role focuses...  ...decision, and writing a concise rationale. Audit word-level audio alignments by confirming that each segment’s start and end... 
    Audio
    Hourly pay
    Remote work

    SaidGig

    United States
    9 days ago
  • $39.5 per hour

     ...Overview Evaluate Italian-language AI training data through detailed audio and transcription audits. This role focuses on careful, evidence...  ...multilingual transcription tasks by listening to human or agent speech, comparing it with an annotator''s transcript, applying the... 
    Audio
    Hourly pay
    Remote work

    SaidGig

    United States
    9 days ago
  • $160k - $185k

     ...long-form understanding, grounding events in time, and holding audio, video, and text together do not fall out of image benchmarks...  ...advancing video and multimodal models, and we are hiring a Research Scientist to own the science behind it. You will partner directly... 
    Audio
    Full time
    Fixed term contract

    Innodata Inc.

    Remote
    4 days ago
  • $231.5k - $405.1k

     ...Description About the team  Our Core AI Research team develops novel methods for...  ...About the role  As a Staff Research Scientist, you will independently lead a major...  ...language, documents, images/video, and speech/audio - and across multilingual or cross-lingual... 
    Audio
    Full time
    Work at office
    Immediate start
    Remote work
    Flexible hours
    Shift work

    ServiceNow

    Santa Clara, CA
    27 days ago
  •  ...Overview We're looking for audio engineers with strong experience editing and quality-checking voice recordings for speech and text-to-speech (TTS) applications , specifically for French-language audio . This role focuses on preparing high-quality audio datasets... 
    Audio

    Mercor

    Remote
    17 days ago
  •  ..., and testing of ML solutions using online code repositories, research publications, or customer specifications. Staying current with...  ...: Image Processing/Computer Vision, ADAS, Anomaly Detection, Audio/Speech Processing, Automatic Speech Recognition, and Time Series Modeling... 
    Audio

    Xforia Inc

    Laguna Woods, CA
    1 day ago
  •  ...development of Spotify's state-of-the-art speech models, contributing to speech...  ...processing and model serving, and capturing audio of outstanding quality from our voice talent...  .... We're looking for a senior applied research scientist with experience in developing novel ML... 
    Audio
    Full time

    Spotify

    Remote
    7 days ago
  • $26 - $28 per hour

     ...individuals to join our team as Data Labeling Analysts, supporting speech and voice AI systems. This is a high-impact production role...  ...datasets that power real-world AI systems. You’ll be working with audio, speech, and language data — helping ensure models are trained on... 
    Audio
    Full time
    Remote work
    Visa sponsorship

    Welo Data

    New York, NY
    3 days ago
  • $26 - $28 per hour

     ...individuals to join our team as Data Labeling Analysts, supporting speech and voice AI systems. This is a high-impact production role...  ...that power real-world AI systems. You’ll be working with audio, speech, and language data — helping ensure models are trained on... 
    Audio
    Full time
    Work experience placement
    Remote work
    Visa sponsorship

    Welocalize

    Washington DC
    5 days ago
  •  ...Workforce Engagement Management (WEM) Application Consultant (Verint Speech Analytics) to help our clients unlock actionable insights from...  ...trusted advisor to contact center leaders, translating complex audio and text data into strategies that improve customer experience,... 
    Audio
    Full time
    Work at office
    Remote work
    Worldwide
    3 days per week

    Five9

    Remote
    18 days ago
  • $224.5k - $256k

     ...Your Role As a Sr. AI Engineer: Speech, you’ll be a senior technical leader on our...  ...across speech recognition, enhancement, audio intelligence, and real-time inference. You...  ...engineering: whether you are strongest in research, systems, or both, you’ll help turn advances... 
    Audio
    Work at office
    Remote work

    Software Engineering, Data Science

    California
    23 days ago
  • $40 - $43 per hour

     ...the way the world experiences healthcare? Look no further, the Speech Language Pathologist is a key member of our team, who provides support...  ...and health related incidents. Alertness to respond to audio and visual cues from participants and their families, other staff... 
    Audio
    Part time
    Casual work
    Work at office
    Local area
    Immediate start
    Work from home
    3 days per week

    catalight

    Hilo, HI
    4 days ago
  •  ...industry-leading services include game development, art production, audio production, quality assurance, localization, localization QA,...  ...Manager who will lead high-profile relationships within the speech and AI ecosystem. This role combines technical fluency, partnership... 
    Audio
    Full time
    Remote work
    Flexible hours

    Side

    Remote
    9 days ago
  • $160k - $185k

     ...outstanding outcomes for our customers. Scope of the Role: As speech and audio models get better, the human role gets harder, not easier...  ...is the applied-expert counterpart to our Speech & Audio Research Scientist. You set the standards the models are trained and measured... 
    Audio
    Full time
    Fixed term contract
    Immediate start

    Innodata Inc.

    Remote
    4 days ago
  • $28.14 - $36.02 per hour

     ...Under the supervision of a credentialed Speech-Language Specialist/Pathologist, assist in...  ...activities such as picture cards, worksheets and audio equipment. Assist with the development...  ...Assist with speech-language pathology research projects, in-service training, and family... 
    Audio
    Hourly pay
    Permanent employment
    Full time
    Contract work
    Apprenticeship
    Internship
    Work at office
    Monday to Friday

    Bassett Unified School District

    La Puente, CA
    1 day ago
  • $49.3 per hour

     ...outstanding opportunity for Speech Pathologist. WORK SCHEDULE...  ...Uses equipment such a digital audio recorders, software/hardware-...  ...needs Assists with research and development. REQUIREMENTS...  ...preparing tomorrow's physicians, scientists and other health... 
    Audio
    Hourly pay
    16 hours
    Full time
    Temporary work
    Part time
    Work at office
    Remote work
    Shift work
    Day shift

    University of Washington

    Seattle, WA
    7 days ago
  • $25 - $30 per hour

     ...media and database tools to conduct initial research and gather intel on subjects....  ...Obtain videotaped documentation, photos, and audio recordings as part of thorough surveillance...  ...with arrest and conviction records. #J-18808-Ljbffr SPEECH EXCHANGE AND LANGUAGE THERAP
    Audio
    Hourly pay
    Part time
    Remote work
    Night shift
    Weekend work

    SPEECH EXCHANGE AND LANGUAGE THERAP

    Victorville, CA
    5 days ago
  • $35 - $45 per hour

     ...A technology firm in the United States is seeking candidates to manage multilingual audio data for AI projects. The ideal candidate must be a native Danish speaker with strong English skills, and the ability to transcribe and annotate audio accurately. The position offers... 
    Audio
    Hourly pay
    Remote work
    Flexible hours

    Pantera Capital

    United States
    4 days ago
  • $50 per hour

    Gridnaut Recruiting is hiring a remote Audio Engineer (Speech / TTS Audio Specialist) - French contractor (pay $50/hr). Contribute to frontier AI research and evaluation work. Ideal candidates: Remote contractor engagement supporting a leading AI lab.; Hands-on technical... 
    Audio
    Hourly pay
    For contractors
    Remote work

    Gridnaut Recruiting

    Remote
    18 days ago
  • $26 - $28 per hour

     ...individuals to join our team as Data Labeling Analysts, supporting speech and voice AI systems. This is a high-impact production role...  ...that power real-world AI systems. You'll be working with audio, speech, and language data - helping ensure models are trained on... 
    Audio
    Full time
    Work experience placement
    Remote work
    Visa sponsorship

    Welocalize

    Sunnyvale, CA
    5 days ago
  •  ...versions of AI Assistants. The Applied Scientist (AS) will help us[KM7] [LY8] [YW9] develop...  ...e.g., statistics, predictive analytics, research). Experience working on successful...  ...Foundation Models Application of Vision, Audio, and Multimodal Foundation Models... 
    Audio
    Remote work

    Yochana

    United States
    3 days ago
  •  ...speaking about American history. Ability to engage guests throughout each cruise. Sense of urgency in all guest, crew, and home office requests. Positive attitude and receptive to continuous performance feedback. Basic knowledge of audio/visual equipment. #J-18808-Ljbffr... 
    Audio
    Casual work
    Home office
    Afternoon shift

    American Cruise Lines

    Rockland, ME
    1 day ago
  • $65k

     ...Grow and expand existing relationships, while also providing research to sales team to help grow their perspective verticals. Strategize...  ...that specializes in direct response advertisements across TV, audio, digital and direct mail. With recent acquisitions of a leading... 
    Audio
    Remote work

    Barrington Media Group

    Shelton, CT
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Research Scientist, Speech & Audio. Be the first to apply!