Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Machine Learning Engineer, Speech - Joint Audio-Video Modeling

$200k - $220k
Full-time

Cantina

Role Description

We're looking for a Research / ML Engineer to join our Speech Team to build state-of-the-art speech and audio generation systems end-to-end from data specs through production inference with a focus on joint audio-video modeling.

You'll own the audio side of multimodal generation:

  • The representations (audio VAEs, neural codecs)
  • The generative backbone (diffusion / flow-matching transformers)
  • The conditioning and alignment machinery that makes characters speak, sing, and emote in sync with what's on screen

This includes:

  • Voice cloning and multi-speaker conditioning inside joint AV models
  • Cinematic dialogue with music and sound design
  • Adjacent speech tasks (controllable TTS, voice conversion) that feed the same stack

You'll drive the model ↔ data ↔ eval flywheel, partnering closely with research, video, data, and infra to ship fast, reliable, and cost-aware models. In this role you'll work at the intersection of cutting-edge research and practical engineering, contributing to the development of safe, steerable, and trustworthy AI systems.

You will thrive in this role if you:

  • See research and engineering as two sides of the same coin and enjoy owning work end-to-end
  • Are excited to work across modalities and collaborate closely with a video generation team rather than staying inside audio
  • Are results-oriented, flexible, and willing to pick up whatever moves the needle
  • Like collaborating closely with infra, data, and product to ship measurable improvements
  • Enjoy designing experiments, listening tests, and metrics that correlate with user-perceived quality
  • Are eager to learn every day, and to find and solve unique large-scale problems

Qualifications

  • Exceptional research/development experience with large-scale audio models (>8B parameters, >500k hours of data)
  • Deep hands-on experience with diffusion and/or flow-matching transformers, including practical knowledge of samplers, schedules, conditioning mechanisms, and distillation
  • Deep hands-on experience training audio VAEs, neural audio codecs, and vocoders latent/tokenizer design, reconstruction and perceptual objectives, adversarial training
  • Strong experience with multi-node, multi-GPU distributed training (FSDP/DeepSpeed or equivalent)
  • Strong software engineering skills with a proven track record of building complex systems
  • Strong with PyTorch and performance work (profiling, CUDA/Triton/C++ as needed) and writing reliable production-quality code
  • Shipped large-scale speech/audio or multimodal generative models to production
  • Background in working with large-scale ML data, and the ability to iterate on data and triangulate quality using both subjective and objective signals
  • Experience with voice cloning, speech control/steerability, or expressive speech generation
  • Notable publications and/or open-source contributions in speech/audio/ML

Requirements

  • Experience with multimodal audio-video modeling: joint AV generation of multi-shot, multi-speaker scenes with dialogue, music, and sound design generated jointly with video, and the cross-modal alignment that keeps them in sync
  • Experience with video generation: video diffusion/flow-matching transformers, video VAEs, conditioned and multi-shot generation, building data pipelines for video models
  • Streaming or real-time generation, causal distillation (e.g., Self Forcing / Self Forcing++)

Benefits

  • Competitive salary and generous company equity
  • Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina
  • 42 days of paid time off, including:
    • 15 PTO days
    • 10 sick days
    • 15 company holidays
    • 2 floating holidays
  • Generous parental leave & fertility support
  • 401(k) retirement savings plan
  • Lifestyle spending account – $500/month to use however you’d like
  • Complimentary lunch and snacks for in-office employees
  • One Medical membership, and more!
Vacancy posted 8 days ago
Similar jobs that could be interesting for youBased on the Machine Learning Engineer, Speech - Joint Audio-Video Modeling in Remote vacancy
  • $175k - $275k

     ...datasets (image, video, multimodal, reasoning...  ...full-cycle data engineering, Abaka AI provides...  ...hiring our first Machine Learning Engineer in the United...  ...roadmap for model development across...  ...across text, image, audio, video, and 3D data...  ...understanding, speech-audio modeling) and... 
    Video
    Audio
    Full time
    Immediate start
    Flexible hours

    Abaka Ai

    Remote
    1 day ago
  •  ...Today, not even the best models can continuously...  ...year-long stream of audio, video and text—1B text tokens...  ...innovation and systems engineering paired with a design-...  ...are searching for a Machine Learning Engineer to own the...  ...Build evaluations of speech models, both via... 
    Video
    Audio
    Full time
    Work at office
    Relocation package

    Cartesia

    Remote
    1 day ago
  •  ...reliability of today’s AI models? If yes, then this...  ...and scoring existing audio samples and providing...  ...constructive feedback to speech recording or...  ...train and improve AI and machine learning models. Typical tasks...  ...categories OR Speech and/or video recording Work benefits... 
    Video
    Audio
    Extra income
    Part time
    Freelance
    Remote work
    Work from home

    Wahojobs

    New York, NY
    2 days ago
  • $230k - $322k

     ...overview: The Engagement Modeling Team at Reddit focuses on building machine learning models to drive on-...  ...of click-throughs and video view-throughs. This role...  ...scientists, and other engineering teams to align on engagement...  ..., Sensory Information (audio/video recording), and... 
    Video
    Audio
    Remote job
    Full time
    For contractors
    Work experience placement
    Work at office
    Home office
    Flexible hours

    Reddit

    Remote
    19 days ago
  • $170k - $225k

     ...commerce ecosystems including video shopping platforms, social...  ...\• Architect and maintain machine learning systems across the full lifecycle...  ...textual data, imagery, and audio at enterprise scale \• Engineer and implement large language model applications utilizing... 
    Video
    Audio
    Full time
    Immediate start

    Rainesdev

    San Francisco, CA
    1 day ago
  • $142.65k - $213.98k

     ...diverse content sources—including **video, images, audio, and documents**—to generate rich...  ...platform unifies multiple ML/AI models to extract curated insights at...  ...looking for a **mid-level Backend Engineer** to join our **Machine Learning Platform team**. This role focuses... 
    Video
    Audio
    Full time
    Work at office
    Remote work
    Worldwide
    Flexible hours

    Comcast

    Washington DC
    1 day ago
  • $216.7k - $303.4k

     ...Efficiency function to make model training and inference materially...  ...person will be a key senior engineer on that team, owning...  ...vacation, and parental leave. To learn more, please visit . To provide...  ..., Sensory Information (audio/video recording), and any other categories... 
    Video
    Audio
    Remote job
    Full time
    For contractors
    Work experience placement
    Work at office
    Flexible hours

    Reddit

    United States
    1 day ago
  • $185.8k - $303.4k

     ...ll join a set of tight-knit engineers working on high-impact, internet...  ...We are hiring Machine Learning Engineers (IC3 and IC4) to build...  ...independently own scoped projects, ship models and services, and contribute...  ..., Sensory Information (audio/video recording), and any other... 
    Video
    Audio
    Full time
    For contractors
    Work experience placement
    Flexible hours

    Reddit

    United States
    1 day ago
  •  ...Machine Learning Engineer Location: US / Canada (Eastern Time) - Home-based Job Type: Full-time...  ...native services as well as custom-built models to deliver predictive insights to our...  ...IoT to unstructured data as images, audio, video and documents, and in between. Use... 
    Video
    Audio
    Remote job
    Permanent employment
    Full time
    Local area
    Work from home

    Allcloud

    New York, NY
    1 day ago
  • $116.84k - $142.81k

    Working Title: Senior Machine Learning Engineer [Hybrid] No Visa Sponsorship is available for this...  ...worldwide to identify birds from images, audio, and video.You will lead the architecture and...  ...full lifecycle of machine learning models, from training pipelines to... 
    Video
    Audio
    Full time
    Contract work
    Work at office
    Local area
    Worldwide
    Relocation
    Visa sponsorship

    Cornell University

    Ithaca, NY
    2 days ago
  • $266k - $372.4k

     ...We’re looking for a Senior Staff Machine Learning Engineer to lead Reddit’s next-generation user...  ...deep expertise in mainstream ML user modeling approaches (e.g., large-scale embeddings...  ...Related Information, Sensory Information (audio/video recording), and any other categories... 
    Video
    Audio
    Full time
    For contractors
    Work experience placement
    Immediate start
    Remote work
    Flexible hours
    Shift work

    Reddit

    United States
    1 day ago
  •  ...record pace, we are seeking experienced machine learning engineers who thrive on solving complex...  ...you will develop deep learning-based models for user understanding, content/query...  ...auto-generated summaries, images, and videos from experience content to improve discoverability... 
    Video
    Full time

    Roblox

    Remote
    1 day ago
  • $216.7k - $303.4k

     ..., Assurance, and Corporate Engineering organization protects Reddit...  ...practical, high-quality ML models that detect and prevent...  ...We are looking for a Senior Machine Learning Engineer to lead model development...  ..., Sensory Information (audio/video recording), and any other... 
    Video
    Audio
    Remote job
    Full time
    For contractors
    Work experience placement
    Flexible hours

    Reddit

    United States
    1 day ago
  • $216.7k - $303.4k

     ...agentic AI platforms that power machine learning development across Reddit. Our...  ...frameworks that enable faster model iterations. We are looking for an experienced engineer with deep expertise in large-...  ..., Sensory Information (audio/video recording), and any other categories... 
    Video
    Audio
    For contractors
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    GrabJobs

    Columbus, OH
    21 hours ago
  • $230k - $322k

     ...teams.  We are looking for a Staff Machine Learning Engineer who will lead the Commercial Content...  ...alignment), and 50% in direct hands-on work (modeling, pipelines, and debugging complex...  ...Information, Sensory Information (audio/video recording), and any other categories... 
    Video
    Audio
    Remote job
    Full time
    For contractors
    Work experience placement
    Flexible hours

    Reddit

    United States
    1 day ago
  • $90k - $110k

    Job Title Machine Learning Engineer Salary $90,000 - $110,000...  ...imagery, drone footage, and audio data. We focus on delivering...  ...and behavior analysis from video collar footage Please click...  ...Scientists to train machine learning models and create robust, scalable... 
    Video
    Audio
    Permanent employment
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    Flexible hours
    Shift work

    Western EcoSystems

    Fort Collins, CO
    1 hour ago
  • $216.7k - $303.4k

     ...Role: The Safety ML team is hiring a Machine Learning Engineer to build and iterate on the next...  ...optimizing state-of-the-art large language models (LLMs) in order to scalably and...  ...Related Information, Sensory Information (audio/video recording), and any other categories... 
    Video
    Audio
    Remote job
    Full time
    For contractors
    Work experience placement
    Work at office

    Reddit

    United States
    1 day ago
  • $140k

     ...Come talk to us to learn more about what it...  ...performance of a team of ML engineers and architects. You...  ..., data pipelines, model lifecycle, and...  ...negotiables) * 10+ years in machine learning or AI, with...  ...Processing, Video Understanding, Speech & Audio, Time Series & Forecasting... 
    Video
    Audio
    Full time
    Work at office
    Immediate start
    Remote work
    Flexible hours

    Caylent

    United States
    5 days ago
  •  ...implementation-focused  Senior Machine Learning Engineer / Applied Data Scientist...  ...machine learning models. The ideal candidate combines...  ...datasets, including  video, image, and audio data. Provide technical...  ...Experience with computer vision, speech/audio processing, and... 
    Video
    Audio
    Remote job

    Jetsoftpro LLC

    Remote
    a month ago
  • $26 - $28 per hour

     ...Data Labeling Analysts, supporting speech and voice AI systems. This is a...  ...systems. You’ll be working with audio, speech, and language data — helping ensure models are trained on accurate, well-structured...  ...employees must complete a live video verification with their selected... 
    Video
    Audio
    Full time
    Remote work
    Visa sponsorship

    Welo Global

    Sunnyvale, CA
    8 hours ago
  • $216.7k

     ...We are looking for a Senior Machine Learning Engineer (IC4) who will act as a key...  ...leverage hosted LLMs versus custom models, help scale content...  ...understanding to new modalities (e.g., video), and drive practical ML...  ..., Sensory Information (audio/video recording), and any other... 
    Video
    Audio
    For contractors
    Work experience placement
    Work at office
    Immediate start
    Remote work
    Flexible hours

    GrabJobs

    Philadelphia, PA
    1 day ago
  • $227.5k - $324.99k

     ...moderation, including detection models, policy enforcement...  ...Do Build and scale machine learning systems for proactive content...  ...models that combine text, audio, image, and video signals for safety and...  ...support other machine learning engineers, helping raise the bar... 
    Video
    Audio
    Full time
    Work from home
    Flexible hours

    Spotify

    Remote
    1 day ago
  • $227.5k - $324.99k

     ...global scale. We're seeking a Staff Machine Learning Engineer to build and scale foundational ML...  ...readable understanding of content across audio, video, text, and images-enabling automation...  ...content across modalities Develop models for classification, tagging, semantic... 
    Video
    Audio
    Work from home
    Worldwide
    Flexible hours

    Spotify

    New York, NY
    1 day ago
  • $153.76k - $225.51k

     ...the country. Our Engineering and Analytics Team Members...  ...in decision science, machine learning, and generative AI...  ...limited to large language models (LLMs), deep learning,...  ...AI (text, image, audio integration) Proficiency...  ...from you. Play the video below to learn more... 
    Video
    Audio
    Casual work
    Work at office
    Local area
    Remote work
    Work from home

    Credit Acceptance

    Southfield, MI
    8 hours ago
  • $253.3k - $354.6k

     ...visit . The LS Embedding Machine Learning team is at the forefront of...  ...expressive machine learning models that power Reddit’s recommendation...  ...a Staff Machine Learning Engineer , you will own the...  ...Information, Sensory Information (audio/video recording), and any other... 
    Video
    Audio
    Remote job
    Full time
    For contractors
    Work experience placement
    Flexible hours

    Reddit

    United States
    1 day ago
  • $185k - $400k

     ...accomplished Research Scientists in Foundation Models with expertise in pre-training and mid-...  ...-training/mid-training (text, image, audio, and video), and drive innovative approaches for...  .... You will collaborate closely with engineering and product teams, shaping the future... 
    Video
    Audio
    Remote work

    Pika

    Palo Alto, CA
    2 days ago
  • $184k - $356.5k

     ...advances world foundation models to enable high-fidelity, temporally stable video and world generation for...  ..., platform, and product engineering to ensure improvements...  ...models for image/video/audio, with strong fundamentals in modern deep learning. ~Hands-on experience... 
    Video
    Audio
    Full time

    NVIDIA

    Remote
    4 days ago
  •  ...human-like AI voice model. Today, we serve...  ...and edit speech, music, image, and video across 70+ languages...  ...our leading AI audio foundational models...  ...researchers, engineers, and operators....  ...responsibilities. Learning & development :...  ...might be running a joint program with a... 
    Video
    Audio
    Immediate start
    Remote work

    GrabJobs

    Charlotte, NC
    1 day ago
  • $266k - $372.4k

     ...quickly build a new kind of search engine. As a Senior Staff, you will...  .... You will also deploy ML models, integrate LLMs, and ensure...  ...vacation, and parental leave. To learn more, please visit . To...  ..., Sensory Information (audio/video recording), and any other categories... 
    Video
    Audio
    Full time
    For contractors
    Work experience placement
    Remote work
    Flexible hours

    Reddit

    United States
    1 day ago
  •  ...NiCE is looking for a Senior Machine Learning Engineer to join NiCE Labs Research (NLR), a team dedicated to model expertise and agent...  ...evaluation and optimization of speech-oriented AI models — covering...  ...field with a focus on speech or audio processing. Three or more... 
    Audio
    Full time
    Work at office
    Remote work
    Flexible hours

    Nice Inc.

    Remote
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Machine Learning Engineer, Speech - Joint Audio-Video Modeling. Be the first to apply!