Machine Learning Engineer, Speech - Joint Audio-Video Modeling
$200k - $220kCantina
Role Description
We're looking for a Research / ML Engineer to join our Speech Team to build state-of-the-art speech and audio generation systems end-to-end from data specs through production inference with a focus on joint audio-video modeling.
You'll own the audio side of multimodal generation:
- The representations (audio VAEs, neural codecs)
- The generative backbone (diffusion / flow-matching transformers)
- The conditioning and alignment machinery that makes characters speak, sing, and emote in sync with what's on screen
This includes:
- Voice cloning and multi-speaker conditioning inside joint AV models
- Cinematic dialogue with music and sound design
- Adjacent speech tasks (controllable TTS, voice conversion) that feed the same stack
You'll drive the model ↔ data ↔ eval flywheel, partnering closely with research, video, data, and infra to ship fast, reliable, and cost-aware models. In this role you'll work at the intersection of cutting-edge research and practical engineering, contributing to the development of safe, steerable, and trustworthy AI systems.
You will thrive in this role if you:
- See research and engineering as two sides of the same coin and enjoy owning work end-to-end
- Are excited to work across modalities and collaborate closely with a video generation team rather than staying inside audio
- Are results-oriented, flexible, and willing to pick up whatever moves the needle
- Like collaborating closely with infra, data, and product to ship measurable improvements
- Enjoy designing experiments, listening tests, and metrics that correlate with user-perceived quality
- Are eager to learn every day, and to find and solve unique large-scale problems
Qualifications
- Exceptional research/development experience with large-scale audio models (>8B parameters, >500k hours of data)
- Deep hands-on experience with diffusion and/or flow-matching transformers, including practical knowledge of samplers, schedules, conditioning mechanisms, and distillation
- Deep hands-on experience training audio VAEs, neural audio codecs, and vocoders latent/tokenizer design, reconstruction and perceptual objectives, adversarial training
- Strong experience with multi-node, multi-GPU distributed training (FSDP/DeepSpeed or equivalent)
- Strong software engineering skills with a proven track record of building complex systems
- Strong with PyTorch and performance work (profiling, CUDA/Triton/C++ as needed) and writing reliable production-quality code
- Shipped large-scale speech/audio or multimodal generative models to production
- Background in working with large-scale ML data, and the ability to iterate on data and triangulate quality using both subjective and objective signals
- Experience with voice cloning, speech control/steerability, or expressive speech generation
- Notable publications and/or open-source contributions in speech/audio/ML
Requirements
- Experience with multimodal audio-video modeling: joint AV generation of multi-shot, multi-speaker scenes with dialogue, music, and sound design generated jointly with video, and the cross-modal alignment that keeps them in sync
- Experience with video generation: video diffusion/flow-matching transformers, video VAEs, conditioned and multi-shot generation, building data pipelines for video models
- Streaming or real-time generation, causal distillation (e.g., Self Forcing / Self Forcing++)
Benefits
- Competitive salary and generous company equity
- Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina
- 42 days of paid time off, including:
- 15 PTO days
- 10 sick days
- 15 company holidays
- 2 floating holidays
- Generous parental leave & fertility support
- 401(k) retirement savings plan
- Lifestyle spending account – $500/month to use however you’d like
- Complimentary lunch and snacks for in-office employees
- One Medical membership, and more!
$175k - $275k
...datasets (image, video, multimodal, reasoning... ...full-cycle data engineering, Abaka AI provides... ...hiring our first Machine Learning Engineer in the United... ...roadmap for model development across... ...across text, image, audio, video, and 3D data... ...understanding, speech-audio modeling) and...VideoAudioFull timeImmediate startFlexible hours- ...Today, not even the best models can continuously... ...year-long stream of audio, video and text—1B text tokens... ...innovation and systems engineering paired with a design-... ...are searching for a Machine Learning Engineer to own the... ...Build evaluations of speech models, both via...VideoAudioFull timeWork at officeRelocation package
- ...reliability of today’s AI models? If yes, then this... ...and scoring existing audio samples and providing... ...constructive feedback to speech recording or... ...train and improve AI and machine learning models. Typical tasks... ...categories OR Speech and/or video recording Work benefits...VideoAudioExtra incomePart timeFreelanceRemote workWork from home
$230k - $322k
...overview: The Engagement Modeling Team at Reddit focuses on building machine learning models to drive on-... ...of click-throughs and video view-throughs. This role... ...scientists, and other engineering teams to align on engagement... ..., Sensory Information (audio/video recording), and...VideoAudioRemote jobFull timeFor contractorsWork experience placementWork at officeHome officeFlexible hours$170k - $225k
...commerce ecosystems including video shopping platforms, social... ...\• Architect and maintain machine learning systems across the full lifecycle... ...textual data, imagery, and audio at enterprise scale \• Engineer and implement large language model applications utilizing...VideoAudioFull timeImmediate start$142.65k - $213.98k
...diverse content sources—including **video, images, audio, and documents**—to generate rich... ...platform unifies multiple ML/AI models to extract curated insights at... ...looking for a **mid-level Backend Engineer** to join our **Machine Learning Platform team**. This role focuses...VideoAudioFull timeWork at officeRemote workWorldwideFlexible hours$216.7k - $303.4k
...Efficiency function to make model training and inference materially... ...person will be a key senior engineer on that team, owning... ...vacation, and parental leave. To learn more, please visit . To provide... ..., Sensory Information (audio/video recording), and any other categories...VideoAudioRemote jobFull timeFor contractorsWork experience placementWork at officeFlexible hours$185.8k - $303.4k
...ll join a set of tight-knit engineers working on high-impact, internet... ...We are hiring Machine Learning Engineers (IC3 and IC4) to build... ...independently own scoped projects, ship models and services, and contribute... ..., Sensory Information (audio/video recording), and any other...VideoAudioFull timeFor contractorsWork experience placementFlexible hours- ...Machine Learning Engineer Location: US / Canada (Eastern Time) - Home-based Job Type: Full-time... ...native services as well as custom-built models to deliver predictive insights to our... ...IoT to unstructured data as images, audio, video and documents, and in between. Use...VideoAudioRemote jobPermanent employmentFull timeLocal areaWork from home
$116.84k - $142.81k
Working Title: Senior Machine Learning Engineer [Hybrid] No Visa Sponsorship is available for this... ...worldwide to identify birds from images, audio, and video.You will lead the architecture and... ...full lifecycle of machine learning models, from training pipelines to...VideoAudioFull timeContract workWork at officeLocal areaWorldwideRelocationVisa sponsorship$266k - $372.4k
...We’re looking for a Senior Staff Machine Learning Engineer to lead Reddit’s next-generation user... ...deep expertise in mainstream ML user modeling approaches (e.g., large-scale embeddings... ...Related Information, Sensory Information (audio/video recording), and any other categories...VideoAudioFull timeFor contractorsWork experience placementImmediate startRemote workFlexible hoursShift work- ...record pace, we are seeking experienced machine learning engineers who thrive on solving complex... ...you will develop deep learning-based models for user understanding, content/query... ...auto-generated summaries, images, and videos from experience content to improve discoverability...VideoFull time
$216.7k - $303.4k
..., Assurance, and Corporate Engineering organization protects Reddit... ...practical, high-quality ML models that detect and prevent... ...We are looking for a Senior Machine Learning Engineer to lead model development... ..., Sensory Information (audio/video recording), and any other...VideoAudioRemote jobFull timeFor contractorsWork experience placementFlexible hours$216.7k - $303.4k
...agentic AI platforms that power machine learning development across Reddit. Our... ...frameworks that enable faster model iterations. We are looking for an experienced engineer with deep expertise in large-... ..., Sensory Information (audio/video recording), and any other categories...VideoAudioFor contractorsWork experience placementWork at officeRemote workFlexible hours$230k - $322k
...teams. We are looking for a Staff Machine Learning Engineer who will lead the Commercial Content... ...alignment), and 50% in direct hands-on work (modeling, pipelines, and debugging complex... ...Information, Sensory Information (audio/video recording), and any other categories...VideoAudioRemote jobFull timeFor contractorsWork experience placementFlexible hours$90k - $110k
Job Title Machine Learning Engineer Salary $90,000 - $110,000... ...imagery, drone footage, and audio data. We focus on delivering... ...and behavior analysis from video collar footage Please click... ...Scientists to train machine learning models and create robust, scalable...VideoAudioPermanent employmentFull timeWork experience placementWork at officeLocal areaRemote workFlexible hoursShift work$216.7k - $303.4k
...Role: The Safety ML team is hiring a Machine Learning Engineer to build and iterate on the next... ...optimizing state-of-the-art large language models (LLMs) in order to scalably and... ...Related Information, Sensory Information (audio/video recording), and any other categories...VideoAudioRemote jobFull timeFor contractorsWork experience placementWork at office$140k
...Come talk to us to learn more about what it... ...performance of a team of ML engineers and architects. You... ..., data pipelines, model lifecycle, and... ...negotiables) * 10+ years in machine learning or AI, with... ...Processing, Video Understanding, Speech & Audio, Time Series & Forecasting...VideoAudioFull timeWork at officeImmediate startRemote workFlexible hours- ...implementation-focused Senior Machine Learning Engineer / Applied Data Scientist... ...machine learning models. The ideal candidate combines... ...datasets, including video, image, and audio data. Provide technical... ...Experience with computer vision, speech/audio processing, and...VideoAudioRemote job
$26 - $28 per hour
...Data Labeling Analysts, supporting speech and voice AI systems. This is a... ...systems. You’ll be working with audio, speech, and language data — helping ensure models are trained on accurate, well-structured... ...employees must complete a live video verification with their selected...VideoAudioFull timeRemote workVisa sponsorship$216.7k
...We are looking for a Senior Machine Learning Engineer (IC4) who will act as a key... ...leverage hosted LLMs versus custom models, help scale content... ...understanding to new modalities (e.g., video), and drive practical ML... ..., Sensory Information (audio/video recording), and any other...VideoAudioFor contractorsWork experience placementWork at officeImmediate startRemote workFlexible hours$227.5k - $324.99k
...moderation, including detection models, policy enforcement... ...Do Build and scale machine learning systems for proactive content... ...models that combine text, audio, image, and video signals for safety and... ...support other machine learning engineers, helping raise the bar...VideoAudioFull timeWork from homeFlexible hours$227.5k - $324.99k
...global scale. We're seeking a Staff Machine Learning Engineer to build and scale foundational ML... ...readable understanding of content across audio, video, text, and images-enabling automation... ...content across modalities Develop models for classification, tagging, semantic...VideoAudioWork from homeWorldwideFlexible hours$153.76k - $225.51k
...the country. Our Engineering and Analytics Team Members... ...in decision science, machine learning, and generative AI... ...limited to large language models (LLMs), deep learning,... ...AI (text, image, audio integration) Proficiency... ...from you. Play the video below to learn more...VideoAudioCasual workWork at officeLocal areaRemote workWork from home$253.3k - $354.6k
...visit . The LS Embedding Machine Learning team is at the forefront of... ...expressive machine learning models that power Reddit’s recommendation... ...a Staff Machine Learning Engineer , you will own the... ...Information, Sensory Information (audio/video recording), and any other...VideoAudioRemote jobFull timeFor contractorsWork experience placementFlexible hours$185k - $400k
...accomplished Research Scientists in Foundation Models with expertise in pre-training and mid-... ...-training/mid-training (text, image, audio, and video), and drive innovative approaches for... .... You will collaborate closely with engineering and product teams, shaping the future...VideoAudioRemote work$184k - $356.5k
...advances world foundation models to enable high-fidelity, temporally stable video and world generation for... ..., platform, and product engineering to ensure improvements... ...models for image/video/audio, with strong fundamentals in modern deep learning. ~Hands-on experience...VideoAudioFull time- ...human-like AI voice model. Today, we serve... ...and edit speech, music, image, and video across 70+ languages... ...our leading AI audio foundational models... ...researchers, engineers, and operators.... ...responsibilities. Learning & development :... ...might be running a joint program with a...VideoAudioImmediate startRemote work
$266k - $372.4k
...quickly build a new kind of search engine. As a Senior Staff, you will... .... You will also deploy ML models, integrate LLMs, and ensure... ...vacation, and parental leave. To learn more, please visit . To... ..., Sensory Information (audio/video recording), and any other categories...VideoAudioFull timeFor contractorsWork experience placementRemote workFlexible hours- ...NiCE is looking for a Senior Machine Learning Engineer to join NiCE Labs Research (NLR), a team dedicated to model expertise and agent... ...evaluation and optimization of speech-oriented AI models — covering... ...field with a focus on speech or audio processing. Three or more...AudioFull timeWork at officeRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Machine Learning Engineer, Speech - Joint Audio-Video Modeling. Be the first to apply!
- senior ml engineer Remote
- machine learning engineer Remote
- graduate machine learning engineer Remote
- data scientist machine learning engineer Remote
- ai ml engineer Remote
- junior machine learning research engineer Remote
- machine learning software engineer Remote
- computer vision machine learning engineer Remote
- machine learning ai engineer Remote
- audio visual Remote



