Machine Learning Engineer, Speech - Joint Audio-Video Modeling
$200k - $220kCantina Llc
About Cantina:
Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning.
If you're excited about the potential AI has to shape human creativity and social interactions, join us in building the future!
About the Role:
We're looking for a Research / ML Engineer to join our Speech Team to build state-of-the-art speech and audio generation systems end-to-end from data specs through production inference with a focus on joint audio-video modeling.
You'll own the audio side of multimodal generation: the representations (audio VAEs, neural codecs), the generative backbone (diffusion / flow-matching transformers), and the conditioning and alignment machinery that makes characters speak, sing, and emote in sync with what's on screen. That includes voice cloning and multi-speaker conditioning inside joint AV models, cinematic dialogue with music and sound design, and adjacent speech tasks (controllable TTS, voice conversion) that feed the same stack.
You'll drive the model ↔ data ↔ eval flywheel, partnering closely with research, video, data, and infra to ship fast, reliable, and cost-aware models. In this role you'll work at the intersection of cutting-edge research and practical engineering, contributing to the development of safe, steerable, and trustworthy AI systems.
You will thrive in this role if you:
See research and engineering as two sides of the same coin and enjoy owning work end-to-end.
Are excited to work across modalities and collaborate closely with a video generation team rather than staying inside audio.
Are results-oriented, flexible, and willing to pick up whatever moves the needle.
Like collaborating closely with infra, data, and product to ship measurable improvements.
Enjoy designing experiments, listening tests, and metrics that correlate with user-perceived quality.
Are eager to learn every day, and to find and solve unique large-scale problems.
What You’ll Do:
Audio Representations: Design, train, and improve the audio VAEs, neural codecs, and vocoders our generative models sit on top of latent design, reconstruction and perceptual objectives, compression-vs-fidelity tradeoffs.
Model Building: Architect, implement, pre-train, fine-tune, and post-train/alignment (e.g., GRPO/DPO) diffusion and flow-matching transformers for large-scale audio and video generation.
Joint Audio-Video Modeling: Design the audio conditioning and cross-modal alignment inside joint AV models, audio latents alongside video latents, reference-audio and multi-speaker conditioning, multi shot generation audio/video modeling.
Experimental Design: Design, run, and analyze scientific experiments to advance our understanding of the models.
Data Ownership: Define data requirements and collaborate on acquisition, curation, AV-sync and quality filtering, annotation quality, and synthetic data strategies for paired audio-video and speech corpora.
Rigorous Evaluation: Design automated objective/subjective evaluations audio fidelity and intelligibility metrics, AV-sync, listening and viewing tests, robustness & bias checks, and red-team studies.
Inference Efficiency: Drive distillation, step-count reduction, quantization, and kernel/memory optimization to meet interactive latency and cost targets.
Pipeline Delivery: Harden the training → evaluation → inference pipeline; profile latency, memory, and cost; and meet production SLAs with robust monitoring and rollback.
GPU Scaling: Partner with infrastructure to run distributed training/inference on cloud fleets and productionize models with reliability and observability.
Project Leadership: Independently lead small research projects while collaborating on larger team initiatives, including cross-team work with video generation.
Tool Development: Develop and improve dev tooling to enhance team productivity.
Safety & Responsibility: Contribute to safety/consent guardrails, watermarking, and misuse/abuse mitigation for responsible voice and likeness technology.
What You’ll Bring:
Exceptional research/development experience with large-scale audio models (>8B parameters, >500k hours of data).
Deep hands-on experience with diffusion and/or flow-matching transformers, including practical knowledge of samplers, schedules, conditioning mechanisms, and distillation.
Deep hands-on experience training audio VAEs, neural audio codecs, and vocoders latent/tokenizer design, reconstruction and perceptual objectives, adversarial training.
Strong experience with multi-node, multi-GPU distributed training (FSDP/DeepSpeed or equivalent).
Strong software engineering skills with a proven track record of building complex systems.
Strong with PyTorch and performance work (profiling, CUDA/Triton/C++ as needed) and writing reliable production-quality code.
Shipped large-scale speech/audio or multimodal generative models to production.
Background in working with large-scale ML data, and the ability to iterate on data and triangulate quality using both subjective and objective signals.
Experience with voice cloning, speech control/steerability, or expressive speech generation.
Notable publications and/or open-source contributions in speech/audio/ML.
Strongly preferred:
Experience with multimodal audio-video modeling : joint AV generation of multi-shot, multi-speaker scenes with dialogue, music, and sound design generated jointly with video, and the cross-modal alignment that keeps them in sync.
Experience with video generation : video diffusion/flow-matching transformers, video VAEs, conditioned and multi-shot generation, building data pipelines for video models.
Streaming or real-time generation, causal distillation (e.g., Self Forcing / Self Forcing++).
Compensation:
The anticipated annual base salary range for this role is between $200,000-$220,000 (€170,000-€190,000). When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.
Benefits for U.S.-based roles:
Competitive salary and generous company equity
Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina
42 days of paid time off, including:
15 PTO days
10 sick days
15 company holidays
2 floating holidays
Generous parental leave & fertility support
401(k) retirement savings plan
Lifestyle spending account – $500/month to use however you’d like
Complimentary lunch and snacks for in-office employees
One Medical membership, and more!
$230k - $322k
...advertising platform. The Measurement Modeling team owns all the Machine Learning solutions for Ads Measurement,... ...We are looking for an IC5 Staff ML Engineer of Ads Identity Modeling, to... ...Information, Sensory Information (audio/video recording), and any other categories...VideoAudioRemote jobFull timeFor contractorsWork experience placementFlexible hours$175k - $275k
...datasets (image, video, multimodal, reasoning... ...full-cycle data engineering, Abaka AI provides... ...hiring our first Machine Learning Engineer in the United... ...roadmap for model development across... ...across text, image, audio, video, and 3D data... ...understanding, speech-audio modeling) and...VideoAudioFull timeImmediate startFlexible hours$200k - $220k
...online—across voice, video, and text. Create yourself... ...for a Research / ML Engineer to join our Speech Team to build state-... .... You’ll drive the model ↔ data ↔ eval... ...quality. Eager to learn every-day, find and solve... ...experience with large-scale audio models (8B parameters...VideoAudioFull timeWork at officeFlexible hours$180k - $270k
...security and privacy protection. To learn more about Plaud, please visit... ..., ultra-low-latency inference engines for large language models or foundational speech models. Understand the intricate... ...To-First-Token (or Time-To-First-Audio) in real-time streaming...AudioFull timeWork at officeWorldwide- ...Today, not even the best models can continuously... ...year-long stream of audio, video and text—1B text tokens... ...innovation and systems engineering paired with a design-... ...are searching for a Machine Learning Engineer to own the... ...Build evaluations of speech models, both via...VideoAudioFull timeWork at officeRelocation package
$230k - $322k
...overview: The Engagement Modeling Team at Reddit focuses on building machine learning models to drive on-... ...of click-throughs and video view-throughs. This role... ...scientists, and other engineering teams to align on engagement... ..., Sensory Information (audio/video recording), and...VideoAudioRemote jobFull timeFor contractorsWork experience placementWork at officeHome officeFlexible hours- ...technologies and solutions on our engineering teams. Our Engineering Team... ..., embedded systems, machine learning, and the use of artificial... ...deploying machine learning models in creative ways while working... ...NLP, SLAM, forecasting, or audio/video processingProficient...VideoAudioFull time
$170k - $225k
...commerce ecosystems including video shopping platforms, social... ...\• Architect and maintain machine learning systems across the full lifecycle... ...textual data, imagery, and audio at enterprise scale \• Engineer and implement large language model applications utilizing...VideoAudioFull timeImmediate start$142.65k - $213.98k
...diverse content sources—including **video, images, audio, and documents**—to generate rich... ...platform unifies multiple ML/AI models to extract curated insights at... ...looking for a **mid-level Backend Engineer** to join our **Machine Learning Platform team**. This role focuses...VideoAudioFull timeWork at officeRemote workWorldwideFlexible hours$152k - $241.5k
...and advances in deep learning, Generative AI, and Cloud... ...of research and engineering, to guide and enable... ...customers adopt NVIDIA AI models and libraries by... ...multimodal, retrieval, speech, content safety, and... ...to multi-modal data (audio, image, and video).Knowledge of NVIDIA...VideoAudioFull time$209k - $313k
...live in the moment, learn about the world, and... ...multimodal LLMs and video generation to real-time... ...including foundational models, efficient... ....We're looking for a Machine Learning Engineer to join our Generative... ...for image, video, and audio generationDeliver generative...VideoAudioFull timeLive inWork at officeLocal areaWorldwide$85k - $125k
...processing (NLP), vision-language models (VLMs), and other... ...data sources like images, video, metadata, audio, and text, and we recognize... ...researchers on projects related to machine learning, artificial intelligence,... ..., Electrical and Computer Engineering, or related...VideoAudioFull timeContract workTemporary workWorldwideFlexible hours$120k - $180k
...best-in-class, pre-trained AI models, serving billions of... ...joining the future of AI! Machine Learning Role In order to execute... ...-in-class machine learning engineers. We are looking for... ...machine model (image, NLP, video, or audio) into production, with measurably...VideoAudioFull time- ...to define a new kind of video game experience. You'll... ...AI researchers and engineers in the world in a small... ...image and video generation model teams in the world on... ...training text-to-motion or audio-to-motion models.... ...if you have worked on speech-driven 3D facial animation...VideoAudioWork at officeVisa sponsorship
$185.8k - $303.4k
...ll join a set of tight-knit engineers working on high-impact, internet... ...We are hiring Machine Learning Engineers (IC3 and IC4) to build... ...independently own scoped projects, ship models and services, and contribute... ..., Sensory Information (audio/video recording), and any other...VideoAudioFull timeFor contractorsWork experience placementFlexible hours$116.84k - $142.81k
Working Title: Senior Machine Learning Engineer [Hybrid] No Visa Sponsorship is available for this... ...worldwide to identify birds from images, audio, and video.You will lead the architecture and... ...full lifecycle of machine learning models, from training pipelines to...VideoAudioFull timeContract workWork at officeLocal areaWorldwideRelocationVisa sponsorship- ...reliability of today’s AI models? If yes, then this... ...and scoring existing audio samples and providing... ...constructive feedback to speech recording or... ...train and improve AI and machine learning models. Typical tasks... ...categories OR Speech and/or video recording Work benefits...VideoAudioExtra incomePart timeFreelanceRemote workWork from home
- ...Francisco startup is seeking a Senior Machine Learning Engineer to build and scale AI systems. You... ...work on impactful projects with models used daily by many users, with... .... Build multimodal ML systems for video, text, images, and audio data. Develop applications using Large...VideoAudio
$250k
...Discord. They are training foundation models from the ground up for full-duplex audiovisual... ...-lived WebRTC connections that keep video and audio smooth Build and orchestrate robust... ...closely with ML researchers and product engineers to build the foundational infrastructure...VideoAudioFull timeH1bVisa sponsorshipRelocation package$230k - $322k
...likely to find them useful. As a Staff Machine Learning Engineer on Shopping Ads, you will lead the... ...technical strategy and execution for the models that power Shopping Ads delivery. You... ...Information, Sensory Information (audio/video recording), and any other categories...VideoAudioFull timeFor contractorsWork experience placementFlexible hoursShift work- ...record pace, we are seeking experienced machine learning engineers who thrive on solving complex... ...you will develop deep learning-based models for user understanding, content/query... ...auto-generated summaries, images, and videos from experience content to improve discoverability...VideoFull time
$266k - $372.4k
...We’re looking for a Senior Staff Machine Learning Engineer to lead Reddit’s next-generation user... ...deep expertise in mainstream ML user modeling approaches (e.g., large-scale embeddings... ...Related Information, Sensory Information (audio/video recording), and any other categories...VideoAudioFull timeFor contractorsWork experience placementImmediate startRemote workFlexible hoursShift work- ...Description Join Runware as a Senior Machine Learning Engineer and be at the forefront of... ...media modalities including text, image, video, 3D, and audio. We're building a powerful AI media... ...Integrate open-source and third-party models into our inference platform ~Lead...VideoAudioFull timeRemote workWork from homeFlexible hours
$216.7k - $303.4k
...agentic AI platforms that power machine learning development across Reddit. Our... ...frameworks that enable faster model iterations. We are looking for an experienced engineer with deep expertise in large-... ..., Sensory Information (audio/video recording), and any other categories...VideoAudioFor contractorsWork experience placementWork at officeRemote workFlexible hours$216.7k - $303.4k
..., Assurance, and Corporate Engineering organization protects Reddit... ...practical, high-quality ML models that detect and prevent... ...We are looking for a Senior Machine Learning Engineer to lead model development... ..., Sensory Information (audio/video recording), and any other...VideoAudioRemote jobFull timeFor contractorsWork experience placementFlexible hours$216.7k - $303.4k
...Efficiency function to make model training and inference materially... ...person will be a key senior engineer on that team, owning... ...vacation, and parental leave. To learn more, please visit . To provide... ..., Sensory Information (audio/video recording), and any other categories...VideoAudioFor contractorsWork experience placementWork at officeRemote workFlexible hours$216.7k - $303.4k
...Role: The Safety ML team is hiring a Machine Learning Engineer to build and iterate on the next... ...optimizing state-of-the-art large language models (LLMs) in order to scalably and... ...Related Information, Sensory Information (audio/video recording), and any other categories...VideoAudioFor contractorsWork experience placementWork at officeRemote work$174.72k - $295.68k
...through cutting-edge R&D in AI, machine learning, and smart connectivity.We are looking... ...a full-time Machine Learning Engineer / Research Scientist to drive the modeling and algorithmic development of... ...unlabeled fleet data (images, video, LiDAR, CAN bus, maps, human driving...VideoFull time$171.6k - $230.1k
...global organization of engineers, product developers, designers... ...working across multiple machine learning areas with primary focus... ...mixed media, language models, and other agentic... ...may include generative video, generative image, generative audio, chatbots, LLM applications...VideoAudio- ...to reinvent the way people learn, starting with language.... ...looking for an experienced Machine Learning Engineer to join our team and help develop cutting-edge speech recognition models that help teach language fluency... ...Experience with speech or audio Office ~ San...AudioFull timeLive inWork at officeWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Machine Learning Engineer, Speech - Joint Audio-Video Modeling. Be the first to apply!
- staff machine learning engineer United States
- entry level machine learning engineer United States
- senior ml engineer United States
- machine learning engineer United States
- graduate machine learning engineer United States
- junior machine learning engineer United States
- data scientist machine learning engineer United States
- ai ml engineer United States
- junior machine learning research engineer United States
- lead machine learning engineer United States



