Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Machine Learning Engineer, Voice AI

$220k - $280k
Full-time

Together AI

About the Role

Together AI is building the best inference infrastructure for voice applications. Our Voice AI platform powers production-grade, real-time voice agents and applications — serving speech-to-text and text-to-speech models with best-in-class latency and reliability.

We're looking for a Staff ML Engineer to drive the model serving layer for voice workloads. You'll work hands-on with inference engines like TRT-LLM and SGLang to optimize how we serve models like Whisper, Parakeet, Orpheus, and Kokoro — pushing latency and throughput to the frontier. You'll profile GPU utilization, design batching strategies for streaming audio, and ensure new model architectures can go from research to production quickly.

This is a foundational hire on a small, high-impact team. Voice inference has unique challenges — streaming audio, tokenization, real-time latency budgets — that require dedicated ML engineering focus. You'll shape how Together serves voice models as the industry moves from pipeline architectures (ASR → LLM → TTS) toward end-to-end speech-to-speech.

  • Own the model serving stack that powers Together's voice platform across STT, TTS, and speech-to-speech.
  • Work directly with state-of-the-art accelerators (H100s, H200s, B200s) to optimize voice model inference.
  • Collaborate with model partners (Cartesia, Deepgram, Rime, and others) to bring their models to production on Together's infrastructure.
  • Build quality evaluation frameworks that guide model selection for customers and inform the roadmap.
  • Join a small, early-stage team with outsized impact on a fast-growing product area.

Responsibilities

  • Own the voice inference roadmap end-to-end — define and execute the technical strategy for optimizing STT, TTS, and speech-to-speech models across Together's infrastructure, with a clear-eyed view of where the field is heading and how to position the platform ahead of it.
  • Drive best-in-class inference performance — architect and implement systems targeting leading TTFB, throughput, and GPU utilization for voice workloads; set the performance bar others in the industry measure against, not just catch up to.
  • Lead productionization of voice models at scale — design the serving architecture for serverless and dedicated endpoints, including batching strategies, streaming inference pipelines, and memory management tailored to real-time audio; own reliability and latency SLAs.
  • Build the voice evaluation platform — design a rigorous, extensible evaluation framework covering WER across accents, languages, and noise conditions for STT; naturalness, latency, and pronunciation fidelity for TTS; establish the internal benchmark methodology that informs model selection and roadmap decisions.
  • Shape the architecture for next-generation model support — anticipate and enable emerging model paradigms — audio-native LLMs, codec-based architectures (SNAC, Encodec), and end-to-end speech-to-speech systems — before they're mainstream, not after.
  • Serve as the technical DRI for model partner integrations — lead deep collaboration with partners such as Cartesia, Deepgram, and Rime; own the full lifecycle from integration to optimization to ongoing performance accountability.
  • Diagnose and resolve the hardest performance problems in the stack — conduct systematic profiling and root-cause analysis from GPU kernel behavior to framework-level bottlenecks; drive shipped improvements with documented, measurable impact.
  • Influence platform architecture across the organization — partner with platform engineering leadership to ensure the serving layer is built for the latency and reliability demands of real-time voice APIs; your technical decisions should raise the ceiling for the whole team.
  • Define and scale voice fine-tuning capabilities — lead the technical direction for enabling customers to fine-tune STT and TTS models on Together's infrastructure, establishing the primitives for differentiated voice experiences.
  • Lay technical foundations for a category-defining product surface — architect systems with enough foresight that they support multiple new voice products with minimal rework; think in terms of platforms, not point solutions.

Requirements

  • 8+ years of ML engineering experience, with a demonstrated focus on model serving, inference optimization, or ML infrastructure at production scale — including systems you've owned from design through live traffic.
  • Deep, practical expertise in LLM serving engines (vLLM, SGLang, TensorRT-LLM, or equivalent) — you've modified engine internals, debugged edge cases under load, and contributed improvements back; you don't stop at the API surface.
  • Expert-level Python and PyTorch proficiency, with a strong command of GPU optimization — CUDA kernels, memory hierarchies, profiling toolchains — and a track record of turning that knowledge into shipped latency or throughput wins.
  • Proven system design judgment — you've made architectural decisions that held up at scale and influenced how a team or platform evolved; you can articulate the tradeoffs you made and why.
  • Strong technical leadership — you operate with high autonomy, define the right problems before solving them, and raise the bar for engineering quality around you without requiring process overhead.
  • Sharp product intuition for developer tooling — you understand what voice application developers actually need to ship great products, and you let that shape your technical priorities, not just the other way around.
  • Proven ability to move fast in ambiguous environments — you've thrived on early-stage or platform teams where scope is wide, ownership is deep, and the roadmap you build is the one you execute.
  • Strong foundation in speech and audio ML (ASR/TTS architectures, audio signal processing) — directly relevant experience is strongly preferred; exceptional ML engineering fundamentals with genuine curiosity about the domain is also considered.
  • Familiarity with audio codec and tokenization schemes (SNAC, Encodec, DAC) is a meaningful plus at this level.
  • Experience training or fine-tuning speech models at scale is a significant advantage.
  • Bachelor's or Master's in Computer Science, Electrical Engineering, or related field — or equivalent depth demonstrated through your work.

About Together AI

Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month.

Compensation

We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $220,000 - $280,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.

Equal Opportunity

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Please see our privacy policy at

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Staff Machine Learning Engineer, Voice AI in San Francisco, CA vacancy
  • $160k - $230k

     ...About the Role Together AI is building the best inference infrastructure for voice applications. Our Voice AI platform...  ...We're looking for a Senior ML Engineer to drive the model serving layer...  ...plus but not required — you can learn this quickly if you have strong... 
    Suggested
    Full time

    Together AI

    San Francisco, CA
    2 days ago
  • $181k - $250k

     ...Aircall is a unicorn, AI-powered customer communications...  ...by bringing voice, SMS, WhatsApp, and AI...  ...ownership, continuous learning, and thoughtful speed....  ...BS in Computer Science, Machine Learning, Statistics, or...  ...years of experience in ML Engineering or Applied ML with 8+ years... 
    Suggested
    Contract work
    Worldwide

    Aircall.io, Inc.

    San Francisco, CA
    9 days ago
  • $250k - $350k

    About ScaleScale’s mission is to develop reliable AI systems for the world’s most important decisions. As the leading...  ...-grade agentic AI looks like.About the RoleAs a Staff Machine Learning Research Engineer, you will operate across the full breadth of AIS’s technical... 
    Suggested
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  • $209k - $313k

     ...themselves, live in the moment, learn about the world, and...  ...digital services.Snap Engineering teams build fun and...  ....We’re looking for a Machine Learning Engineer to join...  ...driven featuresUtilize AI tools to design and...  ...diverse backgrounds and voices working together will enable... 
    Suggested
    Full time
    Live in
    Work at office
    Local area

    Snap

    San Francisco, CA
    20 hours ago
  • $173k - $259k

     ...themselves, live in the moment, learn about the world, and...  ...digital services.Snap Engineering teams build fun and...  ....We’re looking for a Machine Learning Engineer to join...  ...driven featuresUtilize AI tools and high velocity...  ...backgrounds and voices working together will enable... 
    Suggested
    Full time
    Live in
    Work at office
    Local area

    Snap

    San Francisco, CA
    20 hours ago
  • $173k - $259k

     ...themselves, live in the moment, learn about the world, and...  ...services. Snap Engineering ( teams build fun and...  ...We're looking for a Machine Learning Engineer to join...  ...driven features Utilize AI tools and high velocity...  ...backgrounds and voices working together will enable... 
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Snap

    San Francisco, CA
    20 hours ago
  •  ...just another scribe. We're building the AI intelligence platform that restores...  ...started. The Role: As a Senior Machine Learning Engineer at Ambience , you will build and improve...  ...to-Haves Experience with realtime voice, conversational AI, or multimodal systems... 
    Work at office
    Immediate start
    Remote work
    Flexible hours
    3 days per week

    Ambience Healthcare

    San Francisco, CA
    4 days ago
  •  ...build and orchestrate AI workforces. Our AI workers...  ...autonomously across voice, email, and enterprise...  ...runs the real economy. Learn more about our vision in...  ...Collaborate with product and engineering teams to integrate and...  ...Strong experience in machine learning, deep learning... 
    Worldwide
    Shift work

    Happy Robot

    San Francisco, CA
    3 days ago
  •  ...the agentic document platform for leading AI teams who demand enterprise performance...  ..., with high agency, and who doesn't just voice problems but actively jumps in to fix them...  ...customers to shape the product direction and engineering strategy Bonus points if you:... 
    Work at office
    Local area

    Reducto

    San Francisco, CA
    3 days ago
  •  ...Machine Learning Engineer We are looking for a Machine Learning Engineer to join the growing AI and Machine Learning team at Strava. This team is responsible for sophisticated machine...  ...Shape AI at Strava: Be a strong voice on a highly collaborative team with a range... 
    Work at office
    Worldwide
    Flexible hours
    3 days per week

    Strava

    San Francisco, CA
    4 days ago
  • $229k - $343k

     ..., live in the moment, learn about the world, and have...  ...digital services.Snap Engineering teams build fun and...  ....​We’re looking for a Staff Machine Learning Engineer to join...  ...ranking, generative AI, LLM-based ranking,...  ...diverse backgrounds and voices working together will... 
    Full time
    Live in
    Work at office
    Local area

    Snap

    San Francisco, CA
    1 day ago
  •  ...mission is to reinvent the way people learn, starting with language. Learning a...  ...decades. Speak is building a human-level, AI-powered tutor in your pocket: a...  ...role We are looking for an experienced Machine Learning Engineer to join our team and help develop cutting... 
    Full time
    Live in
    Work at office
    Worldwide

    Speak

    San Francisco, CA
    more than 2 months ago
  • $70 per hour

     ...across 15+ U.S. states. Software Engineering builds the brains of Waymo's...  ..., decision-making and deep learning, while collaborating with...  ...program in Computer Science, Machine Learning or a related field,...  ...published papers in top-tier AI/ML, data mining, or computer... 
    Hourly pay
    Full time
    Internship
    Summer internship
    Relocation package

    Waymo

    San Francisco, CA
    2 days ago
  • $70 per hour

     ...simulation across 15+ U.S. states. Software Engineering builds the brains of Waymo's fully...  ...robotics, perception, decision-making and deep learning, while collaborating with hardware and...  ...driving, simulation, evaluation, and physical AI. We prefer: Knowledge/prior... 
    Hourly pay
    Full time
    Internship
    Summer internship

    Waymo

    San Francisco, CA
    2 days ago
  •  ...DoorDash's logistics network. The Drive Machine Learning team builds the prediction and intelligence...  ..., logistics decision-making, and AI-powered delivery quality signals.Drive presents...  ....About the RoleAs a Machine Learning Engineer on the Drive team, you'll own machine learning... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Relocation
    Flexible hours

    Doordash

    San Francisco, CA
    4 days ago
  • $151.8k - $265.35k

     ...enterprise managed-service offering for custom multimedia generative AI — deep-tuned image, video, and 3D models built on each...  ...rapidly into adjacent verticals. We are hiring a Senior Machine Learning Engineer to build the pipelines and services that turn Firefly Foundry... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Francisco, CA
    1 day ago
  •  ...artificial intelligence and advanced ML, deep learning techniques to power decision-making in...  .... About the RoleWe’re looking for a Machine Learning Engineer to help design, build, optimize and...  ...a related field.Proficiency in using AI coding tools (e.g., Claude Code, Codex... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    2 days ago
  • $160k - $240k

     ...As a Fortune 500 company and a leading AI platform for managing people, money, and...  ...challenging problems at the intersection of machine learning, agentic reasoning, and enterprise-scale...  ....About the RoleAs a Machine Learning Engineer on the AI Core team, you will develop tailored... 
    Full time
    Work at office
    Remote work
    Home office
    Flexible hours

    Workday

    San Francisco, CA
    2 days ago
  •  ...SuperhumanGrammarly is now part of Superhuman, the AI productivity platform on a mission to...  ...busywork and focus on what matters. Learn more at superhuman.com and about our...  ...Superhuman ubiquitous UI. As a Machine Learning Engineer on this team, you will be at the heart... 
    Worldwide
    Home office
    Flexible hours

    Superhuman

    San Francisco, CA
    20 hours ago
  • $172.5k - $306.63k

     ...creative ecosystem. Our mission is to employ machine learning to enhance our comprehension of the...  ...What You'll DoAs a Senior Machine Learning Engineer on the Content Intelligence team, you...  ...with the latest advancements in ML and AI to ensure our solutions remain at the forefront... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Francisco, CA
    3 days ago
  • $200k - $235k

     ...our community. The Trust Frontier AI team is where new AI technology for Trust...  ...product managers, data scientists, software engineers, fraud intelligence, and operations...  ...Difference You Will Make: As a Senior Machine Learning Engineer on the Trust Frontier AI team,... 
    Work experience placement
    Casual work
    Live in
    Work at office
    Remote work

    Airbnb

    San Francisco, CA
    1 day ago
  • $161.26k - $332.01k

     ...? It’s Possible.At Pinterest, AI isn't just a feature, it's a powerful...  ..., multimodal representation learning, heterogeneous graph neural...  ...pod is a small group (~6 engineers along with a product prototyping...  ...vision experience.M.S. or PhD in Machine Learning, Computer Science, or... 
    Currently hiring
    Work at office
    Local area
    Remote work
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    1 day ago
  • $244k - $320k

    Attentive is the AI marketing platform for 1:1 personalization redefining the way brands...  ..., our AI-powered personalization engine delivers bespoke experiences that drive...  ...Corporate Equality Index!About the RoleOur Machine Learning Engineering team powers personalized experiences... 
    Full time

    Attentive

    San Francisco, CA
    2 days ago
  • Overview Atlassian is looking for a Senior Machine Learning Engineer to join our Search & Intelligence organization. We build AI-native experiences, agentic systems, models, evaluation frameworks, and data platforms that power Atlassian’s AI products.Working at AtlassianAtlassians... 
    Work at office
    Local area

    Atlassian

    San Francisco, CA
    4 days ago
  • Overview Atlassian is looking for a Senior Machine Learning Engineer to join our Search & Intelligence organization. Our team builds the intelligent...  ...frameworks, and data pipelines that power Atlassian’s AI products and accelerate AI innovation across the company.Working... 
    Work at office
    Local area
    Immediate start
    Worldwide

    Atlassian

    San Francisco, CA
    4 days ago
  • $163.42k - $285.98k

     ...work. Creating a career you love? It’s Possible.At Pinterest, AI isn't just a feature, it's a powerful partner that augments...  ...around the world and 300 billion ideas saved, Pinterest Machine Learning engineers build personalized experiences to help Pinners create a life... 
    Local area
    Relocation package

    Pinterest

    San Francisco, CA
    2 days ago
  • $228.96k - $315.36k

     ...The Fraud Data team at Plaid builds the machine learning systems that power Plaid’s fraud detection...  ...threats.As a Senior Machine Learning Engineer on Plaid's Fraud Data team, you will develop...  ....Explore how LLMs and Generative AI can improve fraud detection, prevention,... 
    Work experience placement
    Work at office
    Local area

    Plaid Financial

    San Francisco, CA
    1 day ago
  •  ...that the next era of teamwork will be shaped by people and AI agents working together, with Rovo as the intelligent teammate...  .... Responsibilities What you’ll doAs a Senior Machine Learning Engineer, you’ll tackle the open-ended technical problems that determine... 
    Work at office
    Local area

    Atlassian

    San Francisco, CA
    1 day ago
  • $200k - $300k

     ...technology company at the top of their industry is seeking a Machine Learning Engineer to join its research and development team. This full-time...  ...invent entirely new sensing capabilities, develop multimodal AI solutions, and influence the future roadmap of emerging health... 
    Full time
    Worldwide

    Motion Recruitment

    San Francisco, CA
    1 day ago
  • $228.96k - $315.36k

     ...within Plaid’s Fraud organization builds the machine learning systems that power Plaid’s fraud...  ...customers.As a Senior Machine Learning Engineer, you will own the development of high-performance...  ...capabilities, while leveraging AI-assisted tools to investigate complex system... 
    Work experience placement
    Work at office
    Local area

    Plaid Financial

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Machine Learning Engineer, Voice AI. Be the first to apply!