Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff ML Engineer Real-Time Voice Inference Architect

Together

Together AI is building the best inference infrastructure for voice applications. We seek a Staff ML Engineer to own the model serving stack and optimize latency and throughput for real-time voice workloads. You'll work with state-of-the-art accelerators and collaborate with model partners to bring models to production on Together's platform. This is a foundational role on a small, high-impact team. #J-18808-Ljbffr Together

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Staff ML Engineer Real-Time Voice Inference Architect in San Francisco, CA vacancy
  • $200k - $280k

     ...A leading voice AI startup is seeking a Founding Senior Machine Learning Engineer to fine-tune and deploy human-like voice agents, handling millions of real-time calls. This role offers a salary range of $20...  ...world experience in deploying ML models and work well in a... 
    Suggested

    Retell AI

    San Francisco, CA
    5 days ago
  • Bluejay is hiring a Member of Technical Staff in San Francisco to architect and build systems that simulate,...  ...evaluate conversational AI agents across voice and multimodal domains. You’ll own...  ...on a high-visibility project with real-time AWS infrastructure for millions of... 
    Suggested

    davidjoseph-co

    San Francisco, CA
    1 day ago
  •  ...is seeking a Senior Member of Technical Staff to architect and develop systems that simulate,...  ...evaluate conversational AI agents across voice and multimodal systems. You will design...  ...AWS infrastructure to handle millions of real-time conversations on a small team with full... 
    Suggested

    davidjoseph-co

    San Francisco, CA
    1 day ago
  • Shipt is seeking a Staff Machine Learning Engineer on the Personalization Platform team to drive...  ...boost user engagement. You will architect and implement scalable ML infrastructure, design data pipelines...  ...team to deliver low-latency, real-time or batch recommendations with... 
    Suggested

    Shipt

    San Francisco, CA
    2 days ago
  • $220k - $280k

     ...AI is building the best inference infrastructure for voice applications. Our Voice...  ...powers production-grade, real-time voice agents and applications...  ...We're looking for a Staff ML Engineer to drive the model...  ...inference performance — architect and implement systems targeting... 
    Suggested
    Full time

    Together Ai

    San Francisco, CA
    10 hours ago
  • Parafin, Inc. is seeking a Software Engineer to lead the evolution of its ML Platform within the Infrastructure...  ..., training, evaluation, inference, and retraining powering underwriting...  ...end-to-end, supporting batch and real-time underwriting infrastructure in collaboration... 

    Parafin, Inc.

    San Francisco, CA
    4 days ago
  • Kindredventures is hiring software engineers to own the path from trained model to customer...  ...production systems that deliver predictions in real time across cloud, VPC, on‑prem, and...  ...strong generalist with experience deploying ML systems in production, comfortable working... 

    Kindredventures

    San Francisco, CA
    10 hours ago
  • David Joseph & Company is seeking a Member of Technical Staff to architect and build systems that simulate, analyze, and evaluate conversational...  ...to-end features while collaborating with major customers on real-time deployments. The role is fully on-site in San Francisco,... 

    David Joseph & Company

    San Francisco, CA
    1 day ago
  • $300k - $375k

     ...leading conversational AI platform in California seeks a Staff Software Engineer focused on Voice Agent. This role involves owning the architecture of...  ...engineering experience with technical leadership in real-time systems. Highly collaborative environment that offers... 
    Work at office

    Decagon

    San Francisco, CA
    4 days ago
  • $203.5k - $299.3k

     ...a Causal Machine Learning Engineer to help build the causal ML foundation behind how DoorDash...  ...ML systems that influence real marketplace decisions...  ...practical experience with causal inference, econometrics,...  ...commuter benefits match, paid time off and paid sick leave in... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    3 days ago
  • Bluejay is building the simulation, observability, and self-improvement layer for Voice AI agents. The role focuses on architecting and delivering scalable infrastructure on AWS to support millions of concurrent conversations while developing advanced AI/evaluation tooling... 
    Work at office

    Bluejay

    San Francisco, CA
    3 days ago
  • $203.5k - $299.3k

     ...a Causal Machine Learning Engineer to help build the causal ML foundation behind how DoorDash...  ...ML systems that influence real marketplace decisions...  ...practical experience with causal inference, econometrics,...  ...commuter benefits match, paid time off and paid sick leave in... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    DoorDash

    San Francisco, CA
    4 days ago
  •  ...a Causal Machine Learning Engineer to help build the causal ML foundation behind how DoorDash...  ...ML systems that influence real marketplace decisions...  ...practical experience with causal inference, econometrics,...  ...commuter benefits match, paid time off and paid sick leave in... 
    Hourly pay
    Work at office
    Local area
    Flexible hours

    DoorDash, Inc.

    San Francisco, CA
    2 days ago
  •  ...hiring for a role focused on the StreamModule pipeline that powers voice AI. You will be responsible for ensuring the pipeline's...  ...architecture. The ideal candidate should have extensive experience with real-time systems and queue architectures like BullMQ and Kafka, along... 

    VAPI

    San Francisco, CA
    4 days ago
  •  ...leading AI evaluation platform in San Francisco is looking for a Senior Software Engineer specializing in ML infrastructure. The successful candidate will design and develop robust real-time data and API systems, enabling insights for researchers and developers. Ideal... 

    LMArena

    San Francisco, CA
    2 days ago
  • $197.3k - $313.7k

     ...Slack is looking for a Staff Machine Learning Engineer with deep expertise...  ...finetuning to join our ML team. You'll design,...  ...and moving on. Other times that means developing...  ...that serve real users at scale — not...  ...model optimization for inference (quantization, pruning... 
    Full time

    Salesforce

    San Francisco, CA
    3 days ago
  •  ...the Waymo Driver to improve mobility and save lives. The role focuses on designing and scaling real-time fleet monitoring and anomaly detection systems, with production-grade ML infrastructure and collaboration with data scientists. The candidate will work with Java/C++... 

    Neura Market

    San Francisco, CA
    10 hours ago
  • $200k - $260k

     ...Together AI is building the best inference infrastructure for voice applications. Our Voice AI platform powers production-grade, real-time voice agents and applications — serving...  ...reliability. We're looking for a Senior ML Engineer to drive the model serving layer for... 
    Full time

    Together Ai

    San Francisco, CA
    10 hours ago
  •  ...decisions and operate across enterprise systems. We seek a senior ML engineer to own end-to-end ML lifecycle—from data ingestion to...  ...intelligence. You’ll help design robust, scalable pipelines and ship real-time models powering our core products. Based in San Francisco,... 

    Happy Robot

    San Francisco, CA
    4 days ago
  • Orb is building the runtime layer for usage-based billing, entitlements, and real-time analytics. You will own a surface end to end, from customer conversations to API design and rollout, shaping primitives that other teams rely on. You will work across the product, design... 

    Orb Staging

    San Francisco, CA
    10 hours ago
  • $155k - $180k

     ...not only product and engineering), so Roboflow...  ...of all of this is inference — one of our most...  ...that validates the real health of every build...  ...computer vision and ML models to our users...  ...being a public-facing voice for a project. ~...  ...hubs more than 3+ times per week. In addition... 
    Full time
    Second job
    Remote work
    Work from home
    Relocation package
    Flexible hours
    Night shift

    Roboflow

    San Francisco, CA
    10 hours ago
  • Jaide Health is seeking an engineer for their Model Efficiency team in San Francisco. The role focuses on building reliable ML systems while enhancing core performance metrics across...  ...C++ or Python and insights into the LLM inference ecosystem. A commitment to diversity and... 
    Remote job

    Jaide Health

    San Francisco, CA
    4 days ago
  • A pioneering research lab is seeking a Senior Software Engineer focused on building next-generation conversational AI interfaces. The...  ...collaboration with cross-functional teams to create scalable real-time interfaces. Join a well-funded startup at the forefront of the... 

    Jack & Jill

    San Francisco, CA
    4 days ago
  • $137.1k - $201.6k

     ...will leverage AI and advanced ML to power decision making in real-time – from personalized sign up promotions...  ...Role We’re looking for a Staff Machine Learning Engineer to drive the design and...  ...will: Contribute to Causal inference modeling to measure the incremental... 
    Hourly pay
    Full time
    Work at office
    Local area
    Remote work
    Flexible hours

    From Restaurants Near You

    San Francisco, CA
    10 hours ago
  • $190k - $222k

    Wheel the World in San Francisco seeks a Senior Machine Learning Engineer to lead the perception stack for their Archimedes project. This...  ...expertise in developing and deploying perception systems under real-world conditions. The successful candidate will have extensive experience... 

    Wheel the World

    San Francisco, CA
    2 days ago
  •  ...changing world around us in real time and make informed...  ...consequences on. As a Staff Machine Learning Engineer, you’ll own AI-driven products...  ...latency, high-concurrency inference (Triton, vLLM, GPU-backed...  ...of shipping and operating ML-driven functionality. ~... 
    Full time
    Contract work
    Remote work
    Flexible hours

    Primer.ai

    San Francisco, CA
    10 hours ago
  •  ...machine learning, and causal inference. We are looking for...  ...data scientists, data engineers and other MLEs to...  ...ionization of machine learning (ML) solutions for complex...  ...Flexible PTO - take time when you need it! Equity...  ...and celebrate in real life! Free Lunch – Grab... 
    Full time
    Work at office
    Work from home
    Worldwide
    Flexible hours

    Haus Analytics

    San Francisco, CA
    3 hours ago
  • $215k - $322k

     ...GoFundMe as our next Staff Machine Learning Engineer (Pricing) . In this role...  ...optimization (one-time and recurring), recurring...  ...building production ML systems (data → training → online inference → measurement) with rigorous...  ...Build low-latency real-time inferencing... 
    Full time
    Temporary work
    Work at office
    Flexible hours

    Gofundme

    San Francisco, CA
    10 hours ago
  •  ...technology company in San Francisco is seeking a Founding Engineer specializing in ML Inference. This highly technical role requires expertise in the...  .... The ideal candidate will drive innovations in real-time model performance, design in-house inference runtimes,... 
    Relocation package

    Reactor.am

    San Francisco, CA
    2 days ago
  • Whatnot is looking for a Software Engineer to join their Fraud Experience team. You'll lead the...  ...and a strong understanding of Python and ML libraries. Whatnot offers a flexible work...  ...inclusive workplace. Benefits include generous time off, health insurance options, and support... 
    Remote job
    Flexible hours

    Whatnot

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff ML Engineer Real-Time Voice Inference Architect. Be the first to apply!