Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Engineer: Speech & LLMs

Full-time

Knowtex

About Knowtex

Knowtex is building the future of voice AI operating systems for clinicians, transforming how healthcare documentation happens at the point of care. We are experiencing rapid growth across both commercial health systems and federal healthcare, with our ambient documentation platform scaling to thousands of clinicians across hundreds of specialties.

We are at an inflection point where advances in speech, language models, and clinical AI can fundamentally change how clinicians interact with technology, giving them more time to focus on what matters most: their patients.

Position Overview

We are hiring two ML Engineers / Researchers to help build the next generation of Knowtex's AI stack.

We are looking for researchers with deep expertise in one of two areas:

  • Speech & Audio: Build state-of-the-art medical speech-to-text systems using our large proprietary dataset of real-world clinical audio, with the goal of bringing more of our speech stack in-house.

  • Large Language Models: Develop and optimize models for clinical documentation and structured clinical reasoning, improving quality, cost, latency, and control.

You do not need to be an expert in both areas. We are looking for exceptional depth in either speech/audio modeling or LLMs.

These are research-heavy roles with a direct path to production. You will design experiments, build datasets and evaluation systems, train and fine-tune models, and work closely with engineering and clinical teams to deploy successful approaches at scale.

This role plays a central part in defining Knowtex's long-term ML strategy.

Key Responsibilities

Speech & Audio

  • Develop and train speech recognition models optimized for medical conversations across hundreds of specialties

  • Leverage Knowtex's large proprietary clinical audio dataset to train and fine-tune domain-specific speech models

  • Research approaches for improving medical terminology recognition, speaker attribution, punctuation, timestamps, and robustness across accents and clinical environments

  • Build rigorous speech evaluation frameworks beyond traditional WER, including medical terminology and clinically significant error measurement

  • Explore modern speech architectures, self-supervised learning, speech foundation models, and audio-language models

  • Optimize models for low-latency, real-time inference at production scale

Large Language Models

  • Develop and optimize models for generating high-quality clinical documentation, including SOAP notes and specialty-specific note formats

  • Build models for downstream clinical tasks such as medication extraction, orders, ICD-10 coding, E&M coding, patient visit summaries, and other structured clinical artifacts

  • Evaluate open-weight and proprietary model architectures and determine where fine-tuning, distillation, structured generation, or task-specific models can outperform general-purpose API-based approaches

  • Fine-tune and post-train models using Knowtex's proprietary clinical datasets

  • Develop rigorous evaluation frameworks for clinical accuracy, hallucinations, completeness, formatting, and clinician preferences

  • Research approaches for reducing inference cost and latency while maintaining or improving clinical quality

Across Both Tracks

  • Move quickly from idea → dataset → experiment → evaluation → production

  • Design experiments that clearly measure whether an approach improves real-world clinical outcomes

  • Build datasets, benchmarks, and evaluation infrastructure that make model improvements measurable and reproducible

  • Collaborate closely with clinicians, applied ML engineers, and platform engineers

  • Take successful research beyond prototypes and help deploy models into production

  • Balance model quality with latency, inference cost, reliability, and scalability

Required Qualifications

  • 2+ years of experience in machine learning research or ML engineering, with deep expertise in speech/audio modeling or large language models

  • Strong expertise in Python and PyTorch

  • Deep understanding of modern transformer architectures and model training techniques

  • Experience training, fine-tuning, or post-training large neural models

  • Strong experimental methodology and ability to independently design and execute research projects

  • Experience working with large-scale datasets and distributed training environments

  • Ability to translate research results into production systems

  • Strong understanding of model evaluation and benchmarking

  • Bachelor’s, Master’s, or PhD in Computer Science, Machine Learning, or a related technical field, or equivalent research experience

Preferred Qualifications

For Speech Researchers

  • Deep experience with automatic speech recognition (ASR)

  • Experience training or fine-tuning Whisper, Conformer, wav2vec, or similar speech architectures

  • Experience with large-scale audio datasets and speech data pipelines

  • Familiarity with speaker diarization, voice activity detection, streaming ASR, or audio-language models

  • Experience optimizing speech models for real-time inference

For LLM Researchers

  • Experience fine-tuning or post-training open-weight LLMs

  • Experience with supervised fine-tuning, distillation, preference optimization, or reinforcement learning

  • Experience building LLM evaluation systems and model benchmarks

  • Experience serving and optimizing open-weight models at scale

  • Experience with structured generation, tool use, or agentic systems

For Either Track

  • Experience in healthcare AI, clinical NLP, or medical speech

  • Familiarity with clinical documentation workflows and medical terminology

  • Knowledge of coding systems such as ICD-10, CPT, E&M, or SNOMED

  • Publications at leading ML, NLP, or speech conferences

  • Experience deploying ML systems in HIPAA-compliant or regulated environments

  • Experience working in fast-moving startup environments where researchers own projects from experimentation through production

Technical Environment

  • AWS

  • Python, PyTorch

  • Transformer-based LLM and speech architectures

  • Open-weight and frontier language models

  • Large-scale clinical audio and text datasets

  • Distributed model training and inference

  • GPU-based model serving and optimization

  • Real-time speech and clinical AI pipelines

  • Structured clinical evaluation and benchmarking infrastructure

Compensation & Benefits

  • Competitive salary

  • Meaningful equity compensation

  • Unlimited PTO

  • Premium health, dental, and vision coverage

  • 401(k) plan

  • Work model: Hybrid In-person

Vacancy posted 7 hours ago
Similar jobs that could be interesting for youBased on the ML Engineer: Speech & LLMs in San Francisco, CA vacancy
  •  ...We are looking for an experienced Machine Learning Engineer to join our team and help develop cutting-edge speech recognition models that help teach language fluency...  ...more. This is an incredibly exciting time to join an ML team designing a personalized learning experience... 
    Suggested
    Full time
    Live in
    Work at office
    Worldwide

    Speak

    San Francisco, CA
    19 hours ago
  • $195k - $365k

     ...building and training large-scale audio or speech models from the ground up, whether that...  ...living at the intersection of research and engineering, eager to design novel sequence modeling...  ...May Also Have Experience With Text-based LLMs: Hands‑on experience with core text‑based... 
    Suggested
    Full time
    Work at office
    Worldwide

    Plaud

    San Francisco, CA
    31 minutes ago
  •  ...This role is a combination of research and engineering. We are looking for someone who's a talented software engineer at their core, but...  ...contributed to AI research, especially in the field of RAG, agents, and LLMs. Role Build AI powered product features Evaluate and... 
    Suggested
    Full time

    Alldus

    San Francisco, CA
    35 minutes ago
  • Pinterest Labs is hiring for an advanced ML researcher to push the boundaries of machine learning and multi-modal LLMs across Pinterest's platform. You will work with a world-class team of researchers and engineers to translate research into practical, impactful solutions... 
    Suggested

    Pinterest

    San Francisco, CA
    5 days ago
  • $180k - $270k

     ...deploying high-throughput, ultra-low-latency inference engines for large language models or foundational speech models. Understand the intricate tradeoffs...  ...will sit at the critical intersection between the core ML training team and the backend infrastructure team.... 
    Suggested
    Full time
    Work at office
    Worldwide

    Plaud

    San Francisco, CA
    19 hours ago
  • $200k - $240k

     ...safer, more secure world for all. The AI Engineering Team is chartered with enabling next-generation...  ...special focus on Large Language Models (LLMs) and agentic systems. Our mission is to...  ...than the market. As a Senior or Staff ML Systems Engineer - LLM , you’ll be at the... 
    Remote work
    Worldwide

    TRM Labs

    San Francisco, CA
    5 days ago
  •  ...with hands-on support from AMD engineers the team is scaling rapidly to...  .... About the role As an ML Engineer at Sciforium, you will...  ...architectures that combine LLMs, retrieval systems, memory, tools...  ...domains (e.g., NLP, vision, speech, generative models). Communication... 
    Full time
    Flexible hours

    Sciforium

    San Francisco, CA
    2 days ago
  •  ...design, build, and automate offline and live evals that keep our speech and multimodal models honest in production. Harness the...  ...Qualifications: Expert-level PyTorch. Proven software engineer who loves ML; comfortable writing production code across the stack.... 
    Full time
    Contract work
    Flexible hours
    Shift work

    Sesame, L.l.c.

    San Francisco, CA
    19 hours ago
  •  ...ll work at the intersection of speech, language, and intelligence,...  ...huge impact as one of our first ML hires, shaping not only the...  ...Collaborate with product and engineering teams to integrate and deploy...  ...processing. Familiarity with LLMs, generative AI, or real-time inference... 
    Full time
    Worldwide
    Shift work

    HappyRobot

    San Francisco, CA
    19 hours ago
  •  ...Training to build systems that transform powerful models. The ideal candidate has a deep understanding of machine learning and strong engineering skills. The role involves collaborating with teams to enhance model capabilities and improve model behavior. This full-time... 
    Full time

    Reflection AI

    San Francisco, CA
    1 day ago
  • B Capital is looking for a talented engineer to join their team in San Francisco, focusing on building systems that transform powerful pre...  ...driving innovative research, and significantly contributing to ML capabilities. We offer top-tier compensation and robust health benefits... 

    B Capital

    San Francisco, CA
    3 days ago
  • $200k - $260k

     ...voice agents and applications — serving speech-to-text and text-to-speech models with best...  .... We're looking for a Senior ML Engineer to drive the model serving layer for voice...  ...the field evolves, including audio-native LLMs, codec-based models (SNAC), and speech-to... 
    Full time

    Together Ai

    San Francisco, CA
    19 hours ago
  •  ...The Role As a Senior Machine Learning Engineer, you will build the intelligence layer that...  ...automation. You will work across LLMs, OCR pipelines, voice AI, evaluation systems...  .... Familiarity with telephony vendors, speech systems, or conversational agent infrastructure... 
    Full time
    Work at office

    Hike Medical

    San Francisco, CA
    19 hours ago
  • $225k - $300k

     ...Role: As a Senior Machine Learning Engineer at Ambience , you will build and...  ...Distill insights from recent research in LLMs, agents, NLP, speech, and multimodal AI and translate promising...  ...AI Experience 5+ years in production ML, research engineering, or applied AI.... 
    Full time
    Work at office
    Immediate start
    Remote work
    Flexible hours
    3 days per week

    Ambience Healthcare

    San Francisco, CA
    19 hours ago
  •  ...ll work at the intersection of speech, language, and intelligence,...  ...huge impact as one of our first ML hires, shaping not only the...  ...Collaborate with product and engineering teams to integrate and deploy...  ...processing. Familiarity with LLMs, generative AI, or real‑time inference... 
    Worldwide
    Shift work

    Happy Robot

    San Francisco, CA
    1 day ago
  •  ...Check out 1962 new Machine Learning Engineer opportunities posted on AI Chopping Block Design...  ...deployment. Develop and optimize end-to-end ML pipelines encompassing data collection,...  ...be familiar with the state of the art in LLMs, RL, and code generation. Develop methods... 
    Flexible hours

    AI Chopping Block, Inc.

    San Francisco, CA
    3 days ago
  • $200k - $235k

     ...managers, data scientists, software engineers, fraud intelligence, and...  ...Together, you'll design and build ML solutions that have direct,...  ...and safety use cases, and use LLMs and AI agents to accelerate...  ...models (vision, document, or speech) is a plus. ~ Exposure to... 
    Work experience placement
    Casual work
    Live in
    Work at office
    Remote work

    Airbnb

    San Francisco, CA
    4 days ago
  • Stealth Startup in San Francisco seeks a Founding Machine Learning Research Engineer to advance real-time AI by exploring state-of-the-art LLMs, speech models, and multimodal AI for human-like voice agents operating in complex environments. You’ll design evaluation frameworks... 

    Stealth Startup

    San Francisco, CA
    4 days ago
  •  ...engagement, working directly with the engineer who leads our embedded...  ...modality (computer vision, ADAS, LLMs, audio). Building applications...  ...embedded-device, or on-device ML context, and the instinct to...  ...Nice-to-have: Audio or speech model experience. Function... 
    Full time
    Internship
    Shift work

    Liquid Ai Inc

    San Francisco, CA
    19 hours ago
  • $220k - $280k

     ...voice agents and applications — serving speech-to-text and text-to-speech models with best...  .... We're looking for a Staff ML Engineer to drive the model serving layer for voice...  ...emerging model paradigms — audio-native LLMs, codec-based architectures (SNAC, Encodec... 
    Full time

    Together Ai

    San Francisco, CA
    19 hours ago
  • $170k - $280k

     ...Our mission is to give leaders clarity and engineers time. We help leaders understand how...  ...the role We're looking for an Applied ML Engineer to help build and improve the machine...  ...learning or reinforcement learning for LLMs (RLHF, RLAIF, GRPO, PPO, DPO, or similar... 
    Odd job
    Full time

    Macroscope Inc

    San Francisco, CA
    19 hours ago
  •  ...Clay is seeking a Machine Learning Engineer to join the Learning Team and help build the intelligence...  ...learning-driven features, design scalable ML systems, and collaborate across product...  ...engineering experience, familiarity with LLMs in production, and a passion for AI... 

    Sapphire Partners

    San Francisco, CA
    31 minutes ago
  • $197.3k - $313.7k

     ...Slack is looking for a Staff Machine Learning Engineer with deep expertise in model training and finetuning to join our ML team. You'll design, train, and ship NLP models...  ...in NLP (or a closely related domain like speech, IR, or multimodal).5+ years of experience with... 
    Full time

    Salesforce

    San Francisco, CA
    19 hours ago
  • Design the intelligence and memory layer. LLMs, embeddings, and systems thinking required. Kodezi isn’t just another dev tool. It’s...  ..., evolves, and documents codebases autonomously. As a Founding ML Engineer , you’ll architect the intelligence powering autonomous pull requests... 
    Remote work
    Flexible hours

    Kodezi Inc.

    San Francisco, CA
    2 days ago
  •  ...instead of type anywhere on your computer or phone, and turns messy speech into polished text that is ready to send. It's scaled from $2M...  ...every person shapes what gets built. About the Role As a ML engineer at Wispr, you’ll play a crucial role in building the first... 
    H1b
    Work at office
    Remote work
    Relocation
    Visa sponsorship
    Flexible hours

    Visa Hunt

    San Francisco, CA
    2 days ago
  • At Dynamo AI, we believe that LLMs must be developed with safety, privacy, and real-world responsibility in mind. Our ML team comes from a culture of academic research driven to democratize AI advancements responsibly. By operating at the intersection of ML research and... 
    Local area
    Shift work

    Capitolis

    San Francisco, CA
    1 day ago
  •  ...voice interface to reach billions of users. We are hiring an ML Engineer to prototype and ship features for our voice interface and to...  ...across a global user base. You will help scale personalization of speech models using fine-tuning and RL, collaborating with a small,... 

    Visa Hunt

    San Francisco, CA
    2 days ago
  •  ...Physics | 5 Days Onsite Machine Learning Engineer Location: Onsite in San Francisco Compensation...  ...About the Role UniversalAGI is hiring an ML Engineer to help ship ML outcomes by...  ...or fine-tuning models (any modality/type - LLMs, computer vision, physics, surrogate models... 
    Work at office
    Flexible hours
    1 day per week

    UniversalAGI

    San Francisco, CA
    2 days ago
  •  ...Design, build, and deploy production‑grade ML systems with end‑to‑end ownership of the...  ...AI‑powered solutions enabling natural speech interaction and real‑time audio understanding...  ...6 years of professional experience in ML engineering. Strong programming skills in Python (... 
    Full time

    Catalyst Labs, LLC

    San Francisco, CA
    4 days ago
  • About The Role Skills: Python, PyTorch, NLP, LLMs, Information Retrieval, Entity Resolution, Text Classification...  ...someone who can push the boundaries of what our ML systems can do. We're hiring a Founding ML Engineer to own the research and engineering behind our core... 

    Crustdata (YC F24)

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Engineer: Speech & LLMs. Be the first to apply!