Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Engineer: Speech & LLMs

Full-time

Knowtex

About Knowtex

Knowtex is building the future of voice AI operating systems for clinicians, transforming how healthcare documentation happens at the point of care. We are experiencing rapid growth across both commercial health systems and federal healthcare, with our ambient documentation platform scaling to thousands of clinicians across hundreds of specialties.

We are at an inflection point where advances in speech, language models, and clinical AI can fundamentally change how clinicians interact with technology, giving them more time to focus on what matters most: their patients.

Position Overview

We are hiring two ML Engineers / Researchers to help build the next generation of Knowtex's AI stack.

We are looking for researchers with deep expertise in one of two areas:

  • Speech & Audio: Build state-of-the-art medical speech-to-text systems using our large proprietary dataset of real-world clinical audio, with the goal of bringing more of our speech stack in-house.

  • Large Language Models: Develop and optimize models for clinical documentation and structured clinical reasoning, improving quality, cost, latency, and control.

You do not need to be an expert in both areas. We are looking for exceptional depth in either speech/audio modeling or LLMs.

These are research-heavy roles with a direct path to production. You will design experiments, build datasets and evaluation systems, train and fine-tune models, and work closely with engineering and clinical teams to deploy successful approaches at scale.

This role plays a central part in defining Knowtex's long-term ML strategy.

Key Responsibilities

Speech & Audio

  • Develop and train speech recognition models optimized for medical conversations across hundreds of specialties

  • Leverage Knowtex's large proprietary clinical audio dataset to train and fine-tune domain-specific speech models

  • Research approaches for improving medical terminology recognition, speaker attribution, punctuation, timestamps, and robustness across accents and clinical environments

  • Build rigorous speech evaluation frameworks beyond traditional WER, including medical terminology and clinically significant error measurement

  • Explore modern speech architectures, self-supervised learning, speech foundation models, and audio-language models

  • Optimize models for low-latency, real-time inference at production scale

Large Language Models

  • Develop and optimize models for generating high-quality clinical documentation, including SOAP notes and specialty-specific note formats

  • Build models for downstream clinical tasks such as medication extraction, orders, ICD-10 coding, E&M coding, patient visit summaries, and other structured clinical artifacts

  • Evaluate open-weight and proprietary model architectures and determine where fine-tuning, distillation, structured generation, or task-specific models can outperform general-purpose API-based approaches

  • Fine-tune and post-train models using Knowtex's proprietary clinical datasets

  • Develop rigorous evaluation frameworks for clinical accuracy, hallucinations, completeness, formatting, and clinician preferences

  • Research approaches for reducing inference cost and latency while maintaining or improving clinical quality

Across Both Tracks

  • Move quickly from idea → dataset → experiment → evaluation → production

  • Design experiments that clearly measure whether an approach improves real-world clinical outcomes

  • Build datasets, benchmarks, and evaluation infrastructure that make model improvements measurable and reproducible

  • Collaborate closely with clinicians, applied ML engineers, and platform engineers

  • Take successful research beyond prototypes and help deploy models into production

  • Balance model quality with latency, inference cost, reliability, and scalability

Required Qualifications

  • 2+ years of experience in machine learning research or ML engineering, with deep expertise in speech/audio modeling or large language models

  • Strong expertise in Python and PyTorch

  • Deep understanding of modern transformer architectures and model training techniques

  • Experience training, fine-tuning, or post-training large neural models

  • Strong experimental methodology and ability to independently design and execute research projects

  • Experience working with large-scale datasets and distributed training environments

  • Ability to translate research results into production systems

  • Strong understanding of model evaluation and benchmarking

  • Bachelor’s, Master’s, or PhD in Computer Science, Machine Learning, or a related technical field, or equivalent research experience

Preferred Qualifications

For Speech Researchers

  • Deep experience with automatic speech recognition (ASR)

  • Experience training or fine-tuning Whisper, Conformer, wav2vec, or similar speech architectures

  • Experience with large-scale audio datasets and speech data pipelines

  • Familiarity with speaker diarization, voice activity detection, streaming ASR, or audio-language models

  • Experience optimizing speech models for real-time inference

For LLM Researchers

  • Experience fine-tuning or post-training open-weight LLMs

  • Experience with supervised fine-tuning, distillation, preference optimization, or reinforcement learning

  • Experience building LLM evaluation systems and model benchmarks

  • Experience serving and optimizing open-weight models at scale

  • Experience with structured generation, tool use, or agentic systems

For Either Track

  • Experience in healthcare AI, clinical NLP, or medical speech

  • Familiarity with clinical documentation workflows and medical terminology

  • Knowledge of coding systems such as ICD-10, CPT, E&M, or SNOMED

  • Publications at leading ML, NLP, or speech conferences

  • Experience deploying ML systems in HIPAA-compliant or regulated environments

  • Experience working in fast-moving startup environments where researchers own projects from experimentation through production

Technical Environment

  • AWS

  • Python, PyTorch

  • Transformer-based LLM and speech architectures

  • Open-weight and frontier language models

  • Large-scale clinical audio and text datasets

  • Distributed model training and inference

  • GPU-based model serving and optimization

  • Real-time speech and clinical AI pipelines

  • Structured clinical evaluation and benchmarking infrastructure

Compensation & Benefits

  • Competitive salary

  • Meaningful equity compensation

  • Unlimited PTO

  • Premium health, dental, and vision coverage

  • 401(k) plan

  • Work model: Hybrid In-person

Vacancy posted 13 hours ago
Similar jobs that could be interesting for youBased on the ML Engineer: Speech & LLMs in San Francisco, CA vacancy
  • Deloitte is seeking a Global Data ML Engineer for Multilingual Speech AI in the San Francisco area to design and manage scalable data pipelines on AWS and Snowflake. You will transform source data into analytics-ready datasets and build robust ETL/ELT processes across teams... 
    Suggested

    Cartesia, Inc.

    San Francisco, CA
    3 days ago
  • $102.75k - $171.25k

    Global Data Ml Engineer For Multilingual Speech AI Are you an experienced, passionate pioneer in technology who wants to work in a collaborative environment? As an experienced Data Engineer you will have the ability to share new ideas and collaborate on projects as a consultant... 
    Suggested
    Visa sponsorship

    Cartesia

    San Francisco, CA
    3 days ago
  • This role is a combination of research and engineering. We are looking for someone who's a talented software engineer at their core, but has...  ...to AI research, especially in the field of RAG, agents, and LLMs. Role Build AI powered product features Evaluate and enhance the... 
    Suggested
    Full time

    Alldus

    San Francisco, CA
    1 day ago
  • Pinterest Labs is hiring for an advanced ML researcher to push the boundaries of machine learning and multi-modal LLMs across Pinterest's platform. You will work with a world-class team of researchers and engineers to translate research into practical, impactful solutions... 
    Suggested

    Pinterest

    San Francisco, CA
    1 day ago
  • $180k - $270k

     ...deploying high-throughput, ultra-low-latency inference engines for large language models or foundational speech models. Understand the intricate tradeoffs...  ...will sit at the critical intersection between the core ML training team and the backend infrastructure team.... 
    Suggested
    Full time
    Work at office
    Worldwide

    Plaud

    San Francisco, CA
    13 hours ago
  • AI Talent Now, LLC in San Francisco, CA, is seeking a Senior Machine Learning Research Engineer to own the full lifecycle of ML research and development for speech and audio models. You will design, train, deploy, and fine-tune state-of-the-art systems with end-to-end... 
    Work at office

    AI Talent Now

    San Francisco, CA
    4 days ago
  • YO AI Labs seeks an experienced AI/ML Engineer to build and deploy secure, scalable AI solutions for mission-critical initiatives while contributing...  ...to proprietary AI infrastructure. You will work with LLMs, RAG, prompt engineering, multi-agent systems, and cloud AI platforms... 
    Remote job

    YO AI Labs

    San Francisco, CA
    4 days ago
  •  ...Training to build systems that transform powerful models. The ideal candidate has a deep understanding of machine learning and strong engineering skills. The role involves collaborating with teams to enhance model capabilities and improve model behavior. This full-time... 
    Full time

    Reflection AI

    San Francisco, CA
    2 days ago
  •  ...design, build, and automate offline and live evals that keep our speech and multimodal models honest in production. Harness the...  ...Qualifications: Expert-level PyTorch. Proven software engineer who loves ML; comfortable writing production code across the stack.... 
    Full time
    Contract work
    Flexible hours
    Shift work

    Sesame, L.l.c.

    San Francisco, CA
    13 hours ago
  •  ...ll work at the intersection of speech, language, and intelligence,...  ...huge impact as one of our first ML hires, shaping not only the...  ...Collaborate with product and engineering teams to integrate and deploy...  ...processing. Familiarity with LLMs, generative AI, or real-time inference... 
    Full time
    Worldwide
    Shift work

    HappyRobot

    San Francisco, CA
    13 hours ago
  • YO AI Labs is seeking an experienced AI/ML Engineer to build and deploy secure, scalable AI solutions for mission-critical initiatives. You will work with LLMs, RAG, prompt engineering, multi-agent systems, and cloud AI platforms to develop production-grade machine learning... 
    Remote job

    YO AI Labs

    San Francisco, CA
    4 days ago
  • B Capital is looking for a talented engineer to join their team in San Francisco, focusing on building systems that transform powerful pre...  ...driving innovative research, and significantly contributing to ML capabilities. We offer top-tier compensation and robust health benefits... 

    B Capital

    San Francisco, CA
    4 days ago
  • $200k - $260k

     ...voice agents and applications — serving speech-to-text and text-to-speech models with best...  .... We're looking for a Senior ML Engineer to drive the model serving layer for voice...  ...the field evolves, including audio-native LLMs, codec-based models (SNAC), and speech-to... 
    Full time

    Together Ai

    San Francisco, CA
    13 hours ago
  • $225k - $300k

     ...Role: As a Senior Machine Learning Engineer at Ambience , you will build and...  ...Distill insights from recent research in LLMs, agents, NLP, speech, and multimodal AI and translate promising...  ...AI Experience 5+ years in production ML, research engineering, or applied AI.... 
    Full time
    Work at office
    Immediate start
    Remote work
    Flexible hours
    3 days per week

    Ambience Healthcare

    San Francisco, CA
    13 hours ago
  •  ...instead of type anywhere on your computer or phone, and turns messy speech into polished text that is ready to send. It's scaled from $2M...  ...that every person shapes what gets built. About the Role As a ML engineer at Wispr, you’ll play a crucial role in building the first... 
    H1b
    Work at office
    Remote work
    Relocation
    Visa sponsorship
    Flexible hours

    Visa Hunt

    San Francisco, CA
    2 days ago
  • $195k - $365k

     ...building and training large-scale audio or speech models from the ground up, whether that...  ...living at the intersection of research and engineering, eager to design novel sequence modeling...  ...May Also Have Experience With Text-based LLMs: Hands‑on experience with core text‑based... 
    Full time
    Work at office
    Worldwide

    Plaud

    San Francisco, CA
    3 days ago
  • $200k - $235k

     ...managers, data scientists, software engineers, fraud intelligence, and...  ...Together, you'll design and build ML solutions that have direct,...  ...and safety use cases, and use LLMs and AI agents to accelerate...  ...models (vision, document, or speech) is a plus. ~ Exposure to... 
    Work experience placement
    Casual work
    Live in
    Work at office
    Remote work

    Airbnb

    San Francisco, CA
    1 day ago
  •  ...engagement, working directly with the engineer who leads our embedded...  ...modality (computer vision, ADAS, LLMs, audio). Building applications...  ...embedded-device, or on-device ML context, and the instinct to...  ...Nice-to-have: Audio or speech model experience. Function... 
    Full time
    Internship
    Shift work

    Liquid Ai Inc

    San Francisco, CA
    13 hours ago
  • $150k - $225k

     ...re looking for an Founding AI Engineer to design, build, and scale...  ...and experiment with multiple LLMs and APIs, choosing the right tool...  ...building and scaling data or ML pipelines in production environments...  ...modal models (text, image, or speech) Knowledge of GCP, Firebase,... 
    Full time
    Work from home
    Flexible hours

    Hellobabs

    San Francisco, CA
    3 days ago
  • $170k - $280k

     ...Our mission is to give leaders clarity and engineers time. We help leaders understand how...  ...the role We're looking for an Applied ML Engineer to help build and improve the machine...  ...learning or reinforcement learning for LLMs (RLHF, RLAIF, GRPO, PPO, DPO, or similar... 
    Odd job
    Full time

    Macroscope Inc

    San Francisco, CA
    13 hours ago
  • $220k - $280k

     ...voice agents and applications — serving speech-to-text and text-to-speech models with best...  .... We're looking for a Staff ML Engineer to drive the model serving layer for voice...  ...emerging model paradigms — audio-native LLMs, codec-based architectures (SNAC, Encodec... 
    Full time

    Together Ai

    San Francisco, CA
    13 hours ago
  • Design the intelligence and memory layer. LLMs, embeddings, and systems thinking required. Kodezi isn’t just another dev tool. It’s...  ..., evolves, and documents codebases autonomously. As a Founding ML Engineer , you’ll architect the intelligence powering autonomous pull requests... 
    Remote work
    Flexible hours

    Kodezi Inc.

    San Francisco, CA
    3 days ago
  • $150k - $300k

     ...records from across the web. You will own the ML systems that turn that raw, multilingual,...  ..., classifiers, and embedding models. Use LLMs for structured extraction, classification...  ...dilution. Own the full ML research and engineering cycle, from prototype to production.... 

    Open Select

    San Francisco, CA
    1 day ago
  •  ...voice interface to reach billions of users. We are hiring an ML Engineer to prototype and ship features for our voice interface and to...  ...across a global user base. You will help scale personalization of speech models using fine-tuning and RL, collaborating with a small,... 

    Visa Hunt

    San Francisco, CA
    3 days ago
  • At Dynamo AI, we believe that LLMs must be developed with safety, privacy, and real-world responsibility in mind. Our ML team comes from a culture of academic research driven to democratize AI advancements responsibly. By operating at the intersection of ML research and... 
    Local area
    Shift work

    Capitolis

    San Francisco, CA
    2 days ago
  • About The Role Skills: Python, PyTorch, NLP, LLMs, Information Retrieval, Entity Resolution, Text Classification...  ...someone who can push the boundaries of what our ML systems can do. We're hiring a Founding ML Engineer to own the research and engineering behind our core... 

    Crustdata (YC F24)

    San Francisco, CA
    2 days ago
  •  ...Design, build, and deploy production‑grade ML systems with end‑to‑end ownership of the...  ...AI‑powered solutions enabling natural speech interaction and real‑time audio understanding...  ...6 years of professional experience in ML engineering. Strong programming skills in Python (... 
    Full time

    Catalyst Labs, LLC

    San Francisco, CA
    5 days ago
  • $197.3k - $313.7k

     ...Slack is looking for a Staff Machine Learning Engineer with deep expertise in model training and finetuning to join our ML team. You'll design, train, and ship NLP models...  ...in NLP (or a closely related domain like speech, IR, or multimodal).5+ years of experience with... 
    Full time

    Salesforce

    San Francisco, CA
    1 day ago
  • Abacus is hiring its first Founding Engineer to own and expand the core machine learning engine behind Abacus. You’ll report...  ...direction. In your first 3-6 months, you’ll scale the ML extraction engine using LLMs and OCR, own backend architecture, and ship greenfield features... 
    Work at office

    Praxis

    San Francisco, CA
    1 day ago
  • Position: Senior ML Performance Engineer Location: SF Bay Area (US) or Toronto (Canada) - Hybrid Employment Type: Full-Time Industry: AI Infrastructure...  ...and optimizing the performance of large language models (LLMs) before and after compiler optimization on modern GPU... 
    Full time

    Amadeus Search

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Engineer: Speech & LLMs. Be the first to apply!