ML Engineer: Speech & LLMs
Knowtex
About Knowtex
Knowtex is building the future of voice AI operating systems for clinicians, transforming how healthcare documentation happens at the point of care. We are experiencing rapid growth across both commercial health systems and federal healthcare, with our ambient documentation platform scaling to thousands of clinicians across hundreds of specialties.
We are at an inflection point where advances in speech, language models, and clinical AI can fundamentally change how clinicians interact with technology, giving them more time to focus on what matters most: their patients.
Position Overview
We are hiring two ML Engineers / Researchers to help build the next generation of Knowtex's AI stack.
We are looking for researchers with deep expertise in one of two areas:
Speech & Audio: Build state-of-the-art medical speech-to-text systems using our large proprietary dataset of real-world clinical audio, with the goal of bringing more of our speech stack in-house.
Large Language Models: Develop and optimize models for clinical documentation and structured clinical reasoning, improving quality, cost, latency, and control.
You do not need to be an expert in both areas. We are looking for exceptional depth in either speech/audio modeling or LLMs.
These are research-heavy roles with a direct path to production. You will design experiments, build datasets and evaluation systems, train and fine-tune models, and work closely with engineering and clinical teams to deploy successful approaches at scale.
This role plays a central part in defining Knowtex's long-term ML strategy.
Key Responsibilities
Speech & Audio
Develop and train speech recognition models optimized for medical conversations across hundreds of specialties
Leverage Knowtex's large proprietary clinical audio dataset to train and fine-tune domain-specific speech models
Research approaches for improving medical terminology recognition, speaker attribution, punctuation, timestamps, and robustness across accents and clinical environments
Build rigorous speech evaluation frameworks beyond traditional WER, including medical terminology and clinically significant error measurement
Explore modern speech architectures, self-supervised learning, speech foundation models, and audio-language models
Optimize models for low-latency, real-time inference at production scale
Large Language Models
Develop and optimize models for generating high-quality clinical documentation, including SOAP notes and specialty-specific note formats
Build models for downstream clinical tasks such as medication extraction, orders, ICD-10 coding, E&M coding, patient visit summaries, and other structured clinical artifacts
Evaluate open-weight and proprietary model architectures and determine where fine-tuning, distillation, structured generation, or task-specific models can outperform general-purpose API-based approaches
Fine-tune and post-train models using Knowtex's proprietary clinical datasets
Develop rigorous evaluation frameworks for clinical accuracy, hallucinations, completeness, formatting, and clinician preferences
Research approaches for reducing inference cost and latency while maintaining or improving clinical quality
Across Both Tracks
Move quickly from idea → dataset → experiment → evaluation → production
Design experiments that clearly measure whether an approach improves real-world clinical outcomes
Build datasets, benchmarks, and evaluation infrastructure that make model improvements measurable and reproducible
Collaborate closely with clinicians, applied ML engineers, and platform engineers
Take successful research beyond prototypes and help deploy models into production
Balance model quality with latency, inference cost, reliability, and scalability
Required Qualifications
2+ years of experience in machine learning research or ML engineering, with deep expertise in speech/audio modeling or large language models
Strong expertise in Python and PyTorch
Deep understanding of modern transformer architectures and model training techniques
Experience training, fine-tuning, or post-training large neural models
Strong experimental methodology and ability to independently design and execute research projects
Experience working with large-scale datasets and distributed training environments
Ability to translate research results into production systems
Strong understanding of model evaluation and benchmarking
Bachelor’s, Master’s, or PhD in Computer Science, Machine Learning, or a related technical field, or equivalent research experience
Preferred Qualifications
For Speech Researchers
Deep experience with automatic speech recognition (ASR)
Experience training or fine-tuning Whisper, Conformer, wav2vec, or similar speech architectures
Experience with large-scale audio datasets and speech data pipelines
Familiarity with speaker diarization, voice activity detection, streaming ASR, or audio-language models
Experience optimizing speech models for real-time inference
For LLM Researchers
Experience fine-tuning or post-training open-weight LLMs
Experience with supervised fine-tuning, distillation, preference optimization, or reinforcement learning
Experience building LLM evaluation systems and model benchmarks
Experience serving and optimizing open-weight models at scale
Experience with structured generation, tool use, or agentic systems
For Either Track
Experience in healthcare AI, clinical NLP, or medical speech
Familiarity with clinical documentation workflows and medical terminology
Knowledge of coding systems such as ICD-10, CPT, E&M, or SNOMED
Publications at leading ML, NLP, or speech conferences
Experience deploying ML systems in HIPAA-compliant or regulated environments
Experience working in fast-moving startup environments where researchers own projects from experimentation through production
Technical Environment
AWS
Python, PyTorch
Transformer-based LLM and speech architectures
Open-weight and frontier language models
Large-scale clinical audio and text datasets
Distributed model training and inference
GPU-based model serving and optimization
Real-time speech and clinical AI pipelines
Structured clinical evaluation and benchmarking infrastructure
Compensation & Benefits
Competitive salary
Meaningful equity compensation
Unlimited PTO
Premium health, dental, and vision coverage
401(k) plan
Work model: Hybrid In-person
- Deloitte is seeking a Global Data ML Engineer for Multilingual Speech AI in the San Francisco area to design and manage scalable data pipelines on AWS and Snowflake. You will transform source data into analytics-ready datasets and build robust ETL/ELT processes across teams...Suggested
$102.75k - $171.25k
Global Data Ml Engineer For Multilingual Speech AI Are you an experienced, passionate pioneer in technology who wants to work in a collaborative environment? As an experienced Data Engineer you will have the ability to share new ideas and collaborate on projects as a consultant...SuggestedVisa sponsorship- This role is a combination of research and engineering. We are looking for someone who's a talented software engineer at their core, but has... ...to AI research, especially in the field of RAG, agents, and LLMs. Role Build AI powered product features Evaluate and enhance the...SuggestedFull time
- Pinterest Labs is hiring for an advanced ML researcher to push the boundaries of machine learning and multi-modal LLMs across Pinterest's platform. You will work with a world-class team of researchers and engineers to translate research into practical, impactful solutions...Suggested
$180k - $270k
...deploying high-throughput, ultra-low-latency inference engines for large language models or foundational speech models. Understand the intricate tradeoffs... ...will sit at the critical intersection between the core ML training team and the backend infrastructure team....SuggestedFull timeWork at officeWorldwide- AI Talent Now, LLC in San Francisco, CA, is seeking a Senior Machine Learning Research Engineer to own the full lifecycle of ML research and development for speech and audio models. You will design, train, deploy, and fine-tune state-of-the-art systems with end-to-end...Work at office
- YO AI Labs seeks an experienced AI/ML Engineer to build and deploy secure, scalable AI solutions for mission-critical initiatives while contributing... ...to proprietary AI infrastructure. You will work with LLMs, RAG, prompt engineering, multi-agent systems, and cloud AI platforms...Remote job
- ...Training to build systems that transform powerful models. The ideal candidate has a deep understanding of machine learning and strong engineering skills. The role involves collaborating with teams to enhance model capabilities and improve model behavior. This full-time...Full time
- ...design, build, and automate offline and live evals that keep our speech and multimodal models honest in production. Harness the... ...Qualifications: Expert-level PyTorch. Proven software engineer who loves ML; comfortable writing production code across the stack....Full timeContract workFlexible hoursShift work
- ...ll work at the intersection of speech, language, and intelligence,... ...huge impact as one of our first ML hires, shaping not only the... ...Collaborate with product and engineering teams to integrate and deploy... ...processing. Familiarity with LLMs, generative AI, or real-time inference...Full timeWorldwideShift work
- YO AI Labs is seeking an experienced AI/ML Engineer to build and deploy secure, scalable AI solutions for mission-critical initiatives. You will work with LLMs, RAG, prompt engineering, multi-agent systems, and cloud AI platforms to develop production-grade machine learning...Remote job
- B Capital is looking for a talented engineer to join their team in San Francisco, focusing on building systems that transform powerful pre... ...driving innovative research, and significantly contributing to ML capabilities. We offer top-tier compensation and robust health benefits...
$200k - $260k
...voice agents and applications — serving speech-to-text and text-to-speech models with best... .... We're looking for a Senior ML Engineer to drive the model serving layer for voice... ...the field evolves, including audio-native LLMs, codec-based models (SNAC), and speech-to...Full time$225k - $300k
...Role: As a Senior Machine Learning Engineer at Ambience , you will build and... ...Distill insights from recent research in LLMs, agents, NLP, speech, and multimodal AI and translate promising... ...AI Experience 5+ years in production ML, research engineering, or applied AI....Full timeWork at officeImmediate startRemote workFlexible hours3 days per week- ...instead of type anywhere on your computer or phone, and turns messy speech into polished text that is ready to send. It's scaled from $2M... ...that every person shapes what gets built. About the Role As a ML engineer at Wispr, you’ll play a crucial role in building the first...H1bWork at officeRemote workRelocationVisa sponsorshipFlexible hours
$195k - $365k
...building and training large-scale audio or speech models from the ground up, whether that... ...living at the intersection of research and engineering, eager to design novel sequence modeling... ...May Also Have Experience With Text-based LLMs: Hands‑on experience with core text‑based...Full timeWork at officeWorldwide$200k - $235k
...managers, data scientists, software engineers, fraud intelligence, and... ...Together, you'll design and build ML solutions that have direct,... ...and safety use cases, and use LLMs and AI agents to accelerate... ...models (vision, document, or speech) is a plus. ~ Exposure to...Work experience placementCasual workLive inWork at officeRemote work- ...engagement, working directly with the engineer who leads our embedded... ...modality (computer vision, ADAS, LLMs, audio). Building applications... ...embedded-device, or on-device ML context, and the instinct to... ...Nice-to-have: Audio or speech model experience. Function...Full timeInternshipShift work
$150k - $225k
...re looking for an Founding AI Engineer to design, build, and scale... ...and experiment with multiple LLMs and APIs, choosing the right tool... ...building and scaling data or ML pipelines in production environments... ...modal models (text, image, or speech) Knowledge of GCP, Firebase,...Full timeWork from homeFlexible hours$170k - $280k
...Our mission is to give leaders clarity and engineers time. We help leaders understand how... ...the role We're looking for an Applied ML Engineer to help build and improve the machine... ...learning or reinforcement learning for LLMs (RLHF, RLAIF, GRPO, PPO, DPO, or similar...Odd jobFull time$220k - $280k
...voice agents and applications — serving speech-to-text and text-to-speech models with best... .... We're looking for a Staff ML Engineer to drive the model serving layer for voice... ...emerging model paradigms — audio-native LLMs, codec-based architectures (SNAC, Encodec...Full time- Design the intelligence and memory layer. LLMs, embeddings, and systems thinking required. Kodezi isn’t just another dev tool. It’s... ..., evolves, and documents codebases autonomously. As a Founding ML Engineer , you’ll architect the intelligence powering autonomous pull requests...Remote workFlexible hours
$150k - $300k
...records from across the web. You will own the ML systems that turn that raw, multilingual,... ..., classifiers, and embedding models. Use LLMs for structured extraction, classification... ...dilution. Own the full ML research and engineering cycle, from prototype to production....- ...voice interface to reach billions of users. We are hiring an ML Engineer to prototype and ship features for our voice interface and to... ...across a global user base. You will help scale personalization of speech models using fine-tuning and RL, collaborating with a small,...
- At Dynamo AI, we believe that LLMs must be developed with safety, privacy, and real-world responsibility in mind. Our ML team comes from a culture of academic research driven to democratize AI advancements responsibly. By operating at the intersection of ML research and...Local areaShift work
- About The Role Skills: Python, PyTorch, NLP, LLMs, Information Retrieval, Entity Resolution, Text Classification... ...someone who can push the boundaries of what our ML systems can do. We're hiring a Founding ML Engineer to own the research and engineering behind our core...
- ...Design, build, and deploy production‑grade ML systems with end‑to‑end ownership of the... ...AI‑powered solutions enabling natural speech interaction and real‑time audio understanding... ...6 years of professional experience in ML engineering. Strong programming skills in Python (...Full time
$197.3k - $313.7k
...Slack is looking for a Staff Machine Learning Engineer with deep expertise in model training and finetuning to join our ML team. You'll design, train, and ship NLP models... ...in NLP (or a closely related domain like speech, IR, or multimodal).5+ years of experience with...Full time- Abacus is hiring its first Founding Engineer to own and expand the core machine learning engine behind Abacus. You’ll report... ...direction. In your first 3-6 months, you’ll scale the ML extraction engine using LLMs and OCR, own backend architecture, and ship greenfield features...Work at office
- Position: Senior ML Performance Engineer Location: SF Bay Area (US) or Toronto (Canada) - Hybrid Employment Type: Full-Time Industry: AI Infrastructure... ...and optimizing the performance of large language models (LLMs) before and after compiler optimization on modern GPU...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Engineer: Speech & LLMs. Be the first to apply!
- machine learning engineer San Francisco, CA
- machine learning software engineer San Francisco, CA
- ai ml engineer San Francisco, CA
- entry level machine learning engineer San Francisco, CA
- junior machine learning research engineer San Francisco, CA
- machine learning ai engineer San Francisco, CA
- computer vision machine learning engineer San Francisco, CA
- senior ml engineer San Francisco, CA
- machine learning intern San Francisco, CA
- intern - quantum machine learning for quantum computing San Francisco, CA


