Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco

$180k - $270k
Full-time

Plaud

About Plaud Inc.

Plaud is building the world's most trusted AI work companion for professionals to elevate productivity and performance through note-taking solutions, loved by over 1,500,000 users worldwide since 2023. With a mission to amplify human intelligence, Plaud is building the next-generation intelligence infrastructure and interfaces to capture, extract, and utilize what you say, hear, see, and think.

 

Plaud Inc. is a Delaware-incorporated, San Francisco-based company pushing the boundary of human–AI intelligence through a hardware–software combination. With SOC 2, HIPAA, GDPR, ISO27001, ISO27701, and EN18031 compliance, Plaud is committed to the highest standards of data security and privacy protection.

To learn more about Plaud, please visit and follow along on Instagram , X , Facebook , LinkedIn , and YouTube

 

Why You Should Join Us

Plaud is building the next generation intelligence infrastructure and interfaces to capture, extract, and utilize intelligence from what people say, hear, see, and think.

  • Plaud is a bootstrapped, skyrocketing, profitable company with a $250M revenue run rate achieved in just three years.

  • Define the next-gen paradigm for human-AI interaction.

  • Gain exposure to cutting-edge AI for Pro tools and play a direct role in our global expansion.

  • Work with passionate teammates who value innovation, collaboration, and customer success.

  • Grow your career in a culture that champions continuous learning and fast career development.

  • Market-competitive compensation, global exposure, and a vibrant, creativity-fueled work atmosphere.

 

You may be a good fit if you:

  • Have hands-on experience building and deploying high-throughput, ultra-low-latency inference engines for large language models or foundational speech models.

  • Understand the intricate tradeoffs between latency, throughput, and Time-To-First-Token (or Time-To-First-Audio) in real-time streaming environments.

  • Have practical experience with continuous batching, KV cache management (e.g., PagedAttention), and stateful connections necessary for real-time conversational AI.

  • Possess a deep understanding of GPU architectures (NVIDIA Ampere/Hopper) and the memory hierarchy, allowing you to identify and eliminate hardware bottlenecks.

  • Communicate clearly and collaborate effectively, as you will sit at the critical intersection between the core ML training team and the backend infrastructure team.

  • Thrive in fast-moving environments and genuinely enjoy the systems-engineering challenge of squeezing every last drop of performance out of a cluster of GPUs.

  • Are obsessed with building AI systems that natively understand and generate speech, ultimately creating a hardware-software AI companion that amplifies human productivity.

 

Strong candidates may also have experience with:

  • Frontier Serving Frameworks: Deep, under-the-hood familiarity with modern LLM serving frameworks like vLLM, TensorRT-LLM, SGLang, or NVIDIA Triton Inference Server (bonus points for active open-source contributions to these repositories).

  • Real-Time Audio Streaming: Experience handling continuous audio streams over WebSockets or WebRTC, deploying neural audio codecs, and managing chunked audio generation to minimize conversational latency.

  • Advanced Inference Techniques: Implementing cutting-edge generation algorithms such as speculative decoding, lookahead decoding, or chunked prefill.

  • Model Compression & Quantization: Hands-on experience with post-training quantization (PTQ), deploying models in FP8, INT8, AWQ, or GPTQ, without degrading audio naturalness or ASR accuracy.

  • Large-Scale Distributed Systems: Deploying multi-GPU (Tensor Parallelism) and multi-node inference pipelines, and managing autoscaling infrastructure using Kubernetes.

 

What We Offer

  • Founding Team Initiative: Opportunity to be an early, foundational member of our core SpeechLLM lab, with meaningful ownership and impact on a fast-growing startup.

  • Competitive Compensation: $180K - $270K base salary + performance bonus + Equity.

  • Comprehensive Benefits: Top-tier healthcare for employees and dependents, including dental and vision, and a generous employer subsidy.

  • Retirement Planning: 401(k) plan for full-time employees with company matching.

  • Paid Time Off: Unlimited PTO, plus 13 paid holidays.

  • New Parent Leave: 12 weeks of paid time off to spend time with your new family, regardless of gender.

  • Hybrid Office: Minimum of 3x in-office per week to foster highly collaborative, fast-paced research.

  • Gear & Perks: Choice of top-of-the-line laptops/workstations, annual offsites, and a fully stocked office.

 

Plaud is and will continue to be an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristics.

Vacancy posted 17 hours ago
Similar jobs that could be interesting for youBased on the Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco in San Francisco, CA vacancy
  •  ...Machine Learning Engineer (Operations) Location: South San Francisco CA (Hybrid, 3 days/week) (Not remote) Duration: Long term...  ...storing metadata, features, or serving model predictions where applicable...  ...cases. Ability to monitor LLM performance, fine-tune... 
    Suggested
    Full time
    3 days per week

    Esrhealthcare

    San Bruno, CA
    17 hours ago
  • $155k - $180k

     ..., use Roboflow’s machine learning open source and hosted...  ...only product and engineering), so Roboflow...  ...all of this is inference — one of our...  ...so they can self-serve and go deeper, and...  ..., vLLM (or other LLM/model deployment...  ...New York City and San Francisco (and plan to open... 
    Suggested
    Full time
    Second job
    Remote work
    Work from home
    Relocation package
    Flexible hours
    Night shift

    Roboflow

    San Francisco, CA
    17 hours ago
  • $209k - $313k

     ...live in the moment, learn about the world,...  ...services.Snap Engineering teams build fun and...  ...re looking for a Machine Learning Engineer...  ...of causal inference and modern approaches...  ...our values, and serve our community, customers...  ...of the San Francisco Fair Chance Ordinance... 
    Suggested
    Full time
    Live in
    Work at office
    Local area

    Snap

    San Francisco, CA
    2 days ago
  • $222.72k - $389.75k

     ...We are looking for a Staff Machine Learning Engineer to lead the technical vision...  ...understand intention and infer interests from online activity...  ....Familiarity with LLM-powered productivity tools...  ...of the following offices: San Francisco, Palo Alto, Seattle.#LI-HYBRID... 
    Suggested
    Part time
    Work at office
    Local area
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    2 days ago
  •  ...the way people learn, starting with...  ...app, and we now serve learners across...  ...distributed team across San Francisco, Seoul, Tokyo,...  ...hiring an ML Engineer, Assessments...  ...Design) , Machine Learning, Product...  ...training → inference → feedback generation...  ...with speech/audio ML Experience... 
    Suggested
    Full time
    Live in
    Immediate start

    Speak

    San Francisco, CA
    17 hours ago
  •  ..., and we’re proud to serve 90+ leading health systems...  ...We're looking for a Machine Learning Engineer to design, build, and...  ..., monitoring, and inference Build intelligent...  ...services using modern NLP, LLM, classification,...  ...office presence in San Francisco and New York. R&D roles... 
    Full time
    Work at office
    Remote work
    Flexible hours
    2 days per week

    Plenful

    San Francisco, CA
    17 hours ago
  • $220k - $320k

     ...meet you. About Inference.net Inference....  ...-person team of engineers who work in-person in downtown San Francisco on difficult, high...  ...we can train and serve them, and how smoothly...  ...~ Experience with LLM-specific training...  ...the ability to learn quickly matter more... 
    Full time
    Work at office

    Inference

    San Francisco, CA
    17 hours ago
  •  ...Primer has offices in San Francisco, Pasadena, CA and...  ...Arlington, VA. Learn more at primer....  .... As a Staff Machine Learning Engineer, you’ll own AI-...  ...bets across our LLM, agentic, and NLP...  ...high-concurrency inference (Triton, vLLM, GPU-backed serving) that stays fast... 
    Full time
    Contract work
    Remote work
    Flexible hours

    Primer.ai

    San Francisco, CA
    17 hours ago
  • $197.3k - $225.1k

    Lead Machine Learning Engineer At Capital One, we are creating...  ...to reimagine how we serve our customers and...  ...large language model inference, similarity search, model...  ...state-of-the-art LLM optimization techniques...  ...Learning Engineer San Francisco, CA: $215,200 - $245,... 
    Full time
    Part time
    Internship
    H1b
    Local area

    Capital One Financial Corporation

    San Francisco, CA
    17 hours ago
  •  ...Staff Machine Learning Engineer About Sprinter Health At Sprinter...  ..., retrain, and serve machine learning models...  ...our training and inference pipelines, serving patterns...  ...with offices in both San Francisco and Menlo Park. We...  ...feature infrastructure, LLM infrastructure, or... 
    Full time
    Temporary work
    Work at office
    Relocation package
    Monday to Friday
    Monday to Thursday
    Flexible hours

    Sprinter Health

    San Francisco, CA
    17 hours ago
  • $200k - $260k

     ...is building the best inference infrastructure for voice...  ...and applications — serving speech-to-text and text-to-speech...  ...for a Senior ML Engineer to drive the model serving...  ...inference engines like TRT-LLM and SGLang to optimize...  ...required — you can learn this quickly if you... 
    Full time

    Together Ai

    San Francisco, CA
    17 hours ago
  • $140k - $200k

     ...’re looking for an ML Engineer to build the production...  ...monitor, retrain, and serve our machine-learning models reliably. You...  ...will build training and inference pipelines, serve...  ...0K - $200KLocationSan Francisco, CAAddress394 Pacific Avenue , San Francisco, California,... 
    Temporary work
    Work at office
    Monday to Friday
    Monday to Thursday

    Sprinter Health

    San Francisco, CA
    1 day ago
  • $220k - $280k

     ...Together AI is building the best inference infrastructure for voice...  ...voice agents and applications — serving speech-to-text and text-to-speech...  ...We're looking for a Staff ML Engineer to drive the model serving layer...  ...inference engines like TRT-LLM and SGLang to optimize how we... 
    Full time

    Together Ai

    San Francisco, CA
    17 hours ago
  • $215k - $322k

     ...GoFundMe as our next Staff Machine Learning Engineer (Pricing) . In this role,...  ...(data → training → online inference → measurement) with...  ...role will be located in the San Francisco, Bay Area. There will be an...  ...deploying real-time model serving (sub-100ms to low-hundreds... 
    Full time
    Temporary work
    Work at office
    Flexible hours

    Gofundme

    San Francisco, CA
    17 hours ago
  • $295k - $405.5k

     ...tech, data, and machine learning to connect this thriving...  ...Platform Engineer, you will own the...  ...including training, inference, feature...  ...authorityExperience integrating LLM workflows into...  ...problems that serve customers around...  ...headquarters in San Francisco and Kitchener-Waterloo... 
    Work experience placement
    Work at office
    Local area
    Remote work
    Monday to Friday
    Flexible hours
    3 days per week

    Faire

    San Francisco, CA
    1 day ago
  • $229k - $343k

     ...live in the moment, learn about the world,...  ...digital services.Snap Engineering teams build fun...  ...for a Staff Machine Learning Engineer...  ...training infrastructure, serving systems, and model...  ..., generative AI, LLM-based ranking, and...  ...requirements of the San Francisco Fair Chance... 
    Full time
    Live in
    Work at office
    Local area

    Snap

    San Francisco, CA
    3 days ago
  • $229k - $343k

     ...live in the moment, learn about the world,...  ...digital services.Snap Engineering teams build fun...  ...'re looking for a Machine Learning...  ...Experience with building LLM based information...  ...reinforce our values, and serve our community,...  ...of the San Francisco Fair Chance Ordinance... 
    Full time
    Live in
    Work at office
    Local area

    Snap

    San Francisco, CA
    2 days ago
  • $197.3k - $313.7k

     ...looking for a Staff Machine Learning Engineer with deep expertise...  ...Slack ship models that serve millions of users...  ...related domain like speech, IR, or multimodal)....  ...model optimization for inference (quantization,...  ...following link: to the San Francisco Fair Chance Ordinance... 
    Full time

    Salesforce

    San Francisco, CA
    3 days ago
  •  ...San Francisco, CA About the role A San Francisco startup is seeking a Senior Machine Learning Engineer to build and scale AI systems. You’ll work on impactful...  ...data pipelines, model serving, monitoring, and...  ...AWS, GCP) with scalable inference Why This Role is Exciting... 

    Twenty80 LLC

    San Francisco, CA
    1 day ago
  •  ...Machine Learning Engineer San Francisco, CA – Full-time, Mid to Senior, On-site Compensation Competitive salary Plus meaningful equity About This...  ..., JAX, or similar frameworks. Experience with model serving and inference optimization. Published work, competition results,... 
    Full time

    Zof AI

    San Francisco, CA
    1 day ago
  • $264.8k - $331k

    Machine Learning Systems Research Engineer, Agent Post-training - Enterprise GenAI...  ...and resources that serve all of our...  ...optimize our training and inference framework. Post-train...  ...least 1-3 years of LLM training in a...  ...in the locations of San Francisco, New York, Seattle... 
    Full time
    Contract work
    For contractors
    For subcontractor
    Work at office

    Scale LLP

    San Francisco, CA
    17 hours ago
  • $187.9k - $252k

    Job Posting Title:Lead Machine Learning EngineerReq ID:101546...  ...organization of engineers, product developers,...  ...consumer media touch points serving millions of people...  ...experience (RecSys, ML, AI/LLM) who can help bridge...  ...for this position in San Francisco, CA is $187,900.00 -... 
    Full time

    Hulu

    San Francisco, CA
    17 hours ago
  • $225k - $300k

     ...Role: As a Senior Machine Learning Engineer at Ambience , you...  ...working onsite at our San Francisco office three days per...  ...evaluation pipelines for LLM and agentic systems,...  ...LLMs, agents, NLP, speech, and multimodal AI...  ...evaluation, orchestration, serving, and observability,... 
    Full time
    Work at office
    Immediate start
    Remote work
    Flexible hours
    3 days per week

    Ambience Healthcare

    San Francisco, CA
    2 days ago
  • $268.08k

     ...our recruiting process here.We are looking for a Sr. Staff Machine Learning Engineer to be the Technical Lead for the Content Quality who will build...  ..., debugging, testing, and refactoring.Familiarity with LLM-powered productivity tools for documentation search, experiment... 
    Part time
    Work at office
    Local area
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    2 days ago
  • $151.8k - $265.35k

     ...verticals. We are hiring a Senior Machine Learning Engineer to build the pipelines and...  ..., all while ensuring served quality matches the training...  ...of production ML or inference services at scale. Strong Python...  ...liability.SummaryLocation: San Jose; Seattle; San FranciscoType... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Francisco, CA
    3 days ago
  • $200.8k - $251k

    A leading AI technology company in San Francisco seeks a team member to build and optimize a machine learning framework for large language models. Candidates should...  ...system optimization experience and solid software engineering skills, particularly in tools like CUDA and... 
    Full time

    Scale AI

    San Francisco, CA
    1 day ago
  •  ...to reinvent the way people learn, starting with language. We...  ...more than 90 based throughout San Francisco, Seoul, Tokyo, Taipei, and Ljubljana...  ...looking for an experienced Machine Learning Engineer to join our team and help develop cutting-edge speech recognition models that help... 
    Full time
    Live in
    Work at office
    Worldwide

    Speak

    San Francisco, CA
    17 hours ago
  • $293.6k - $335.1k

    Director, Machine Learning Engineer As a Capital One Machine Learning Engineer, you'll be providing...  ...within an Agile environment, you'll serve as a technical domain expert in...  ...to be regularly worked. San Francisco, CA: $293,600 - $335,100 for Director... 
    Full time
    Part time
    Local area

    Capital One Financial Corporation

    San Francisco, CA
    17 hours ago
  • $246.5k - $339k

     ...tech, data, and machine learning to connect this thriving...  ...Platform Engineer, you will help design...  ...training and inference workloadsConfigure...  ...2Salary RangeSan Francisco: the pay range for...  ...meaningful problems that serve customers around...  ...headquarters in San Francisco and... 
    Work experience placement
    Work at office
    Local area
    Remote work
    Monday to Friday
    Flexible hours
    3 days per week

    Faire

    San Francisco, CA
    4 days ago
  • $170k - $225k

     ...requiring 2 days in office at our San Francisco hub every Tuesday &...  ...). About the RoleMachine Learning is a cornerstone at Taskrabbit...  ...we’re looking for a Staff Machine Learning Engineer to take technical...  ...engineers around you. You’ll also serve as the primary technical... 
    H1b
    Work at office
    Immediate start
    Flexible hours

    Taskrabbit

    San Francisco, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco. Be the first to apply!