Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco

$180k - $270k
Full-time

Plaud

About Plaud Inc.

Plaud is building the world's most trusted AI work companion for professionals to elevate productivity and performance through note-taking solutions, loved by over 1,500,000 users worldwide since 2023. With a mission to amplify human intelligence, Plaud is building the next-generation intelligence infrastructure and interfaces to capture, extract, and utilize what you say, hear, see, and think.

 

Plaud Inc. is a Delaware-incorporated, San Francisco-based company pushing the boundary of human–AI intelligence through a hardware–software combination. With SOC 2, HIPAA, GDPR, ISO27001, ISO27701, and EN18031 compliance, Plaud is committed to the highest standards of data security and privacy protection.

To learn more about Plaud, please visit and follow along on Instagram , X , Facebook , LinkedIn , and YouTube

 

Why You Should Join Us

Plaud is building the next generation intelligence infrastructure and interfaces to capture, extract, and utilize intelligence from what people say, hear, see, and think.

  • Plaud is a bootstrapped, skyrocketing, profitable company with a $250M revenue run rate achieved in just three years.

  • Define the next-gen paradigm for human-AI interaction.

  • Gain exposure to cutting-edge AI for Pro tools and play a direct role in our global expansion.

  • Work with passionate teammates who value innovation, collaboration, and customer success.

  • Grow your career in a culture that champions continuous learning and fast career development.

  • Market-competitive compensation, global exposure, and a vibrant, creativity-fueled work atmosphere.

 

You may be a good fit if you:

  • Have hands-on experience building and deploying high-throughput, ultra-low-latency inference engines for large language models or foundational speech models.

  • Understand the intricate tradeoffs between latency, throughput, and Time-To-First-Token (or Time-To-First-Audio) in real-time streaming environments.

  • Have practical experience with continuous batching, KV cache management (e.g., PagedAttention), and stateful connections necessary for real-time conversational AI.

  • Possess a deep understanding of GPU architectures (NVIDIA Ampere/Hopper) and the memory hierarchy, allowing you to identify and eliminate hardware bottlenecks.

  • Communicate clearly and collaborate effectively, as you will sit at the critical intersection between the core ML training team and the backend infrastructure team.

  • Thrive in fast-moving environments and genuinely enjoy the systems-engineering challenge of squeezing every last drop of performance out of a cluster of GPUs.

  • Are obsessed with building AI systems that natively understand and generate speech, ultimately creating a hardware-software AI companion that amplifies human productivity.

 

Strong candidates may also have experience with:

  • Frontier Serving Frameworks: Deep, under-the-hood familiarity with modern LLM serving frameworks like vLLM, TensorRT-LLM, SGLang, or NVIDIA Triton Inference Server (bonus points for active open-source contributions to these repositories).

  • Real-Time Audio Streaming: Experience handling continuous audio streams over WebSockets or WebRTC, deploying neural audio codecs, and managing chunked audio generation to minimize conversational latency.

  • Advanced Inference Techniques: Implementing cutting-edge generation algorithms such as speculative decoding, lookahead decoding, or chunked prefill.

  • Model Compression & Quantization: Hands-on experience with post-training quantization (PTQ), deploying models in FP8, INT8, AWQ, or GPTQ, without degrading audio naturalness or ASR accuracy.

  • Large-Scale Distributed Systems: Deploying multi-GPU (Tensor Parallelism) and multi-node inference pipelines, and managing autoscaling infrastructure using Kubernetes.

 

What We Offer

  • Founding Team Initiative: Opportunity to be an early, foundational member of our core SpeechLLM lab, with meaningful ownership and impact on a fast-growing startup.

  • Competitive Compensation: $180K - $270K base salary + performance bonus + Equity.

  • Comprehensive Benefits: Top-tier healthcare for employees and dependents, including dental and vision, and a generous employer subsidy.

  • Retirement Planning: 401(k) plan for full-time employees with company matching.

  • Paid Time Off: Unlimited PTO, plus 13 paid holidays.

  • New Parent Leave: 12 weeks of paid time off to spend time with your new family, regardless of gender.

  • Hybrid Office: Minimum of 3x in-office per week to foster highly collaborative, fast-paced research.

  • Gear & Perks: Choice of top-of-the-line laptops/workstations, annual offsites, and a fully stocked office.

 

Plaud is and will continue to be an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristics.

Vacancy posted 7 hours ago
Similar jobs that could be interesting for youBased on the Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco in San Francisco, CA vacancy
  • $180k - $250k

    A tech company in San Francisco is looking for a Staff Software Engineer to enhance model performance for generative media models. The ideal candidate will design advanced model serving architectures, develop performance tools, and collaborate with the Applied ML team.... 
    Suggested
    Full time

    fal

    San Francisco, CA
    1 hour ago
  •  ...Senior Staff Machine Learning Engineer, Community Support Engineering Airbnb was...  ...three guests to their San Francisco home, and has since grown to...  ...services and tools including LLM fine-tuning and optimization...  ...model development, low-latency serving and ease of model quality... 
    Suggested
    Full time
    Work experience placement
    Casual work
    Live in
    Work at office
    Remote work

    airbnb, Inc.

    San Francisco, CA
    1 hour ago
  •  ...Machine Learning Engineer (Operations) Location: South San Francisco CA (Hybrid, 3 days/week) (Not remote) Duration: Long term...  ...storing metadata, features, or serving model predictions where applicable...  ...cases. Ability to monitor LLM performance, fine-tune... 
    Suggested
    Full time
    3 days per week

    Esrhealthcare

    San Bruno, CA
    7 hours ago
  • A cutting-edge technology company in San Francisco is looking for a Founding ML Research Engineer to develop the infrastructure for training large speech models. This entry-level position involves designing a production-grade training stack, building scalable data pipelines... 
    Suggested
    Full time

    Kalpa Labs (YC F25)

    San Francisco, CA
    1 hour ago
  •  ...Machine Learning Infrastructure Engineer Join to apply for the Machine Learning Infrastructure Engineer role...  ...and maintaining training and serving infrastructure for ML research....  ...000.00-$207,000.00 2 weeks ago San Francisco, CA $130,000.00-$230,000.00 5 months... 
    Suggested
    Full time
    Internship

    Character.AI

    San Francisco, CA
    1 hour ago
  •  ...in our Palo Alto or San Francisco offices and will require...  ...harness cutting‑edge machine learning to redefine how...  ...recommendation systems to serve millions, balancing...  ..., collaborating with engineering, data science and product...  ...and maintaining LLM workflows for nuanced... 
    Full time
    Casual work
    Work at office
    Immediate start
    Flexible hours

    Grindr LLC

    San Francisco, CA
    1 hour ago
  • $155k - $180k

     ..., use Roboflow’s machine learning open source and hosted...  ...only product and engineering), so Roboflow...  ...all of this is inference — one of our...  ...so they can self-serve and go deeper, and...  ..., vLLM (or other LLM/model deployment...  ...New York City and San Francisco (and plan to open... 
    Full time
    Second job
    Remote work
    Work from home
    Relocation package
    Flexible hours
    Night shift

    Roboflow

    San Francisco, CA
    7 hours ago
  • Senior ML Systems Engineer, Frameworks & Tooling at Cohere Our mission...  ...is to scale intelligence to serve humanity. We’re training and...  ...responsible for large-scale LLM training. Design...  ...offices in Toronto, New York, San Francisco, London and Paris, as well as... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Cohere

    San Francisco, CA
    1 hour ago
  •  ...Staff Machine Learning Engineer, Listings and Host Tools Data and AI Airbnb was born in 2007 when...  ...hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million...  ...models and will build services for serving that are used in the above areas.... 
    Full time
    Work experience placement

    airbnb, Inc.

    San Francisco, CA
    1 hour ago
  • $100k - $300k

     ...scale through data-driven machine learning is the key to unlocking these...  ...for a Machine Learning Engineer to be responsible for designing...  ...effectively with inference, application, and deployment...  ...$202,000.00 2 weeks ago San Francisco, CA $130,000.00-$230,000.00... 
    Full time

    Skild AI

    San Francisco, CA
    1 hour ago
  • $133.5k - $212k

     ...presence—including offices in San Francisco, New York, Denver, London,...  ...encourage you to apply. Learn more about our story and mission...  ...are looking for a Senior Machine Learning Engineer to build the core Machine...  ...evaluation frameworks for LLM- and agent-based features,... 
    Full time
    Contract work
    Local area
    Immediate start
    Remote work
    Worldwide
    Home office
    Flexible hours

    Iterable

    San Francisco, CA
    1 hour ago
  •  ...We are a small, fast-growing team of engineers in San Francisco powering Fortune 100 enterprises, YC...  ...our San Francisco office Eager to learn and adapt quickly Prior startup or...  ...active learning pipelines Optimize inference, batching, and quantization on GPU... 
    Full time
    Work at office
    Visa sponsorship
    Relocation package

    Trypulse

    San Francisco, CA
    1 hour ago
  •  ...Silicon Valley, a small team of engineers is working on what could...  ...platforms, MLOps tools, data/LLM infrastructure). You will...  ...expected to work from our downtown San Francisco office 2 to 3 days per week....  ...etc. Deep understanding of machine learning algorithms and the model... 
    Full time
    Work at office
    Flexible hours
    2 days per week
    3 days per week

    Sailplane

    San Francisco, CA
    1 hour ago
  • $140k - $240k

     ...Job Title: Senior Machine Learning Engineer Salary: $140,000 – $240,000 + Equity + Benefits Job Type...  ...join a fast-growing AI startup based in San Francisco , focused on transforming the way...  ...decision-making. Projects may include: GPT/LLM-powered sales call analysis and... 
    Full time

    Willing Care Recruitment

    San Francisco, CA
    1 hour ago
  • $200k - $400k

     ...corporation headquartered in San Francisco with a team of the world’s...  ...researchers and engineers from organizations like OpenAI...  ...the role We’re looking for Machine Learning Engineers to help build our...  ...interpretability, training, and inference. Integrate new machine learning... 
    Full time

    Goodfire

    San Francisco, CA
    1 hour ago
  •  ...Principal Data Scientist to serve as the statistical...  .... Partner with engineering VPs, product leaders,...  ...forecasting systems, causal inference approaches — to match...  ..., optimization, and machine learning. ~ Exceptional...  ...is headquartered in San Francisco, with offices around... 
    Part time
    Remote work
    Worldwide

    Cacheflow

    San Francisco, CA
    1 hour ago
  • $180.6k - $315k

     ...tools, and resources that serve all of our enterprise...  ...that makes the whole machine move. This includes...  ...to use to improve an LLM/Agent ~ Publications...  ...retirement benefits, a learning and development stipend...  ...position in the locations of San Francisco, New York, Seattle is:... 
    Full time

    Scale AI, Inc.

    San Francisco, CA
    1 hour ago
  • $300k

     ...experimentation, full-scale model training, or inference. Our client operates high-...  ..., tune, and operate inference engines such as vLLM, SGLang, and TensorRT-LLM across multiple model types....  ...up to date with the latest models, serving frameworks, and optimisation techniques... 
    Full time
    Worldwide

    Hamilton Barnes Associates Limited

    San Francisco, CA
    1 hour ago
  •  ...the way people learn, starting with...  ...app, and we now serve learners across...  ...distributed team across San Francisco, Seoul, Tokyo,...  ...hiring an ML Engineer, Assessments...  ...Design) , Machine Learning, Product...  ...training → inference → feedback generation...  ...with speech/audio ML Experience... 
    Full time
    Live in
    Immediate start

    Speak

    San Francisco, CA
    7 hours ago
  • $220k - $260k

     ...We are looking for a Founding Machine Learning Engineer to lead the development of scalable ML systems for RNA biology at EPM Scientific....  ...Level Entry level Employment Type Full-time Job Function Research Location: San Francisco, CA #J-18808-Ljbffr
    Full time

    EPM Scientific

    San Francisco, CA
    1 hour ago
  • $177.31k - $310.29k

     ...Sr. Machine Learning Engineer, Monetization Engineering Join to apply for the Sr. Machine Learning Engineer, Monetization Engineering role...  ...AI/ML Recommendations, Rankings, Predictions, YouTube San Francisco, CA $130,000.00-$230,000.00 5 months ago Staff... 
    Full time
    Local area
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    1 hour ago
  • $200k - $300k

     ...experience — talk with your recruiter to learn more. Base pay range $200,000.00...  ...the job poster from Willing Tech Machine Learning Engineer – Scientific Visualisation Platform...  ...Machine Learning Engineer jobs in San Francisco, CA . San Francisco, CA $135,000.... 
    Full time
    Remote work

    Willing Tech

    San Francisco, CA
    1 hour ago
  •  ...Python and Kubernetes Software Engineer - Data, AI/ML & Analytics...  ...composed of popular, open-source, machine learning tools, such as Kubeflow,...  ..., or as web services. We serve the needs of individuals and...  ...Internship (12 months) San Francisco, CA $110,000.00-$180,000.00... 
    Full time
    Freelance
    Internship
    Local area
    Remote work
    Work from home
    Worldwide

    Canonical

    San Francisco, CA
    1 hour ago
  • ML/AI Research Engineer — Agentic AI Lab (Founding Team) Location: San Francisco Bay Area Type: Full-Time Compensation...  ..., and reinforcement learning — building the...  ...evaluation harnesses for LLM and agent performance,...  ...alignment Optimize inference latency and GPU... 
    Full time

    Fabrion

    San Francisco, CA
    1 hour ago
  • $99.6k - $234.6k

     ...Principal AI Agent / ML Software Engineer is a Senior Staff-level, hands...  ...workflows, scalable inference infrastructure, and enterprise...  ...organization. Responsibilities Serve as a senior technical owner for...  .... ~ Deep understanding of LLM application patterns, including... 
    Full time
    Temporary work
    Flexible hours

    Oracle

    San Francisco, CA
    1 hour ago
  • $220k - $320k

     ...meet you. About Inference.net Inference....  ...-person team of engineers who work in-person in downtown San Francisco on difficult, high...  ...we can train and serve them, and how smoothly...  ...~ Experience with LLM-specific training...  ...the ability to learn quickly matter more... 
    Full time
    Work at office

    Inference

    San Francisco, CA
    7 hours ago
  • $100k - $400k

     ...talk with your recruiter to learn more. Base pay range $100,00...  ...company to recruit a Staff Machine Learning Engineer for their groundbreaking brain...  ....00-$170,000.00 2 weeks ago San Jose, CA $137,500.00-$236,50...  ...- Safety Response San Francisco Bay Area $140,000.00-$157,50... 
    Full time
    Immediate start

    Metric Bio

    San Francisco, CA
    1 hour ago
  •  ...delightful way. Our $1B+ learning platform serves tens of millions of students...  ...cognitive science with machine learning to personalize and...  ...collaborating closely with engineering, product, design, and data...  ...Townsend Street, Suite 600, San Francisco, CA 94107. Salary: $194,83... 
    Full time
    Immediate start
    Remote work

    Quizlet

    San Francisco, CA
    1 hour ago
  • $212k - $276.5k

     ...deeply personalized and adaptive user experiences. As a Staff Machine Learning Engineer, you will lead complex AI initiatives, architect cutting-...  ...analytics to continuously improve AI performance. Serve as a technical mentor, conducting knowledge-sharing sessions... 
    Full time

    Airbnb

    San Francisco, CA
    1 hour ago
  • $200k - $350k

     ...company that’s redefining how models learn to understand subjective quality...  ...creativity . The Role As a Machine Learning Research Engineer , you’ll own end-to-end research...  ...founding team in Jackson Square, San Francisco Competitive salary and meaningful... 
    Full time

    Coders Connect

    San Francisco, CA
    1 hour ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco. Be the first to apply!