Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Inference Engineer - Speech

$151.8k - $332.2k
Full-time

Zoom Video Communications, Inc.

What you can expect

We are looking for an AI Inference Engineer with a solid background in speech recognition and model inference. In this role, you will develop state-of-the-art automatic speech recognition system and ship it to various Zoom products. You will work on the most cutting edge speech modeling and inference technologies with world-class speech scientists. This role will include collaboration with cross-functional teams, including product, science engineering teams, and infrastructure teams, to deliver high-impact projects from the ground up.

About the Team

Zoom's AI Speech Team is developing speech recognition technologies to improve Zoom's conversational AI experience. This work impacts various products, like Zoom AI Companion, Zoom Meetings and Workplace, Zoom Contact Center, Zoom Phone, Zoom Revenue Accelerator, etc. Our team's mission is to equip the powerful AI brain with human-level listening and understanding undefined for voice input.

As an AI Inference Engineer, you will develop novel speech model inference solutions on modern AI inference hardware, such as GPU, TPU and AI-specific chips. Our goal is to deliver the most unique AI-powered collaboration platform to users across the globe.

Responsibilities

  • Developing state-of-the-art speech services for Zoom products. Devising novel techniques where off-the-shelf solutions are not available.
  • Optimizing ASR inference systems for production deployment, including inference latency, throughput, memory footprint, and resource utilization.
  • Optimizing model inference performance by diving deep into the lower stack of inference frameworks, with a focus on hardware-specific optimizations for Nvidia GPUs.
  • Proposing new model structures by joint optimization of model accuracy and inference speed.
  • Designing and developing ASR systems with low latency and high accuracy requirements, while ensuring scalability of GPU infrastructure and improving throughput of ASR service.
  • Profiling and debugging ASR runtime performance bottlenecks across different deployment hardware and environments.

What we’re looking for

  • Possess a Master's in Computer Science, Electrical Engineering or related fields with 3+ years of experience in speech recognition, speech-llm or AI model inference.
  • Display knowledge in deep learning and hands-on programming skills in Python, shell scripts, C/C++; familiarity with ML frameworks such as PyTorch and TensorFlow.
  • Demonstrate deep understanding of transformer encoder-decoder frameworks for speech recognition, including attention mechanisms, beam search and sequence-to-sequence modeling for end-to-end ASR systems.
  • Understand recent advancements in speech foundation models and speech-LLMs that integrate acoustic and linguistic representations, enabling unified modeling for speech understanding and transcription tasks.
  • Have experience in optimizing deep learning model inference on NVIDIA GPUs, including profiling and accelerating AI models using CUDA, TensorRT, and mixed-precision computation to achieve low latency, high-throughput performance.
  • Have experience developing and tuning custom CUDA kernels, leveraging CUDA Graphs for efficient execution scheduling, and minimizing kernel launch overhead to maximize GPU utilization.
  • Be proficient in end-to-end performance analysis, memory optimization, and deployment of largescale ML models on GPU clusters. Experienced with stream management, asynchronous execution, and integrating frameworks such as PyTorch and TensorFlow for real-time inference.

Salary Range or On Target Earnings:

Minimum:

$151,800.00

Maximum:

$332,200.00

In addition to the base salary and/or OTE listed Zoom has a Total Direct Compensation philosophy that takes into consideration; base salary, bonus and equity value.

Note: Starting pay will be based on a number of factors and commensurate with qualifications & experience.

We also have a location based compensation structure; there may be a different range for candidates in this and other locations

At Zoom, we offer a window of at least 5 days for you to apply because we believe in giving you every opportunity. Below is the potential closing date, just in case you want to mark it on your calendar. We look forward to receiving your application!

Anticipated Position Close Date:

09/01/26

Ways of Working
Our structured hybrid approach is centered around our offices and remote work environments. The work style of each role, Hybrid, Remote, or In-Person is indicated in the job description/posting.

Benefits
As part of our award-winning workplace culture and commitment to delivering happiness, our benefits program offers a variety of perks, benefits, and options to help employees maintain their physical, mental, emotional, and financial health; support work-life balance; and contribute to their community in meaningful ways. Click Learnfor more information.

About Us
Zoomies help people stay connected so they can get more done together. We set out to build the best collaboration platform for the enterprise, and today help people communicate better with products like Zoom Contact Center, Zoom Phone, Zoom Events, Zoom Apps, Zoom Rooms, and Zoom Webinars.
We’re problem-solvers, working at a fast pace to design solutions with our customers and users in mind. Find room to grow with opportunities to stretch your skills and advance your career in a collaborative, growth-focused environment.


Our Commitment

At Zoom, we believe great work happens when people feel supported and empowered. We’re committed to fair hiring practices that ensure every candidate is evaluated based on skills, experience, and potential. If you require an accommodation during the hiring process, let us know—we’re here to support you at every step.


If you need assistance navigating the interview process due to a medical disability, please submit an Accommodations Request Form and someone from our team will reach out soon. This form is solely for applicants who require an accommodation due to a qualifying medical disability. Non-accommodation-related requests, such as application follow-ups or technical issues, will not be addressed.

Our interviews are supported by BrightHire, a tool that helps us create a consistent and thoughtful interview experience and may include recordings. Please refer to our candidate privacy statement for more information of how we use your data.

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the AI Inference Engineer - Speech in Seattle, WA vacancy
  • $176.6k - $265k

     ...customers, better. And it means we prioritize a diverse F5 community where each individual can thrive. Job Description The AI Inference Engineer plays a critical role in the AI lifecycle by bridging the gap between high-performance model development and optimized... 
    Suggested
    Full time
    Local area
    Immediate start

    F5

    Seattle, WA
    2 days ago
  • $188k - $275k

     ...CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave...  ...Do Description of the team: The Inference team is responsible for delivering high-performance...  ...: We are looking for an Applied AI Engineer to help us understand, measure, and... 
    Suggested
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Bellevue, WA
    2 days ago
  • NVIDIA in Seattle, WA seeks outstanding AI systems engineers to advance the inference software stack for AI workloads. You will develop libraries, code generators, and GPU kernel technologies for NVIDIA hardware, including new abstractions and runtimes for large language... 
    Suggested

    NVIDIA

    Seattle, WA
    3 days ago
  • $152k - $241.5k

     ...deep learning ignited modern AI — the next era of computing —...  ...is hiring a Senior AI Compiler Engineer. GPUs are driving rapid progress...  ...recommendation, vision, and speech. On this team, you’ll build an...  ...compiler that powers NVIDIA’s inference engine end to end, with a focus... 
    Suggested
    Full time
    Remote work

    Nvidia

    Seattle, WA
    2 days ago
  • $134.96k - $188.95k

     ...manual monitoring and reactive processes. We are seeking an AI Engineer II to join the Blue Ring software team. You will contribute to...  ...model compression, edge deployment, or resource-constrained inference Base Pay Range for: VA applicants is $134,961.00 - $188,... 
    Suggested
    Permanent employment
    Full time
    Temporary work
    Local area

    BLUE ORIGIN

    Seattle, WA
    3 days ago
  • $177.1k - $387.5k

     ...the platform that powers Zoom AI Services, enabling AI capabilities...  ..., and AI platform engineering to build reliable, high-performance services that power speech, translation, summarization, reasoning...  ..., including model serving, AI inference platforms, GPU/CPU resource management... 
    Full time
    Work at office
    Remote work

    Zoom

    Seattle, WA
    5 days ago
  •  ...As a Model Optimization & Deployment Engineer, you will focus on bringing highly efficient...  ...CUDA kernels, and build highly concurrent inference code to ensure real-time, deterministic execution...  ...latency and maximize memory bandwidth on AI accelerators. Write production-level,... 
    Temporary work
    Relocation package

    Zoox

    Seattle, WA
    21 days ago
  •  ...Senior Staff AI Engineer The AI Scientist will work in teams addressing statistical, machine learning and data understanding problems...  ...clinical reports, and other healthcare datasets. Develop robust inference pipelines, model serving infrastructure, and AI services... 
    Worldwide
    Flexible hours

    GE

    Bellevue, WA
    3 days ago
  • $154.56k - $193.2k

     ...hyperscaler for the edge, delivering modular AI infrastructure from first deployment to...  ...Armada is seeking exceptional AI Engineers to build and deploy intelligent systems at...  ...analysis, anomaly detection, or distributed AI inference. This role is intended for engineers... 
    Work at office
    Remote work
    Flexible hours

    Armada

    Bellevue, WA
    4 days ago
  • $168.1k - $227.4k

     ...machinelearning accelerators. This role is for a senior software engineer in the Machine Learning Inference Applications team. This role is responsible for...  ...deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware.Amazon is an equal opportunity... 
    Work experience placement
    Flexible hours

    Amazon

    Seattle, WA
    6 days ago
  • $92k - $135k

     ...CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave...  ...more at What You’ll Do: Join the Inference team to ship production features that improve...  ...quickly with mentorship from experienced engineers. About the role: Implement... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Internship
    Work at office
    Flexible hours

    CoreWeave

    Bellevue, WA
    2 days ago
  • $139k - $204k

     ...CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave...  ...Learn more at What You’ll Do: Senior engineers are area owners who lead designs, raise...  ...hardware teams to evolve our Kubernetes-native inference platform and meet strict P99 SLAs at... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours
    Shift work

    CoreWeave

    Bellevue, WA
    2 days ago
  • $193.3k - $261.5k

     ...cloud-scale machine learning. This senior software engineering role is part of the Machine Learning Inference Applications team and focuses on delivering high-performance...  ...in production on GPUs, AWS Neuron, TPUs, or other AI accelerator hardware- Experience extending or... 
    Work experience placement
    Local area
    Flexible hours

    Amazon

    Seattle, WA
    5 days ago
  • $152k - $241.5k

    We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join our AI Efficiency efforts. As an Engineer, you will...  ...tools Experience investigating, and resolving, training & inference performance end to endDebugging and optimization... 
    Full time
    Remote work

    Nvidia

    Seattle, WA
    5 days ago
  • $140k - $180k

     ...WashingtonValorem Reply Azure SI, US - Data + AI /Full Time /HybridValorem Reply is an...  ...clients do business. As a Senior AI Engineer, you will understand how AI is positioned...  ...generation (RAG), and Azure AI Services (OpenAI, Speech, Vision, etc.) using tools like Azure AI... 
    Full time

    Valorem Reply

    Seattle, WA
    5 days ago
  • $151.8k - $332.2k

    What you can expectWe are seeking an experienced AI Infrastructure Engineer to join our AI Incubation team. You will be focused on building and...  ...workflowsCollaborating with AI researchers to implement efficient training and inference pipelinesWhat we’re looking forHave a bachelor's degree in... 
    Full time
    Work at office
    Remote work

    Zoom

    Seattle, WA
    1 day ago
  • $151.8k - $332.2k

     ...What You Can ExpectYou'll design, implement, and own the inference systems that serve Zoom's AI models at production scale -- across real-time...  ...implementationsDrive technical design and set the bar for inference engineering practices across the teamWhat We're Looking ForA... 
    Full time
    Work at office
    Remote work

    Zoom

    Seattle, WA
    5 days ago
  • $177.1k - $387.5k

    What you can expectAs an AI Engineer specializing in Agentic AI, you will develop intelligent agents that can autonomously perceive, reason...  ...various AI components including LLM, RAG, TTS (Text-to-Speech), and ASR (Automatic Speech Recognition) systems.Writing and optimizing... 
    Full time
    Work at office
    Remote work

    Zoom

    Seattle, WA
    1 day ago
  • We Are:The Global AI Infrastructure team is at the center of enabling infrastructure reinvention for the next era of digital solutions...  ...Manager (BCM), NGC, NCCL, NVLink, and CUDA along with LLM inference engines (TensorRT-LLM), production serving frameworks (vLLM, SGLang),... 
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    Seattle, WA
    2 days ago
  • $163.2k - $220.8k

     ...and career growth. Wilson Sonsini is looking for a Senior AI Security Engineer to join the Security Operations team. The Senior AI Security...  ...secrets management for model API keys, network isolation for AI inference endpoints, and identity-aware proxy patterns for LLM access... 
    Full time
    Work experience placement
    Remote work
    Worldwide
    Shift work

    Wilson Sonsini Goodrich & Rosati

    Seattle, WA
    3 days ago
  • Job Title: Applied AI EngineerLocation: 100% RemoteFull-Time Role Role As an Applied AI Engineer, you will turn model capabilities into real product behavior.You will own...  ...(OpenAI-style APIs, LLaMA, Qwen, etc.)Inference / serving (e.g. vLLM)Vector DB Ideal Experience... 

    Spectraforce Technologies

    Seattle, WA
    1 day ago
  •  ...Senior AI Engineer – Privacy Location: Bellevue WA Must have skills – skill 1 – 7yrs of exp – AI Engineer – Privacy skill 2 – 7yrs...  ...Snowflake, or PySpark to support AI model training, fine-tuning, and inference. Apply prompt engineering, few-shot learning, and fine-... 

    Software Technology Inc

    Bellevue, WA
    15 hours ago
  • $143.7k - $194.4k

     ...commerce.Advertiser Growth Tech (AGT) is an engineering team with the mission to enhance the...  ...scalable code. You will work on Agentic AI initiatives, contributing to systems that...  ..., model training, optimization, and inference infrastructure. Lead technical design discussions... 
    Internship
    Flexible hours

    Amazon

    Seattle, WA
    2 days ago
  • $176.76k - $232k

     ...people.About this team The Enterprise Data & AI team is a strategic and operational...  ....Core responsibilities As a Senior AI/ML Engineer, you will lead the delivery of scalable AI...  ...architectures and system design for serving AI/ML inference solutions in production. You will help... 
    Permanent employment
    Full time
    Contract work
    Part time
    Work visa

    Lululemon Athletica

    Seattle, WA
    3 days ago
  •  ...the forefront of a new era in enterprise AI — one defined not by model capability...  ...of frontier AI research and production engineering — investigating the foundational challenges...  ...persistence architectures, model selection and inference routing strategies, autonomy and goal-... 
    Full time
    Work experience placement
    Live in
    Work at office
    Local area
    Relocation

    Accenture

    Seattle, WA
    2 days ago
  • $220k - $292k

     ...of systems is powered by Lattice OS, an AI-powered operating system that turns thousands...  ..., Motion Planning, Hardware, and Test Engineering to solve some of the hardest problems...  ...TPUs, or custom ASICs) for training and inference workloads. Model Observability at Scale:... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Seattle, WA
    3 days ago
  • $106.9k - $200.6k

     ...better working world. The opportunity We are seeking AI Systems Engineers to own the security and trust fabric of EY’s AI-native...  ...workloads and the specific trust challenges of confidential AI inference (models/secrets inside enclaves). Experience producing... 
    Full time
    Work experience placement
    Summer holiday
    Remote work
    Flexible hours

    EY

    Seattle, WA
    2 days ago
  • $106.9k - $200.6k

     ...working world. The opportunity We are seeking an AI Systems Engineer to own the delivery, model-serving, routing, and observability...  ...automated delivery pipelines, operating high-performance inference (GPUs, model servers, sandboxed execution), and building deep... 
    Full time
    Summer holiday
    Flexible hours

    EY

    Seattle, WA
    2 days ago
  • $220k - $293.33k

     ...software, and applications that power today’s AI stack using sustainable technology...  ...Nscale is looking for a Staff AI Product Engineer to drive technical direction across the...  ...Experience building AI/ML product platforms: inference APIs, fine-tuning UX, model management,... 
    Full time
    Contract work
    Flexible hours

    Nscale

    Seattle, WA
    2 days ago
  •  ...Lead AI Engineer in the Platforms and Products ZS is a place where passion changes lives. As a management consulting and technology...  ...retrieval-augmented generation, fine-tuning workflows, and scalable inference pipelines. Design and implement LLM-powered applications... 
    Local area
    Work from home
    Worldwide
    Flexible hours

    ZS Associates

    Bellevue, WA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Inference Engineer - Speech. Be the first to apply!