Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Inference Engineer - Speech

$151.8k - $332.2k

Zoom Video Communications, Inc.

What you can expect We are looking for an AI Inference Engineer with a solid background in speech recognition and model inference. In this role, you will develop state-of-the-art automatic speech recognition system and ship it to various Zoom products. You will work on the most cutting edge speech modeling and inference technologies with world-class speech scientists. This role will include collaboration with cross-functional teams, including product, science engineering teams, and infrastructure teams, to deliver high-impact projects from the ground up.About the Team Zoom's AI Speech Team is developing speech recognition technologies to improve Zoom's conversational AI experience. This work impacts various products, like Zoom AI Companion, Zoom Meetings and Workplace, Zoom Contact Center, Zoom Phone, Zoom Revenue Accelerator, etc. Our team's mission is to equip the powerful AI brain with human-level listening and understanding undefined for voice input.As an AI Inference Engineer, you will develop novel speech model inference solutions on modern AI inference hardware, such as GPU, TPU and AI-specific chips. Our goal is to deliver the most unique AI-powered collaboration platform to users across the globe.ResponsibilitiesDeveloping state-of-the-art speech services for Zoom products. Devising novel techniques where off-the-shelf solutions are not available.Optimizing ASR inference systems for production deployment, including inference latency, throughput, memory footprint, and resource utilization.Optimizing model inference performance by diving deep into the lower stack of inference frameworks, with a focus on hardware-specific optimizations for Nvidia GPUs.Proposing new model structures by joint optimization of model accuracy and inference speed.Designing and developing ASR systems with low latency and high accuracy requirements, while ensuring scalability of GPU infrastructure and improving throughput of ASR service.Profiling and debugging ASR runtime performance bottlenecks across different deployment hardware and environments.What we’re looking forPossess a Master's in Computer Science, Electrical Engineering or related fields with 3+ years of experience in speech recognition, speech-llm or AI model inference.Display knowledge in deep learning and hands-on programming skills in Python, shell scripts, C/C++; familiarity with ML frameworks such as PyTorch and TensorFlow.Demonstrate deep understanding of transformer encoder-decoder frameworks for speech recognition, including attention mechanisms, beam search and sequence-to-sequence modeling for end-to-end ASR systems.Understand recent advancements in speech foundation models and speech-LLMs that integrate acoustic and linguistic representations, enabling unified modeling for speech understanding and transcription tasks.Have experience in optimizing deep learning model inference on NVIDIA GPUs, including profiling and accelerating AI models using CUDA, TensorRT, and mixed-precision computation to achieve low latency, high-throughput performance.Have experience developing and tuning custom CUDA kernels, leveraging CUDA Graphs for efficient execution scheduling, and minimizing kernel launch overhead to maximize GPU utilization.Be proficient in end-to-end performance analysis, memory optimization, and deployment of largescale ML models on GPU clusters. Experienced with stream management, asynchronous execution, and integrating frameworks such as PyTorch and TensorFlow for real-time inference.Salary Range or On Target Earnings:Minimum:$151,800.00Maximum:$332,200.00In addition to the base salary and/or OTE listed Zoom has a Total Direct Compensation philosophy that takes into consideration; base salary, bonus and equity value.Note: Starting pay will be based on a number of factors and commensurate with qualifications & experience.We also have a location based compensation structure; there may be a different range for candidates in this and other locationsAt Zoom, we offer a window of at least 5 days for you to apply because we believe in giving you every opportunity. Below is the potential closing date, just in case you want to mark it on your calendar. We look forward to receiving your application!Anticipated Position Close Date:08/12/26Ways of WorkingOur structured hybrid approach is centered around our offices and remote work environments. The work style of each role, Hybrid, Remote, or In-Person is indicated in the job description/posting.BenefitsAs part of our award-winning workplace culture and commitment to delivering happiness, our benefits program offers a variety of perks, benefits, and options to help employees maintain their physical, mental, emotional, and financial health; support work-life balance; and contribute to their community in meaningful ways. Click Learnfor more information.About UsZoomies help people stay connected so they can get more done together. We set out to build the best collaboration platform for the enterprise, and today help people communicate better with products like Zoom Contact Center, Zoom Phone, Zoom Events, Zoom Apps, Zoom Rooms, and Zoom Webinars.We’re problem-solvers, working at a fast pace to design solutions with our customers and users in mind. Find room to grow with opportunities to stretch your skills and advance your career in a collaborative, growth-focused environment.Our Commitment​At Zoom, we believe great work happens when people feel supported and empowered. We’re committed to fair hiring practices that ensure every candidate is evaluated based on skills, experience, and potential. If you require an accommodation during the hiring process, let us know—we’re here to support you at every step.If you need assistance navigating the interview process due to a medical disability, please submit an Accommodations Request Form and someone from our team will reach out soon. This form is solely for applicants who require an accommodation due to a qualifying medical disability. Non-accommodation-related requests, such as application follow-ups or technical issues, will not be addressed.Our interviews are supported by BrightHire, a tool that helps us create a consistent and thoughtful interview experience and may include recordings. Please refer to our candidate privacy statement for more information of how we use your data.SummaryLocation: Seattle (WA); San Jose (CA)Type: Full time

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the AI Inference Engineer - Speech in San Jose, CA vacancy
  •  ...fast-moving technology company to find an AI Engineer to build and ship LLM-powered...  ...production services that rely on LLMs and speech models. What you'll do Develop...  ...thousands of users. Build and optimize inference pipelines and real-time speech recognition... 
    Suggested
    Full time
    Work experience placement
    Work at office
    Remote work

    twenty80.io

    San Jose, CA
    1 day ago
  • $152k - $241.5k

     ...into the unlimited potential of AI to define the next era of...  ...an AI & Deep Learning Compiler Engineer. NVIDIA is hiring software engineers...  ..., image classification, speech recognition, etc. With the rapid...  ...been the backbone of NVIDIA’s inference engine, spanning across data centers... 
    Suggested
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    6 days ago
  • $160k - $198k

     ...team members.What You’ll DoAs a Senior AI Systems Engineer, you will architect, deploy, and...  ...for large-scale AI model training and inference. You will ensure our machine learning...  ...QualificationsFamiliarity with audio processing, speech-to-text frameworks, or Automatic... 
    Suggested
    Local area

    Archer Aviation

    San Jose, CA
    3 days ago
  • $117.7k - $221.4k

     ...practical, and cost efficient for embodied AI systems. We believe the next generation...  .... This operating model reflects how Cola engineers think: build durable intermediate artifacts...  ...the data processing, featurization, and inference foundations that power scalable world... 
    Suggested
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    1 day ago
  • $190k - $237k

     ...team members.What You’ll DoAs a Staff AI Systems Engineer, you will architect, deploy, and manage...  ...for large-scale AI model training and inference. You will ensure our machine learning...  ...QualificationsFamiliarity with audio processing, speech-to-text frameworks, or Automatic... 
    Suggested
    Local area

    Archer Aviation

    San Jose, CA
    1 day ago
  • $152k - $241.5k

     ...deep learning ignited modern AI — the next era of computing —...  ...an AI & Deep Learning Compiler Engineer. NVIDIA is hiring software engineers...  ..., image classification, speech recognition, etc. With the rapid...  ...been the backbone of NVIDIA’s inference engine, spanning across data... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $177.1k - $387.5k

     ...the platform that powers Zoom AI Services, enabling AI capabilities...  ..., and AI platform engineering to build reliable, high-performance services that power speech, translation, summarization, reasoning...  ..., including model serving, AI inference platforms, GPU/CPU resource management... 
    Full time
    Work at office
    Remote work

    Zoom

    San Jose, CA
    2 days ago
  •  ...next-generation computing experiences—from AI and data centers, to PCs, gaming and...  ...advance your career. THE ROLEWe are hiring AI Engineers to build recursive self-improvement...  ...JAX, TensorFlow, or distributed training/inference systems.Experience with reinforcement learning... 

    AMD

    Santa Clara, CA
    1 day ago
  •  ...prioritize a diverse F5 community where each individual can thrive.AI Engineer — Customer Success & Services (F5) Location: Hybrid (San Jose...  ...and operate production-quality ML services and APIs (scalable inference, caching, batching, latency SLAs); write performant, well-... 
    Full time
    Local area

    F5 Networks

    San Jose, CA
    1 day ago
  • $100k

     ...is leading the industry on cutting-edge AI technology, revolutionizing performance expectations...  ...Speed Interconnect / Signal Integrity Engineer to design and validate high-bandwidth...  ...technologies for next-generation AI inference and training clusters. This role is on-site... 
    Permanent employment

    Tenstorrent

    Santa Clara, CA
    5 days ago
  • $229.9k - $262.4k

     ...Overview Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning... 
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    more than 2 months ago
  • $152k - $241.5k

    We are now looking for a Senior Software Engineer for Quantized Inference! NVIDIA is seeking software engineers to accelerate the discovery and deployment...  ...fundamentals: concise, well-tested code; fluent with AI-assisted toolingExperience with ML accelerators with a basic... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...forefront of innovation, driving advancements in AI and machine learning to solve some of the...  .... We're seeking talented and motivated engineers to join our TensorRT team in developing the industry-leading deep learning inference software for NVIDIA AI accelerators. As a Senior... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

    We are now looking for a Senior Software Engineer for Deep Learning Inference! Would you like to make a big impact in Deep Learning by helping build a...  ...026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

    NVIDIA is the platform upon which every new AI‑powered application is built. We are seeking a Senior Software Engineer - AI Inference to advance open‑source LLM serving by contributing directly to upstream inference engines like vLLM and SGLang-ensuring they run best‑in... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...next-generation computing experiences—from AI and data centers, to PCs, gaming and...  ...RL training and SOTA LLM and Multimodal inference at scale across multi-GPU and multi-node...  ...software ecosystem. THE PERSON:   Skilled engineer with strong technical and analytical expertise... 

    AMD

    Santa Clara, CA
    6 days ago
  • $100k

     ...AI Optimization Engineer - Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud...  ...minimizing latency, and reducing cost across training and inference workloads for large neural network systems. The role spans... 
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    Santa Clara, CA
    3 days ago
  • $133.2k - $192.8k

    Job Details: Job Description: AI Engineer - Agentic AI Systems Why This Role Matters AI is shifting from models to autonomous systems...  ...are optimized across CPU / GPU / FPGA Explore efficient inference and system-level performance tradeoffs Build Platforms, Not Just... 
    Internship
    Local area
    Shift work

    Altera

    San Jose, CA
    6 hours ago
  •  ...class founding team, we build multi-agent AI systems that can automate complex...  ...level, we're looking for a top-quality AI engineer with a strong focus on AI agents - someone...  ...platforms such as AWS, Azure, or GCP. Deploy inference endpoints and serve AI and LLM... 

    Tessera Labs

    San Jose, CA
    2 days ago
  • $2,000 per month

     ...About Etched Etched is building AI chips that are hard-coded for individual model architectures...  ...agents. Job Summary Etched’s Inference SW team enables optimal mapping of models...  ...seeking a highly skilled and motivated engineer to join our team as we work towards... 
    Full time
    Work at office
    Relocation package

    Etched

    San Jose, CA
    1 day ago
  • $195.2k - $275.58k

    Job Details:Job Description: The Software and AI (SAI) organization is seeking a highly skilled Software Development Engineer to contribute to the development and...  ...optimization to achieve best‑in‑class deep‑learning inference and training throughput on current and next‑generation... 
    Full time
    Local area
    Immediate start
    Remote work
    Worldwide
    Flexible hours
    Shift work

    Intel

    Santa Clara, CA
    2 days ago
  •  ...next-generation computing experiences—from AI and data centers, to PCs, gaming and...  ...looking for a Forward Deployed Research Engineer to build, evaluate, and deploy cutting-edge...  ...ModelingModel or Agent EvaluationAI Training or Inference InfrastructureDeep understanding of... 

    AMD

    Santa Clara, CA
    2 days ago
  • $255.65k - $299k

     ...available for this positionWhat you can expect:As a Senior AI Software Engineer, you will collaborate to design, implement, and optimize AI...  ...algorithms and software applications. You will ensure AI training, inference, deployment, and operation are functional, reliable, and... 
    Full time
    Work at office
    Remote work

    Zoom

    San Jose, CA
    4 days ago
  • $152k - $241.5k

    NVIDIA's GPUs are at the core of modern AI infrastructure, from training large-scale models to running inference in production. That position depends on software as much as hardware, and compiler engineering is a big part of what makes it work.We are looking for outstanding... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    5 days ago
  • $184k - $287.5k

    We're looking for outstanding AI systems engineers to develop groundbreaking technologies in the inference systems software stack! We build innovative AI systems software to accelerate for AI inference. As a member of the team, you'll develop libraries, code generators... 
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  • $264.51k - $332.2k

     ...experience in job offered or related occupation. Must have 4 years of experience in the following:4 years of experience in ASR (Automatic Speech Recognition) for live captioning;4 years of experience in Conformer model structure, RNN-T (Recursive Neural Network – Transducer)... 
    Full time
    Work at office
    Remote work

    Zoom

    San Jose, CA
    6 days ago
  • $135.91k - $223.29k

    Job Title:Senior AI Engineer, MarTechRole Overview:We are looking for a Senior AI Engineer to lead the transformation of our marketing ecosystem...  ...: Strong understanding of Bayesian A/B testing and causal inference to measure the true uplift of AI interventions.Strategic... 
    Full time
    Temporary work
    Relocation package
    Flexible hours

    McAfee

    San Jose, CA
    3 days ago
  • $152k - $241.5k

    We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join our AI Efficiency efforts. As an Engineer, you will...  ...tools Experience investigating, and resolving, training & inference performance end to endDebugging and optimization... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $193.3k - $261.5k

     ...PyTorch and JAX enabling unparalleled ML inference and training performance.The Inference Enablement...  ...till the hardware-software boundary, our engineers build systematic infrastructure, innovate...  ...the boundaries of what's possible in AI acceleration.As part of the broader... 
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $193.3k - $261.5k

     ...accelerators. Join us to optimize the latest models to run really fast on the Trainium hardware.As a Software Development Engineer on the Inference Model Enablement team, you will onboard and optimize state-of-the-art open-source and customer LLMs, both dense and MoE, for... 
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    6 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Inference Engineer - Speech. Be the first to apply!