AI Inference Engineer - Speech
$151.8k - $332.2kZoom Video Communications, Inc.
What you can expect
We are looking for an AI Inference Engineer with a solid background in speech recognition and model inference. In this role, you will develop state-of-the-art automatic speech recognition system and ship it to various Zoom products. You will work on the most cutting edge speech modeling and inference technologies with world-class speech scientists. This role will include collaboration with cross-functional teams, including product, science engineering teams, and infrastructure teams, to deliver high-impact projects from the ground up.
About the Team
Zoom's AI Speech Team is developing speech recognition technologies to improve Zoom's conversational AI experience. This work impacts various products, like Zoom AI Companion, Zoom Meetings and Workplace, Zoom Contact Center, Zoom Phone, Zoom Revenue Accelerator, etc. Our team's mission is to equip the powerful AI brain with human-level listening and understanding undefined for voice input.
As an AI Inference Engineer, you will develop novel speech model inference solutions on modern AI inference hardware, such as GPU, TPU and AI-specific chips. Our goal is to deliver the most unique AI-powered collaboration platform to users across the globe.
Responsibilities
- Developing state-of-the-art speech services for Zoom products. Devising novel techniques where off-the-shelf solutions are not available.
- Optimizing ASR inference systems for production deployment, including inference latency, throughput, memory footprint, and resource utilization.
- Optimizing model inference performance by diving deep into the lower stack of inference frameworks, with a focus on hardware-specific optimizations for Nvidia GPUs.
- Proposing new model structures by joint optimization of model accuracy and inference speed.
- Designing and developing ASR systems with low latency and high accuracy requirements, while ensuring scalability of GPU infrastructure and improving throughput of ASR service.
- Profiling and debugging ASR runtime performance bottlenecks across different deployment hardware and environments.
What we’re looking for
- Possess a Master's in Computer Science, Electrical Engineering or related fields with 3+ years of experience in speech recognition, speech-llm or AI model inference.
- Display knowledge in deep learning and hands-on programming skills in Python, shell scripts, C/C++; familiarity with ML frameworks such as PyTorch and TensorFlow.
- Demonstrate deep understanding of transformer encoder-decoder frameworks for speech recognition, including attention mechanisms, beam search and sequence-to-sequence modeling for end-to-end ASR systems.
- Understand recent advancements in speech foundation models and speech-LLMs that integrate acoustic and linguistic representations, enabling unified modeling for speech understanding and transcription tasks.
- Have experience in optimizing deep learning model inference on NVIDIA GPUs, including profiling and accelerating AI models using CUDA, TensorRT, and mixed-precision computation to achieve low latency, high-throughput performance.
- Have experience developing and tuning custom CUDA kernels, leveraging CUDA Graphs for efficient execution scheduling, and minimizing kernel launch overhead to maximize GPU utilization.
- Be proficient in end-to-end performance analysis, memory optimization, and deployment of largescale ML models on GPU clusters. Experienced with stream management, asynchronous execution, and integrating frameworks such as PyTorch and TensorFlow for real-time inference.
Salary Range or On Target Earnings:
Minimum:
$151,800.00Maximum:
$332,200.00In addition to the base salary and/or OTE listed Zoom has a Total Direct Compensation philosophy that takes into consideration; base salary, bonus and equity value.
Note: Starting pay will be based on a number of factors and commensurate with qualifications & experience.
We also have a location based compensation structure; there may be a different range for candidates in this and other locations
At Zoom, we offer a window of at least 5 days for you to apply because we believe in giving you every opportunity. Below is the potential closing date, just in case you want to mark it on your calendar. We look forward to receiving your application!
Anticipated Position Close Date:
09/01/26 Ways of Working
Our structured hybrid approach is centered around our offices and remote work environments. The work style of each role, Hybrid, Remote, or In-Person is indicated in the job description/posting.
Benefits
As part of our award-winning workplace culture and commitment to delivering happiness, our benefits program offers a variety of perks, benefits, and options to help employees maintain their physical, mental, emotional, and financial health; support work-life balance; and contribute to their community in meaningful ways. Click Learnfor more information.
About Us
Zoomies help people stay connected so they can get more done together. We set out to build the best collaboration platform for the enterprise, and today help people communicate better with products like Zoom Contact Center, Zoom Phone, Zoom Events, Zoom Apps, Zoom Rooms, and Zoom Webinars.
We’re problem-solvers, working at a fast pace to design solutions with our customers and users in mind. Find room to grow with opportunities to stretch your skills and advance your career in a collaborative, growth-focused environment.
Our Commitment
At Zoom, we believe great work happens when people feel supported and empowered. We’re committed to fair hiring practices that ensure every candidate is evaluated based on skills, experience, and potential. If you require an accommodation during the hiring process, let us know—we’re here to support you at every step.
If you need assistance navigating the interview process due to a medical disability, please submit an Accommodations Request Form and someone from our team will reach out soon. This form is solely for applicants who require an accommodation due to a qualifying medical disability. Non-accommodation-related requests, such as application follow-ups or technical issues, will not be addressed.
Our interviews are supported by BrightHire, a tool that helps us create a consistent and thoughtful interview experience and may include recordings. Please refer to our candidate privacy statement for more information of how we use your data.
$176.6k - $265k
...customers, better. And it means we prioritize a diverse F5 community where each individual can thrive. Job Description The AI Inference Engineer plays a critical role in the AI lifecycle by bridging the gap between high-performance model development and optimized...SuggestedFull timeLocal areaImmediate start$188k - $275k
...CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave... ...Do Description of the team: The Inference team is responsible for delivering high-performance... ...: We are looking for an Applied AI Engineer to help us understand, measure, and...SuggestedPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours- NVIDIA in Seattle, WA seeks outstanding AI systems engineers to advance the inference software stack for AI workloads. You will develop libraries, code generators, and GPU kernel technologies for NVIDIA hardware, including new abstractions and runtimes for large language...Suggested
$152k - $241.5k
...deep learning ignited modern AI — the next era of computing —... ...is hiring a Senior AI Compiler Engineer. GPUs are driving rapid progress... ...recommendation, vision, and speech. On this team, you’ll build an... ...compiler that powers NVIDIA’s inference engine end to end, with a focus...SuggestedFull timeRemote work$134.96k - $188.95k
...manual monitoring and reactive processes. We are seeking an AI Engineer II to join the Blue Ring software team. You will contribute to... ...model compression, edge deployment, or resource-constrained inference Base Pay Range for: VA applicants is $134,961.00 - $188,...SuggestedPermanent employmentFull timeTemporary workLocal area$177.1k - $387.5k
...the platform that powers Zoom AI Services, enabling AI capabilities... ..., and AI platform engineering to build reliable, high-performance services that power speech, translation, summarization, reasoning... ..., including model serving, AI inference platforms, GPU/CPU resource management...Full timeWork at officeRemote work- ...As a Model Optimization & Deployment Engineer, you will focus on bringing highly efficient... ...CUDA kernels, and build highly concurrent inference code to ensure real-time, deterministic execution... ...latency and maximize memory bandwidth on AI accelerators. Write production-level,...Temporary workRelocation package
- ...Senior Staff AI Engineer The AI Scientist will work in teams addressing statistical, machine learning and data understanding problems... ...clinical reports, and other healthcare datasets. Develop robust inference pipelines, model serving infrastructure, and AI services...WorldwideFlexible hours
$154.56k - $193.2k
...hyperscaler for the edge, delivering modular AI infrastructure from first deployment to... ...Armada is seeking exceptional AI Engineers to build and deploy intelligent systems at... ...analysis, anomaly detection, or distributed AI inference. This role is intended for engineers...Work at officeRemote workFlexible hours$168.1k - $227.4k
...machinelearning accelerators. This role is for a senior software engineer in the Machine Learning Inference Applications team. This role is responsible for... ...deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware.Amazon is an equal opportunity...Work experience placementFlexible hours$92k - $135k
...CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave... ...more at What You’ll Do: Join the Inference team to ship production features that improve... ...quickly with mentorship from experienced engineers. About the role: Implement...Permanent employmentFull timeTemporary workCasual workInternshipWork at officeFlexible hours$139k - $204k
...CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave... ...Learn more at What You’ll Do: Senior engineers are area owners who lead designs, raise... ...hardware teams to evolve our Kubernetes-native inference platform and meet strict P99 SLAs at...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hoursShift work$193.3k - $261.5k
...cloud-scale machine learning. This senior software engineering role is part of the Machine Learning Inference Applications team and focuses on delivering high-performance... ...in production on GPUs, AWS Neuron, TPUs, or other AI accelerator hardware- Experience extending or...Work experience placementLocal areaFlexible hours$152k - $241.5k
We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join our AI Efficiency efforts. As an Engineer, you will... ...tools Experience investigating, and resolving, training & inference performance end to endDebugging and optimization...Full timeRemote work$140k - $180k
...WashingtonValorem Reply Azure SI, US - Data + AI /Full Time /HybridValorem Reply is an... ...clients do business. As a Senior AI Engineer, you will understand how AI is positioned... ...generation (RAG), and Azure AI Services (OpenAI, Speech, Vision, etc.) using tools like Azure AI...Full time$151.8k - $332.2k
What you can expectWe are seeking an experienced AI Infrastructure Engineer to join our AI Incubation team. You will be focused on building and... ...workflowsCollaborating with AI researchers to implement efficient training and inference pipelinesWhat we’re looking forHave a bachelor's degree in...Full timeWork at officeRemote work$151.8k - $332.2k
...What You Can ExpectYou'll design, implement, and own the inference systems that serve Zoom's AI models at production scale -- across real-time... ...implementationsDrive technical design and set the bar for inference engineering practices across the teamWhat We're Looking ForA...Full timeWork at officeRemote work$177.1k - $387.5k
What you can expectAs an AI Engineer specializing in Agentic AI, you will develop intelligent agents that can autonomously perceive, reason... ...various AI components including LLM, RAG, TTS (Text-to-Speech), and ASR (Automatic Speech Recognition) systems.Writing and optimizing...Full timeWork at officeRemote work- We Are:The Global AI Infrastructure team is at the center of enabling infrastructure reinvention for the next era of digital solutions... ...Manager (BCM), NGC, NCCL, NVLink, and CUDA along with LLM inference engines (TensorRT-LLM), production serving frameworks (vLLM, SGLang),...Full timeWork experience placementLive inWork at officeLocal area
$163.2k - $220.8k
...and career growth. Wilson Sonsini is looking for a Senior AI Security Engineer to join the Security Operations team. The Senior AI Security... ...secrets management for model API keys, network isolation for AI inference endpoints, and identity-aware proxy patterns for LLM access...Full timeWork experience placementRemote workWorldwideShift work- Job Title: Applied AI EngineerLocation: 100% RemoteFull-Time Role Role As an Applied AI Engineer, you will turn model capabilities into real product behavior.You will own... ...(OpenAI-style APIs, LLaMA, Qwen, etc.)Inference / serving (e.g. vLLM)Vector DB Ideal Experience...
- ...Senior AI Engineer – Privacy Location: Bellevue WA Must have skills – skill 1 – 7yrs of exp – AI Engineer – Privacy skill 2 – 7yrs... ...Snowflake, or PySpark to support AI model training, fine-tuning, and inference. Apply prompt engineering, few-shot learning, and fine-...
$143.7k - $194.4k
...commerce.Advertiser Growth Tech (AGT) is an engineering team with the mission to enhance the... ...scalable code. You will work on Agentic AI initiatives, contributing to systems that... ..., model training, optimization, and inference infrastructure. Lead technical design discussions...InternshipFlexible hours$176.76k - $232k
...people.About this team The Enterprise Data & AI team is a strategic and operational... ....Core responsibilities As a Senior AI/ML Engineer, you will lead the delivery of scalable AI... ...architectures and system design for serving AI/ML inference solutions in production. You will help...Permanent employmentFull timeContract workPart timeWork visa- ...the forefront of a new era in enterprise AI — one defined not by model capability... ...of frontier AI research and production engineering — investigating the foundational challenges... ...persistence architectures, model selection and inference routing strategies, autonomy and goal-...Full timeWork experience placementLive inWork at officeLocal areaRelocation
$220k - $292k
...of systems is powered by Lattice OS, an AI-powered operating system that turns thousands... ..., Motion Planning, Hardware, and Test Engineering to solve some of the hardest problems... ...TPUs, or custom ASICs) for training and inference workloads. Model Observability at Scale:...Full timeWork experience placementImmediate start$106.9k - $200.6k
...better working world. The opportunity We are seeking AI Systems Engineers to own the security and trust fabric of EY’s AI-native... ...workloads and the specific trust challenges of confidential AI inference (models/secrets inside enclaves). Experience producing...Full timeWork experience placementSummer holidayRemote workFlexible hours$106.9k - $200.6k
...working world. The opportunity We are seeking an AI Systems Engineer to own the delivery, model-serving, routing, and observability... ...automated delivery pipelines, operating high-performance inference (GPUs, model servers, sandboxed execution), and building deep...Full timeSummer holidayFlexible hours$220k - $293.33k
...software, and applications that power today’s AI stack using sustainable technology... ...Nscale is looking for a Staff AI Product Engineer to drive technical direction across the... ...Experience building AI/ML product platforms: inference APIs, fine-tuning UX, model management,...Full timeContract workFlexible hours- ...Lead AI Engineer in the Platforms and Products ZS is a place where passion changes lives. As a management consulting and technology... ...retrieval-augmented generation, fine-tuning workflows, and scalable inference pipelines. Design and implement LLM-powered applications...Local areaWork from homeWorldwideFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Inference Engineer - Speech. Be the first to apply!

