AI Inference Engineer - Speech
$151.8k - $332.2kZoom Video Communications, Inc.
What you can expect We are looking for an AI Inference Engineer with a solid background in speech recognition and model inference. In this role, you will develop state-of-the-art automatic speech recognition system and ship it to various Zoom products. You will work on the most cutting edge speech modeling and inference technologies with world-class speech scientists. This role will include collaboration with cross-functional teams, including product, science engineering teams, and infrastructure teams, to deliver high-impact projects from the ground up.About the Team Zoom's AI Speech Team is developing speech recognition technologies to improve Zoom's conversational AI experience. This work impacts various products, like Zoom AI Companion, Zoom Meetings and Workplace, Zoom Contact Center, Zoom Phone, Zoom Revenue Accelerator, etc. Our team's mission is to equip the powerful AI brain with human-level listening and understanding undefined for voice input.As an AI Inference Engineer, you will develop novel speech model inference solutions on modern AI inference hardware, such as GPU, TPU and AI-specific chips. Our goal is to deliver the most unique AI-powered collaboration platform to users across the globe.ResponsibilitiesDeveloping state-of-the-art speech services for Zoom products. Devising novel techniques where off-the-shelf solutions are not available.Optimizing ASR inference systems for production deployment, including inference latency, throughput, memory footprint, and resource utilization.Optimizing model inference performance by diving deep into the lower stack of inference frameworks, with a focus on hardware-specific optimizations for Nvidia GPUs.Proposing new model structures by joint optimization of model accuracy and inference speed.Designing and developing ASR systems with low latency and high accuracy requirements, while ensuring scalability of GPU infrastructure and improving throughput of ASR service.Profiling and debugging ASR runtime performance bottlenecks across different deployment hardware and environments.What we’re looking forPossess a Master's in Computer Science, Electrical Engineering or related fields with 3+ years of experience in speech recognition, speech-llm or AI model inference.Display knowledge in deep learning and hands-on programming skills in Python, shell scripts, C/C++; familiarity with ML frameworks such as PyTorch and TensorFlow.Demonstrate deep understanding of transformer encoder-decoder frameworks for speech recognition, including attention mechanisms, beam search and sequence-to-sequence modeling for end-to-end ASR systems.Understand recent advancements in speech foundation models and speech-LLMs that integrate acoustic and linguistic representations, enabling unified modeling for speech understanding and transcription tasks.Have experience in optimizing deep learning model inference on NVIDIA GPUs, including profiling and accelerating AI models using CUDA, TensorRT, and mixed-precision computation to achieve low latency, high-throughput performance.Have experience developing and tuning custom CUDA kernels, leveraging CUDA Graphs for efficient execution scheduling, and minimizing kernel launch overhead to maximize GPU utilization.Be proficient in end-to-end performance analysis, memory optimization, and deployment of largescale ML models on GPU clusters. Experienced with stream management, asynchronous execution, and integrating frameworks such as PyTorch and TensorFlow for real-time inference.Salary Range or On Target Earnings:Minimum:$151,800.00Maximum:$332,200.00In addition to the base salary and/or OTE listed Zoom has a Total Direct Compensation philosophy that takes into consideration; base salary, bonus and equity value.Note: Starting pay will be based on a number of factors and commensurate with qualifications & experience.We also have a location based compensation structure; there may be a different range for candidates in this and other locationsAt Zoom, we offer a window of at least 5 days for you to apply because we believe in giving you every opportunity. Below is the potential closing date, just in case you want to mark it on your calendar. We look forward to receiving your application!Anticipated Position Close Date:08/12/26Ways of WorkingOur structured hybrid approach is centered around our offices and remote work environments. The work style of each role, Hybrid, Remote, or In-Person is indicated in the job description/posting.BenefitsAs part of our award-winning workplace culture and commitment to delivering happiness, our benefits program offers a variety of perks, benefits, and options to help employees maintain their physical, mental, emotional, and financial health; support work-life balance; and contribute to their community in meaningful ways. Click Learnfor more information.About UsZoomies help people stay connected so they can get more done together. We set out to build the best collaboration platform for the enterprise, and today help people communicate better with products like Zoom Contact Center, Zoom Phone, Zoom Events, Zoom Apps, Zoom Rooms, and Zoom Webinars.We’re problem-solvers, working at a fast pace to design solutions with our customers and users in mind. Find room to grow with opportunities to stretch your skills and advance your career in a collaborative, growth-focused environment.Our CommitmentAt Zoom, we believe great work happens when people feel supported and empowered. We’re committed to fair hiring practices that ensure every candidate is evaluated based on skills, experience, and potential. If you require an accommodation during the hiring process, let us know—we’re here to support you at every step.If you need assistance navigating the interview process due to a medical disability, please submit an Accommodations Request Form and someone from our team will reach out soon. This form is solely for applicants who require an accommodation due to a qualifying medical disability. Non-accommodation-related requests, such as application follow-ups or technical issues, will not be addressed.Our interviews are supported by BrightHire, a tool that helps us create a consistent and thoughtful interview experience and may include recordings. Please refer to our candidate privacy statement for more information of how we use your data.SummaryLocation: Seattle (WA); San Jose (CA)Type: Full time
$177.1k - $387.5k
...the platform that powers Zoom AI Services, enabling AI capabilities... ..., and AI platform engineering to build reliable, high-performance services that power speech, translation, summarization, reasoning... ..., including model serving, AI inference platforms, GPU/CPU resource management...SuggestedFull timeWork at officeRemote work- ...prioritize a diverse F5 community where each individual can thrive.AI Engineer — Customer Success & Services (F5) Location: Hybrid (San Jose... ...and operate production-quality ML services and APIs (scalable inference, caching, batching, latency SLAs); write performant, well-...SuggestedFull timeLocal area
$188k - $275k
...Description CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers,... ...'ll Do Description of the team: The Inference team is responsible for delivering high-performance... ...: We are looking for an Applied AI Engineer to help us understand, measure, and...SuggestedPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours- ...As a Model Optimization & Deployment Engineer, you will focus on bringing highly efficient... ...CUDA kernels, and build highly concurrent inference code to ensure real-time, deterministic execution... ...latency and maximize memory bandwidth on AI accelerators. Write production-level,...SuggestedTemporary workRelocation package
$151.8k
What you can expect We are seeking an experienced AI Infrastructure Engineer to join our AI Incubation team. You will be focused on building and... ...with AI researchers to implement efficient training and inference pipelines What we’re looking for Have a bachelor's degree...SuggestedFull timeWork at officeRemote work$168.1k - $227.4k
...machinelearning accelerators. This role is for a senior software engineer in the Machine Learning Inference Applications team. This role is responsible for... ...deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware.Amazon is an equal opportunity...Work experience placementFlexible hours$92k - $135k
...CoreWeave is the AI Hyperscaler™, delivering a cloud platform of cutting edge services... .... What You’ll Do: Join the Inference team to ship production features that improve... ...quickly with mentorship from experienced engineers. About the role: Implement well...Permanent employmentFull timeTemporary workCasual workInternshipWork at officeRemote workFlexible hours$300k
...create reliable, interpretable, and steerable AI systems. We want AI to be safe and... ...growing group of committed researchers, engineers, policy experts, and business leaders working... .... About the role Our Inference team is responsible for building and maintaining...Full timeWork at officeWorldwideVisa sponsorshipFlexible hours$320k
...create reliable, interpretable, and steerable AI systems. We want AI to be safe and... ...growing group of committed researchers, engineers, policy experts, and business leaders working... ...the Role Our mandate is to make inference deployment boring and unattended. Anthropic...Full timeWork at officeVisa sponsorshipFlexible hoursShift work$140k - $180k
...WashingtonValorem Reply Azure SI, US - Data + AI /Full Time /HybridValorem Reply is an... ...clients do business. As a Senior AI Engineer, you will understand how AI is positioned... ...generation (RAG), and Azure AI Services (OpenAI, Speech, Vision, etc.) using tools like Azure AI...Full time$152k - $241.5k
We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join our AI Efficiency efforts. As an Engineer, you will... ...tools Experience investigating, and resolving, training & inference performance end to endDebugging and optimization...Full timeRemote work$163.2k - $220.8k
...and career growth. Wilson Sonsini is looking for a Senior AI Security Engineer to join the Security Operations team. The Senior AI Security... ...secrets management for model API keys, network isolation for AI inference endpoints, and identity-aware proxy patterns for LLM access...Full timeWork experience placementRemote workWorldwideShift work- Job Title: Applied AI EngineerLocation: 100% RemoteFull-Time Role Role As an Applied AI Engineer, you will turn model capabilities into real product behavior.You will own... ...(OpenAI-style APIs, LLaMA, Qwen, etc.)Inference / serving (e.g. vLLM)Vector DB Ideal Experience...
- We Are:The Global AI Infrastructure team is at the center of enabling infrastructure reinvention for the next era of digital solutions... ...Manager (BCM), NGC, NCCL, NVLink, and CUDA along with LLM inference engines (TensorRT-LLM), production serving frameworks (vLLM, SGLang),...Full timeWork experience placementLive inWork at officeLocal area
$151.8k - $332.2k
...What You Can ExpectYou'll design, implement, and own the inference systems that serve Zoom's AI models at production scale -- across real-time... ...implementationsDrive technical design and set the bar for inference engineering practices across the teamWhat We're Looking ForA...Full timeWork at officeRemote work$177.1k - $387.5k
What you can expectAs an AI Engineer specializing in Agentic AI, you will develop intelligent agents that can autonomously perceive, reason... ...various AI components including LLM, RAG, TTS (Text-to-Speech), and ASR (Automatic Speech Recognition) systems.Writing and optimizing...Full timeWork at officeRemote work$176.76k - $232k
...people.About this team The Enterprise Data & AI team is a strategic and operational... ....Core responsibilities As a Senior AI/ML Engineer, you will lead the delivery of scalable AI... ...architectures and system design for serving AI/ML inference solutions in production. You will help...Permanent employmentFull timeContract workPart timeWork visa$143.7k - $194.4k
...commerce.Advertiser Growth Tech (AGT) is an engineering team with the mission to enhance the... ...scalable code. You will work on Agentic AI initiatives, contributing to systems that... ..., model training, optimization, and inference infrastructure. Lead technical design discussions...InternshipFlexible hours$168.75k
....Your ImpactAs a Staff Embedded Software Engineer, you will lead critical software engineering... ...on connected devices, influencing how AI models are trained, deployed, evaluated,... ...including AI-enabled systems spanning on-device inference and cloud-assisted workflows.Lead the...Work experience placementWork at officeRemote work- ...in how businesses manage networks. There AI Core group pioneers’ platforms across Generative... .... The Role As one of our AI ML Engineer’s, you'll be a key technical leader and... ...web services Build real-time inference pipelines for complex models using Triton...Full timeShift work
$91.1k - $179.5k
...significant impact on our clients’ success. We are hiring an AI Engineer to build and operate the data, features, and GenAI foundations... ...pipelines and services that support model training, real-time inference, and LLM applications using Claude-, GPT/Codex-, and Gemini-class...Local area$300k
...create reliable, interpretable, and steerable AI systems. We want AI to be safe and... ...growing group of committed researchers, engineers, policy experts, and business leaders working... .... About the Role The Cloud Inference team scales and optimizes Claude to serve...Full timeWork at officeVisa sponsorshipFlexible hours$130k - $170k
CompanyFounded by CPAs, tax attorneys, and engineers, Taxbit is the leading innovator automating global tax reporting for the digital economy. Taxbit's AI-enabled platform streamlines compliance related to digital assets, payments, and other financial transactions. Its...Work at officeWork from home$119.1k - $184.7k
...agreements with solutions created by the #1 company in e-signature and contract lifecycle management (CLM).What you'll doAs an AI Agentic Engineer at Docusign, you will transform IT operations by designing and deploying autonomous AI agents that proactively resolve user...Permanent employmentFull timeContract workWork at officeLocal areaRemote work2 days per week$171k - $240k
...gain real-time visibility, and control spend effortlessly. Brex’s AI-native automation and world-class service eliminate manual... ...resources, and support you need to grow your career.AI at BrexAI Engineering at Brex is redefining how businesses run their finances by building...Work at officeRemote workWork from homeShift work$157.4k - $236k
...EDS) team designs and delivers enterprise-scale platforms and services across Knowledge Management, Artificial Intelligence (AI), Data Engineering, Business Intelligence, Data Infrastructure, and Health Data Platforms. The EDS team’s mission is to empower the...H1bWork at officeImmediate start- ...Vamsi SattaruCompany: SRI Tech SolutionsPosition Title: Agentic AI Software EngineerLocation: Hybrid, Seattle, WAEmployment Type:... ...Minimum Qualifications: Bachelor’s degree in Computer Science, Engineering, or a closely related discipline. Demonstrated experience in software...Full time
$152k - $241.5k
...intelligence. Our technology powers everything from generative AI to autonomous systems, and we continue to shape the future of computing... ..., platforms, and tools that enable researchers and engineers to develop the next generation of AI/ML systems. By joining us,...Full time$137.4k - $161.7k
Are you ready to make an impact?Senior AI Software Engineer - Technology & Experience (TechEx) West Monroe is seeking a Senior AI Software Engineer to join our Technology & Experience (TechEx) practice. This is a hands-on engineering role focused on building scalable,...Local areaImmediate startFlexible hours- Job ID: 512905Location: Bellevue, Washington, United States of AmericaCompany: Siemens AI Engineer - Principal Job ID 512905 Posted since 18-Jul-2026 Organization Digital Industries Field of work Research & Development Company Siemens Industry...Permanent employmentFull timeWork at officeLocal areaRemote workWork from home
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Inference Engineer - Speech. Be the first to apply!

