Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Inference Engineer - Speech

$151.8k - $332.2k

Zoom Video Communications, Inc.

What you can expect We are looking for an AI Inference Engineer with a solid background in speech recognition and model inference. In this role, you will develop state-of-the-art automatic speech recognition system and ship it to various Zoom products. You will work on the most cutting edge speech modeling and inference technologies with world-class speech scientists. This role will include collaboration with cross-functional teams, including product, science engineering teams, and infrastructure teams, to deliver high-impact projects from the ground up.About the Team Zoom's AI Speech Team is developing speech recognition technologies to improve Zoom's conversational AI experience. This work impacts various products, like Zoom AI Companion, Zoom Meetings and Workplace, Zoom Contact Center, Zoom Phone, Zoom Revenue Accelerator, etc. Our team's mission is to equip the powerful AI brain with human-level listening and understanding undefined for voice input.As an AI Inference Engineer, you will develop novel speech model inference solutions on modern AI inference hardware, such as GPU, TPU and AI-specific chips. Our goal is to deliver the most unique AI-powered collaboration platform to users across the globe.ResponsibilitiesDeveloping state-of-the-art speech services for Zoom products. Devising novel techniques where off-the-shelf solutions are not available.Optimizing ASR inference systems for production deployment, including inference latency, throughput, memory footprint, and resource utilization.Optimizing model inference performance by diving deep into the lower stack of inference frameworks, with a focus on hardware-specific optimizations for Nvidia GPUs.Proposing new model structures by joint optimization of model accuracy and inference speed.Designing and developing ASR systems with low latency and high accuracy requirements, while ensuring scalability of GPU infrastructure and improving throughput of ASR service.Profiling and debugging ASR runtime performance bottlenecks across different deployment hardware and environments.What we’re looking forPossess a Master's in Computer Science, Electrical Engineering or related fields with 3+ years of experience in speech recognition, speech-llm or AI model inference.Display knowledge in deep learning and hands-on programming skills in Python, shell scripts, C/C++; familiarity with ML frameworks such as PyTorch and TensorFlow.Demonstrate deep understanding of transformer encoder-decoder frameworks for speech recognition, including attention mechanisms, beam search and sequence-to-sequence modeling for end-to-end ASR systems.Understand recent advancements in speech foundation models and speech-LLMs that integrate acoustic and linguistic representations, enabling unified modeling for speech understanding and transcription tasks.Have experience in optimizing deep learning model inference on NVIDIA GPUs, including profiling and accelerating AI models using CUDA, TensorRT, and mixed-precision computation to achieve low latency, high-throughput performance.Have experience developing and tuning custom CUDA kernels, leveraging CUDA Graphs for efficient execution scheduling, and minimizing kernel launch overhead to maximize GPU utilization.Be proficient in end-to-end performance analysis, memory optimization, and deployment of largescale ML models on GPU clusters. Experienced with stream management, asynchronous execution, and integrating frameworks such as PyTorch and TensorFlow for real-time inference.Salary Range or On Target Earnings:Minimum:$151,800.00Maximum:$332,200.00In addition to the base salary and/or OTE listed Zoom has a Total Direct Compensation philosophy that takes into consideration; base salary, bonus and equity value.Note: Starting pay will be based on a number of factors and commensurate with qualifications & experience.We also have a location based compensation structure; there may be a different range for candidates in this and other locationsAt Zoom, we offer a window of at least 5 days for you to apply because we believe in giving you every opportunity. Below is the potential closing date, just in case you want to mark it on your calendar. We look forward to receiving your application!Anticipated Position Close Date:09/23/26Ways of WorkingOur structured hybrid approach is centered around our offices and remote work environments. The work style of each role, Hybrid, Remote, or In-Person is indicated in the job description/posting.BenefitsAs part of our award-winning workplace culture and commitment to delivering happiness, our benefits program offers a variety of perks, benefits, and options to help employees maintain their physical, mental, emotional, and financial health; support work-life balance; and contribute to their community in meaningful ways. Click Learnfor more information.About UsZoomies help people stay connected so they can get more done together. We set out to build the best collaboration platform for the enterprise, and today help people communicate better with products like Zoom Contact Center, Zoom Phone, Zoom Events, Zoom Apps, Zoom Rooms, and Zoom Webinars.We’re problem-solvers, working at a fast pace to design solutions with our customers and users in mind. Find room to grow with opportunities to stretch your skills and advance your career in a collaborative, growth-focused environment.Our CommitmentAt Zoom, we believe great work happens when people feel supported and empowered. We’re committed to fair hiring practices that ensure every candidate is evaluated based on skills, experience, and potential. If you require an accommodation during the hiring process, let us know—we’re here to support you at every step.If you need assistance navigating the interview process due to a medical disability, please submit an Accommodations Request Form and someone from our team will reach out soon. This form is solely for applicants who require an accommodation due to a qualifying medical disability. Non-accommodation-related requests, such as application follow-ups or technical issues, will not be addressed.Our interviews are supported by BrightHire, a tool that helps us create a consistent and thoughtful interview experience and may include recordings. Please refer to our candidate privacy statement for more information of how we use your data.SummaryLocation: Seattle (WA); San Jose (CA)Type: Full time

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the AI Inference Engineer - Speech in San Jose, CA vacancy
  •  ...their customers, better. And it means we prioritize a diverse F5 community where each individual can thrive.Job DescriptionThe AI Inference Engineer plays a critical role in the AI lifecycle by bridging the gap between high-performance model development and optimized... 
    Suggested
    Full time
    Local area
    Immediate start

    F5 Networks

    San Jose, CA
    5 days ago
  • $152k - $241.5k

     ...deep learning ignited modern AI — the next era of computing —...  ...is hiring a Senior AI Compiler Engineer. GPUs are driving rapid progress...  ...recommendation, vision, and speech. On this team, you’ll build an...  ...compiler that powers NVIDIA’s inference engine end to end, with a focus... 
    Suggested
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $274k - $300k

     ...Job Description Job Description Saviynt's AI-powered identity platform manages and governs human and non-human access...  .... For more information, please visit  AI Platform Engineer – Training & Inference Saviynt's AI-powered identity platform manages and governs... 
    Suggested

    Saviynt

    Milpitas, CA
    11 days ago
  • $152k - $241.5k

     ...into the unlimited potential of AI to define the next era of...  ...an AI & Deep Learning Compiler Engineer. NVIDIA is hiring software engineers...  ..., image classification, speech recognition, etc. With the rapid...  ...been the backbone of NVIDIA’s inference engine, spanning across data centers... 
    Suggested
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...d-Matrix, headquartered in Santa Clara, CA, seeks a Principal System Software Engineer for AI Inference Execution. You will join the software team to productize the AI compute engine's SW stack, developing deployment software and collaborating with ML, compiler, and hardware... 
    Suggested

    Jobleads-US

    Santa Clara, CA
    6 days ago
  • $229.9k - $262.4k

     ...Overview AI Engineer 5 (FM Hosting, LLM Inference) At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer... 
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    2 days ago
  • $160k - $198k

     ...team members.What You’ll DoAs a Senior AI Systems Engineer, you will architect, deploy, and...  ...for large-scale AI model training and inference. You will ensure our machine learning...  ...QualificationsFamiliarity with audio processing, speech-to-text frameworks, or Automatic... 
    Local area

    Archer Aviation

    San Jose, CA
    1 day ago
  • $117.7k - $221.4k

     ...practical, and cost efficient for embodied AI systems. We believe the next generation...  .... This operating model reflects how Cola engineers think: build durable intermediate artifacts...  ...the data processing, featurization, and inference foundations that power scalable world... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    22 hours ago
  • $190k - $237k

     ...team members.What You’ll DoAs a Staff AI Systems Engineer, you will architect, deploy, and manage...  ...for large-scale AI model training and inference. You will ensure our machine learning...  ...QualificationsFamiliarity with audio processing, speech-to-text frameworks, or Automatic... 
    Local area

    Archer Aviation

    San Jose, CA
    4 days ago
  • $152k - $241.5k

     ...deep learning ignited modern AI — the next era of computing —...  ...an AI & Deep Learning Compiler Engineer. NVIDIA is hiring software engineers...  ..., image classification, speech recognition, etc. With the rapid...  ...been the backbone of NVIDIA’s inference engine, spanning across data... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...next-generation computing experiences—from AI and data centers, to PCs, gaming and...  ...advance your career. THE ROLEWe are hiring AI Engineers to build recursive self-improvement...  ...JAX, TensorFlow, or distributed training/inference systems.Experience with reinforcement learning... 

    AMD

    Santa Clara, CA
    4 days ago
  • $100k

     ...is leading the industry on cutting-edge AI technology, revolutionizing performance expectations...  ...Speed Interconnect / Signal Integrity Engineer to design and validate high-bandwidth...  ...technologies for next-generation AI inference and training clusters. This role is on-site... 
    Permanent employment

    Tenstorrent

    Santa Clara, CA
    3 days ago
  • $215k - $260k

     ...intelligence . As the only vertically integrated AI infrastructure company built from the...  ...in production. That means owning the inference stack end to end: profiling where time...  ...will also work directly with customer engineering teams to tailor deployments to their needs... 
    Temporary work

    Crusoe

    Sunnyvale, CA
    27 days ago
  • $160k - $198k

     ...works across air taxis, UAS, AI and powertrain development — building...  .... We’re seeking exceptional engineers, operators and builders to...  ...large-scale AI model training and inference. You will ensure our machine...  ...with audio processing, speech-to-text frameworks, or Automatic... 
    Local area
    Visa sponsorship
    Night shift

    Archer Aviation

    San Jose, CA
    a month ago
  •  ...Job Description In this AI/ML ASIC Performance Engineer position, you will develop AI Storage...  ...bandwidth Architect memory-efficient inference/training systems utilizing techniques...  ...multiple modalities (text, vision, speech) KV cache optimization, Flash Attention... 
    Full time
    Temporary work
    Remote work
    Flexible hours
    Shift work
    Night shift

    Sandisk

    Milpitas, CA
    11 days ago
  • $2,000 per month

     ...About Etched Etched is building AI chips that are hard-coded for individual model...  ...continuous batching and real time inference Implement inference-time acceleration...  ...person team in Cupertino, and greatly value engineering skills. We do not have boundaries between... 
    Full time
    Work at office
    Relocation package

    Etched

    Cupertino, CA
    1 day ago
  •  ...Embodied AI Engineer UnitX builds the world's leading physical AI systems to automate repetitive visual tasks in factories. UnitX is...  ...deploying, profiling, and optimizing ML models for real-time inference on robotic hardware (e.g., NVIDIA Jetson, TensorRT, CUDA). ~... 

    UnitX

    Milpitas, CA
    3 days ago
  • $184k - $287.5k

    NVIDIA seeks a Senior Software Engineer specializing in Deep Learning Inference for our growing team. As a key contributor, you will help design, build, and...  ...accelerated software that powers today’s most sophisticated AI applications. Our team is responsible for developing... 
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  • $152k - $241.5k

    We are looking for outstanding Senior High Performance AI Engineers to build the next generation of agentic AI systems for the CUDA ecosystem...  ...across NVIDIA's software and hardware stack, from models and inference through compilers, runtimes, libraries, kernels, and GPUs.... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...next-generation computing experiences—from AI and data centers, to PCs, gaming and...  ...looking for a Forward Deployed Research Engineer to build, evaluate, and deploy cutting-edge...  ...ModelingModel or Agent EvaluationAI Training or Inference InfrastructureDeep understanding of... 

    AMD

    Santa Clara, CA
    5 days ago
  • $152k - $241.5k

    We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join our AI Efficiency efforts. As an Engineer, you will...  ...tools Experience investigating, and resolving, training & inference performance end to endDebugging and optimization... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

    We are looking for a strong engineer to join the DRIVE Road Structure / Online Mapping / Context...  ...you interested in inventing human-level AI for navigation in the unconstrained world...  ...-precision techniques, and efficient GPU inference using NVIDIA software and hardware is... 
    Full time
    Remote work
    Shift work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $2,000 per month

     ...Applied AI Engineer, Silicon Engineering About Etched Etched is building AI chips that are hard-coded for individual model architectures...  ...ship yourself. It is not a customer-facing role and not about inference serving — it's AI applied to how we build the chip itself. You... 
    Full time
    Work at office
    Relocation package
    Night shift

    Etched

    San Jose, CA
    1 day ago
  • $255.65k - $299k

     ...available for this positionWhat you can expect:As a Senior AI Software Engineer, you will collaborate to design, implement, and optimize AI...  ...algorithms and software applications. You will ensure AI training, inference, deployment, and operation are functional, reliable, and... 
    Full time
    Work at office
    Remote work

    Zoom

    San Jose, CA
    3 days ago
  • $2,000 per month

     ...Our first products are heavily focused on inference . Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is redefining the infrastructure...  ...performance breakthroughs will come from AI systems that can understand model... 
    Full time
    Work at office
    Relocation package

    Etched

    San Jose, CA
    1 day ago
  • $136.3k - $231.7k

     ...into R&D. Our expert teams of physicists, engineers, data scientists and problem-solvers work...  ...QualificationsKLA is seeking a motivated AI Engineer with a growth mindset to join a...  ...analysisImprove model accuracy; optimize models for inference throughput, including GPU-accelerated and... 
    Minimum wage
    Full time
    Work experience placement
    Flexible hours

    KLA-Tencor

    Milpitas, CA
    4 days ago
  • $264.51k - $332.2k

     ...experience in job offered or related occupation. Must have 4 years of experience in the following:4 years of experience in ASR (Automatic Speech Recognition) for live captioning;4 years of experience in Conformer model structure, RNN-T (Recursive Neural Network – Transducer)... 
    Full time
    Work at office
    Remote work

    Zoom

    San Jose, CA
    4 days ago
  • $250.8k - $286.2k

    Senior Lead AI Engineer (MLX Emerging AI Patterns) Overview: At Capital One, we are creating responsible and reliable AI systems, changing...  ...including foundation model training, large language model inference, similarity search, guardrails, model evaluation, experimentation... 
    Full time
    Part time
    Local area

    Capital One Financial Corporation

    San Jose, CA
    1 day ago
  • $193.3k - $261.5k

     ...accelerators. Join us to optimize the latest models to run really fast on the Trainium hardware.As a Sr. Software Development Engineer on the Inference Model Enablement team, you will onboard and optimize state-of-the-art open-source and customer LLMs, both dense and MoE,... 
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    4 days ago
  • $165.2k - $223.6k

     ...PyTorch and JAX enabling unparalleled ML inference and training performance.The Inference Enablement...  ...till the hardware-software boundary, our engineers build systematic infrastructure, innovate...  ...the boundaries of what's possible in AI acceleration.As part of the broader... 
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Inference Engineer - Speech. Be the first to apply!