Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Embedded AI Engineer, On-Device Models

Full-time

Deepgram

Company Overview

Deepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for speech-to-text (STT), text-to-speech (TTS), and building production-grade voice agents at scale. More than 200,000 developers and 1,300+ organizations build voice offerings that are ‘Powered by Deepgram’, including Twilio, Cloudflare, Sierra, Decagon, Vapi, Daily, Cresta, Granola, and Jack in the Box. Deepgram’s voice-native foundation models are accessed through cloud APIs or as self-hosted and on-premises software, with unmatched accuracy, low latency, and cost efficiency. Backed by a recent Series C led by leading global investors and strategic partners, Deepgram has processed over 50,000 years of audio and transcribed more than 1 trillion words. There is no organization in the world that understands voice better than Deepgram.

Company Operating Rhythm

At Deepgram, we expect an AI-first mindset—AI use and comfort aren’t optional, they’re core to how we operate, innovate, and measure performance.

Every team member who works at Deepgram is expected to actively use and experiment with advanced AI tools, and even build your own into your everyday work. We measure how effectively AI is applied to deliver results, and consistent, creative use of the latest AI capabilities is key to success here. Candidates should be comfortable adopting new models and modes quickly, integrating AI into their workflows, and continuously pushing the boundaries of what these technologies can do.

Additionally, we move at the pace of AI. Change is rapid, and you can expect your day-to-day work to evolve just as quickly. This may not be the right role if you’re not excited to experiment, adapt, think on your feet, and learn constantly, or if you’re seeking something highly prescriptive with a traditional 9-to-5.

About the Role

Deepgram's speech AI models are among the fastest and most accurate in the world — and the next wave of voice experiences won't live only in the cloud. They'll run directly on the small, low-power devices people carry, wear, and keep around their homes: phones, earbuds, wearables, appliances, cameras, and purpose-built consumer hardware. Putting state-of-the-art speech models on devices with tight memory, compute, thermal, and battery budgets is a fundamentally different engineering problem, and it's one of the most important frontiers for bringing voice AI to everyone.

As an Embedded AI Engineer , you will take Deepgram's models and make them run — fast, accurately, and efficiently — on resource-constrained embedded and edge platforms. You'll work across the stack: optimizing and compiling models for on-device inference, writing performance-critical runtime code, and squeezing every last millisecond and milliwatt out of a wide range of mobile application processors, embedded SoCs, microcontrollers, and dedicated AI accelerators. Your work directly enables a new class of private, offline-capable, real-time voice experiences on the devices closest to the user.

This role is a great fit whether you're a hands-on senior embedded engineer who wants to go deep on a hard problem, or a staff-level technical leader who wants to define how Deepgram's voice AI gets onto consumer hardware and raise the bar for the engineers around you. We'll set the level to your experience.

What You'll Do

  • Take Deepgram's Speech and Conversational models and get them running on embedded and low-power consumer hardware — defining the architecture for on-device, real-time inference across a diverse range of processors and accelerators.

  • Optimize models for constrained targets through quantization, pruning, distillation, operator fusion, and architecture-specific compilation to meet strict latency, memory, power, and thermal budgets.

  • Write and optimize performance-critical runtime code (C, C++, and/or Rust) for embedded environments, including bare-metal and real-time operating systems such as FreeRTOS and Zephyr.

  • Integrate with industry-standard edge inference runtimes and vendor NPU/DSP toolchains, mapping model graphs efficiently onto on-device accelerators and CPU/GPU/NPU heterogeneity.

  • Build the on-device runtime plumbing: model packaging, deployment pipelines, over-the-air update mechanisms, and lightweight telemetry for devices operating with limited or intermittent connectivity.

  • Establish repeatable benchmarking and validation across target hardware — measuring latency, accuracy, power consumption, memory footprint, and resource utilization — and catch regressions before they ship.

  • Partner with silicon and device vendors on SDK integration and performance tuning, getting our models to run efficiently on new chipsets and reference platforms.

  • Collaborate with Research and Engine teams to influence model architectures toward edge-friendly designs from the start, reducing the optimization burden at deployment time.

You'll Love This Role If You

  • Find deep satisfaction in making a large model run on a tiny device — and still hit accuracy and latency targets.

  • Want to work at the intersection of AI and hardware, where optimization isn't optional but existential.

  • Are energized by the back-and-forth of getting a model to sing on a new chipset, runtime, or accelerator.

  • Believe on-device AI is the next major deployment frontier and want to define how speech AI gets there for consumers.

  • Prefer hard, constrained, ship-it problems over open-ended research — you want to see your work running in people's hands.

  • Care about the details that don't show up in a cloud benchmark: cold-start time, power draw, thermals, and memory fragmentation.

It's Important To Us That You Have

  • Experience delivering production systems on resource-constrained hardware — embedded systems, mobile, edge AI, or small low-power devices.

  • Strong proficiency in C, C++, and/or Rust, with experience writing performance-critical code for constrained environments.

  • Hands-on experience with model optimization for on-device deployment, including quantization, pruning, knowledge distillation, or architecture-specific compilation.

  • Familiarity with edge inference runtimes (e.g., ONNX Runtime, TensorRT, TFLite, ExecuTorch) and/or vendor-specific NPU/DSP toolchains.

  • A strong understanding of hardware-software interaction — CPU/GPU/NPU/DSP architectures, memory hierarchies, fixed-point/integer arithmetic, and power management — and how they affect inference performance.

  • Experience working close to the metal: bare-metal or RTOS environments (e.g., FreeRTOS, Zephyr), embedded Linux, or microcontroller and edge SoC development.

  • Strong communication skills and a builder mindset — you can scope an ambiguous optimization problem, drive it to a measurable result, and explain the tradeoffs clearly.

It Would Be Great if You Had

  • Experience with real-time audio processing on embedded platforms — DSP pipelines, audio codec optimization, wake-word or always-on listening, or streaming inference on microcontrollers and edge SoCs.

  • Depth in ML optimization techniques — custom quantization schemes, mixed-precision inference, or neural architecture search for edge targets.

  • Background in hardware evaluation and benchmarking — systematically comparing accelerators, SoCs, or GPUs for specific workload profiles.

  • Experience shipping AI features in consumer products at scale, and the instinct for what "production quality" means on a battery-powered device.

  • Familiarity with model compilation and optimization toolchains and their tradeoffs across hardware targets.

  • Experience with secure, robust on-device deployment practices — code signing, encrypted model storage, and safe update mechanisms.

Vacancy posted 14 hours ago
Similar jobs that could be interesting for youBased on the Embedded AI Engineer, On-Device Models in United States vacancy
  • $219.3k - $274.1k

     ...Description Deepgram's speech AI models are among the fastest and...  ...directly on the small, low-power devices people carry, wear, and keep...  ...a fundamentally different engineering problem, and it's one of the...  ...AI to everyone. As an Embedded AI Engineer , you will take... 
    Suggested
    Full time
    Remote work
    Flexible hours

    Deepgram

    Remote
    2 days ago
  • $100k - $150k

    Role Description We are looking for an Embedded AI Engineer to design, optimize, and deploy machine learning models that run efficiently on resource-constrained edge devices, including mobile platforms, embedded systems, and specialized accelerators. The role requires... 
    Suggested
    Full time
    Local area
    Immediate start

    Bright Vision Technologies

    Remote
    1 day ago
  • $150k - $225k

     ...our overall quality of life. As a Senior AI / Embedded Engineer, you will be responsible for the full lifecycle...  ...hardware. This includes data ingestion, model development, optimization, and deployment on embedded devices. This role is critical for building reliable... 
    Suggested
    Full time
    Work at office
    Immediate start
    Visa sponsorship
    Night shift

    E-Space

    Saratoga, CA
    14 hours ago
  •  ...Focused on optimizing Deepgram's speech models for low-power consumer hardware, the full-time Embedded AI Engineer will enhance on-device inference across various processors and accelerators while working remotely. Key responsibilities Optimize and compile speech models... 
    Suggested
    Full time
    Remote work

    Virtual Vocations Inc

    United States
    2 days ago
  • Figure is an AI robotics company developing autonomous general-purpose humanoid robots. Our goal is to build embodied AI systems...  ...that power humanoid autonomy. We are looking for a Helix AI Engineer, Modeling to design and advance the core model architectures and learning... 
    Suggested
    Full time
    Work at office

    Figure

    San Jose, CA
    14 hours ago
  •  ...Fathom to eliminate the needless overhead of meetings. Our AI assistant captures, summarizes, and organizes the key moments...  ...Sign up today (it’s free)! ROLE OVERVIEW We're hiring a Model Performance Engineer to own the speed, cost, and reliability of our model inference... 
    Full time
    Remote work

    Fathom

    Remote
    14 hours ago
  • $70k - $300k

     ...Field AI  is transforming how robots interact with the real world. We are building risk...  ...real-world results and rapidly improving models through real-field applications. Are...  ...decision-making? We are seeking a Robotics AI Engineer – Field Foundation Models & Dynamics... 
    Full time
    Remote work

    Field AI

    Remote
    14 hours ago
  • $197.3k - $225.1k

    Lead AI Engineer (Vision model customization, VML) At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer... 
    Full time
    Part time
    Local area

    Capital One Financial Corporation

    New York, NY
    14 hours ago
  •  ...At Sonatus, we’re driving the transformation to AI-enabled software-defined vehicles. Traditional automotive...  ...Sonatus is seeking a highly motivated Staff AI Engineer with expertise in data analytics and modeling to build AI-based software for next-generation AI-driven... 
    Full time
    Work at office
    Worldwide
    Flexible hours
    Shift work

    The Posted Salary Range

    Sunnyvale, CA
    14 hours ago
  • $40 per hour

     ...startup is seeking experienced cybersecurity professionals to evaluate AI-generated content and solve technical problems. This role offers...  ...to work remotely while contributing to innovative security AI models. Candidates should have 2+ years of experience in areas like... 
    Hourly pay
    Remote work

    DataAnnotation

    Bismarck, ND
    2 days ago
  • $40 per hour

    A technology firm specializing in cybersecurity is seeking experienced cybersecurity professionals to help train AI models. In this role, you will evaluate AI-generated security content and solve technical problems to strengthen AI's reasoning. Candidates should have 2+... 
    Hourly pay
    Remote work

    DataAnnotation

    Boston, MA
    3 days ago
  • $40 per hour

     ...experienced cybersecurity professionals to join our team to help train AI models. In this role, you will evaluate AI-generated security content...  ...testing, red teaming, incident response, detection engineering, DFIR, malware analysis, threat intelligence, or similar) ~ Some... 
    Hourly pay
    Full time
    Part time
    Remote work

    DataAnnotation

    Hartford, CT
    2 days ago
  • $197.3k - $225.1k

     ...Lead AI Engineer (Vision model customization, VLM) Overview At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized... 
    Full time
    Part time
    Local area

    Capital One

    New York, NY
    4 days ago
  •  ...next career opportunity now! Position Overview: The Staff AI/ML Engineer (LLMs) will lead the development of Agentic AI capabilities...  ...problems and operational workflows Adapt and fine-tune foundation models for specialized use cases Design and implement retrieval-... 
    Full time
    Temporary work
    Work at office
    Visa sponsorship
    Relocation package
    Flexible hours

    Arka Group, L.p.

    Remote
    14 hours ago
  •  ...'re ALTEN Technology USA, an engineering company helping clients bring...  ...exploration and life-saving medical devices to building autonomous...  ...taking architectural decisions on Embedded Products with evaluating...  ...Enterprise Architectural frameworks, Model-based Systems Engineering... 
    For contractors

    ALTEN Technology USA

    Bartlesville, OK
    4 days ago
  • $197.3k - $225.1k

     ...we are creating responsible and reliable AI systems, changing banking for good. For years...  ...to build world‑class applied science and engineering teams to deliver our industry leading...  ...deliver value to millions of customers. Our AI models and platforms empower teams across... 
    Local area

    Capital One

    McLean, VA
    4 days ago
  • $40 per hour

    A cybersecurity solutions company is seeking experienced professionals to help train AI models by evaluating AI-generated security content and solving technical cybersecurity problems. This position offers flexibility to choose projects and set schedules, with compensation... 
    Hourly pay
    Remote work

    DataAnnotation

    Nevada, IA
    14 hours ago
  • $40 per hour

     ...cybersecurity professionals for a remote position. In this role, you'll evaluate AI-generated cybersecurity content, solve technical problems, and provide valuable feedback to enhance AI models. Ideal candidates should have over 2 years of experience in various... 
    Hourly pay
    Remote work

    DataAnnotation

    New York, NY
    14 hours ago
  • $40 per hour

    A cybersecurity firm is seeking experienced professionals to evaluate AI-generated security content and solve technical problems. In this remote role, you will work with AI models to assess their accuracy and provide valuable feedback. Candidates should have 2+ years of... 
    Hourly pay
    Remote work

    DataAnnotation

    Wisconsin
    14 hours ago
  • $40 per hour

     ...cybersecurity firm in the United States is seeking experienced professionals to evaluate AI-generated security content and solve technical problems. You'll work directly with AI models to enhance their accuracy and improve cybersecurity tools. Ideal candidates have 2+... 
    Hourly pay
    Remote work

    DataAnnotation

    Charleston, WV
    14 hours ago
  • $241k - $326k

     ...TL, L7) with scope over a team of 15-25 engineers. Solutions will vary and can encompass GAI...  ...Feed, and Agentic solutions.  Advanced modeling skills are required, and the work will include...  ...for a team of Senior and Staff/Lead AI Engineers, reporting to a Director of AI... 
    Full time
    For contractors
    Work at office
    Flexible hours

    Linkedin

    Remote
    14 hours ago
  •  ...of a multi-modality foundation model to drive the next generation...  ...Model Optimization & Deployment Engineer, you will focus on bringing...  ...deterministic execution on edge devices. In this role, you will:...  ...maximize memory bandwidth on AI accelerators. Write production... 
    Temporary work
    Relocation package

    Zoox

    San Diego, CA
    16 days ago
  • $45 - $60 per hour

     ...workloads in a wide variety of edge and endpoint devices, ranging from battery operated smart-...  ...to work on site. Responsibilities: Model pruning: Prune the model to speed up...  ...experience working alongside industry experts in AI and semiconductor technology, with access... 
    Hourly pay
    Temporary work
    Internship
    Work at office
    Relocation

    quadric, Inc

    Burlingame, CA
    4 days ago
  • $40 per hour

     ...professionals to join their remote team. In this role, you will evaluate AI-generated security content, design solutions to cybersecurity problems, and provide essential feedback for improving AI models. Candidates should have over 2 years of hands-on experience in... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Kansas City, MO
    3 days ago
  • $190k - $260k

     ...developed an artificial intelligence (AI) powered technology stack...  ...it. Every improvement to our models – from GigaFusionNet to large-...  .... We are looking for engineers who make model training fast:...  ...models that don't fit on a single device Maximize utilization of modern... 
    Temporary work
    Work at office
    Visa sponsorship
    Flexible hours

    Kodiak

    Mountain View, CA
    16 days ago
  •  ...needed to create cutting-edge products & solutions that keep our users Ahead of Ready. Who You Are As a Software Engineer in the Radar Modeling and Simulation group, you will develop high performance C++ software. Your contributions will directly support radar modeling... 

    Lockheed Martin

    Blackwood, NJ
    14 hours ago
  •  ...the Team At OpenAI | Consumer Devices, you won’t just build products...  ...talented minds in research, engineering, and operations to create products...  ..., CA. We use a hybrid work model of four days in the office per...  ...reliable, and easy to adopt. Use AI-native tooling and automation... 
    Full time
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    14 hours ago
  • $40 per hour

     ...innovation company is seeking experienced professionals to evaluate AI-generated security content and solve technical cybersecurity problems. In this remote position, you will work with advanced AI models and contribute to improving cybersecurity tools. The ideal... 
    Remote job
    Hourly pay
    Flexible hours

    DataAnnotation

    Lansing, MI
    14 hours ago
  • A cybersecurity solutions company is looking for experienced cybersecurity professionals to help train AI models. You will work remotely to evaluate AI-generated security content, solve technical problems, and provide feedback to improve AI systems. Ideal candidates have... 
    Remote job
    Flexible hours

    DataAnnotation

    New York, NY
    14 hours ago
  • $40 per hour

     ...cybersecurity firm is looking for experienced cybersecurity professionals to evaluate AI-generated content and solve technical problems. The role involves working with advanced AI models, providing feedback, and contributing to the cybersecurity industry's future. The... 
    Remote job
    Hourly pay
    Flexible hours

    DataAnnotation

    Washington DC
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Embedded AI Engineer, On-Device Models. Be the first to apply!