Embedded AI Engineer, On-Device Models
Deepgram
Company Overview
Deepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for speech-to-text (STT), text-to-speech (TTS), and building production-grade voice agents at scale. More than 200,000 developers and 1,300+ organizations build voice offerings that are ‘Powered by Deepgram’, including Twilio, Cloudflare, Sierra, Decagon, Vapi, Daily, Cresta, Granola, and Jack in the Box. Deepgram’s voice-native foundation models are accessed through cloud APIs or as self-hosted and on-premises software, with unmatched accuracy, low latency, and cost efficiency. Backed by a recent Series C led by leading global investors and strategic partners, Deepgram has processed over 50,000 years of audio and transcribed more than 1 trillion words. There is no organization in the world that understands voice better than Deepgram.
Company Operating Rhythm
At Deepgram, we expect an AI-first mindset—AI use and comfort aren’t optional, they’re core to how we operate, innovate, and measure performance.
Every team member who works at Deepgram is expected to actively use and experiment with advanced AI tools, and even build your own into your everyday work. We measure how effectively AI is applied to deliver results, and consistent, creative use of the latest AI capabilities is key to success here. Candidates should be comfortable adopting new models and modes quickly, integrating AI into their workflows, and continuously pushing the boundaries of what these technologies can do.
Additionally, we move at the pace of AI. Change is rapid, and you can expect your day-to-day work to evolve just as quickly. This may not be the right role if you’re not excited to experiment, adapt, think on your feet, and learn constantly, or if you’re seeking something highly prescriptive with a traditional 9-to-5.
About the Role
Deepgram's speech AI models are among the fastest and most accurate in the world — and the next wave of voice experiences won't live only in the cloud. They'll run directly on the small, low-power devices people carry, wear, and keep around their homes: phones, earbuds, wearables, appliances, cameras, and purpose-built consumer hardware. Putting state-of-the-art speech models on devices with tight memory, compute, thermal, and battery budgets is a fundamentally different engineering problem, and it's one of the most important frontiers for bringing voice AI to everyone.
As an Embedded AI Engineer , you will take Deepgram's models and make them run — fast, accurately, and efficiently — on resource-constrained embedded and edge platforms. You'll work across the stack: optimizing and compiling models for on-device inference, writing performance-critical runtime code, and squeezing every last millisecond and milliwatt out of a wide range of mobile application processors, embedded SoCs, microcontrollers, and dedicated AI accelerators. Your work directly enables a new class of private, offline-capable, real-time voice experiences on the devices closest to the user.
This role is a great fit whether you're a hands-on senior embedded engineer who wants to go deep on a hard problem, or a staff-level technical leader who wants to define how Deepgram's voice AI gets onto consumer hardware and raise the bar for the engineers around you. We'll set the level to your experience.
What You'll Do
Take Deepgram's Speech and Conversational models and get them running on embedded and low-power consumer hardware — defining the architecture for on-device, real-time inference across a diverse range of processors and accelerators.
Optimize models for constrained targets through quantization, pruning, distillation, operator fusion, and architecture-specific compilation to meet strict latency, memory, power, and thermal budgets.
Write and optimize performance-critical runtime code (C, C++, and/or Rust) for embedded environments, including bare-metal and real-time operating systems such as FreeRTOS and Zephyr.
Integrate with industry-standard edge inference runtimes and vendor NPU/DSP toolchains, mapping model graphs efficiently onto on-device accelerators and CPU/GPU/NPU heterogeneity.
Build the on-device runtime plumbing: model packaging, deployment pipelines, over-the-air update mechanisms, and lightweight telemetry for devices operating with limited or intermittent connectivity.
Establish repeatable benchmarking and validation across target hardware — measuring latency, accuracy, power consumption, memory footprint, and resource utilization — and catch regressions before they ship.
Partner with silicon and device vendors on SDK integration and performance tuning, getting our models to run efficiently on new chipsets and reference platforms.
Collaborate with Research and Engine teams to influence model architectures toward edge-friendly designs from the start, reducing the optimization burden at deployment time.
You'll Love This Role If You
Find deep satisfaction in making a large model run on a tiny device — and still hit accuracy and latency targets.
Want to work at the intersection of AI and hardware, where optimization isn't optional but existential.
Are energized by the back-and-forth of getting a model to sing on a new chipset, runtime, or accelerator.
Believe on-device AI is the next major deployment frontier and want to define how speech AI gets there for consumers.
Prefer hard, constrained, ship-it problems over open-ended research — you want to see your work running in people's hands.
Care about the details that don't show up in a cloud benchmark: cold-start time, power draw, thermals, and memory fragmentation.
It's Important To Us That You Have
Experience delivering production systems on resource-constrained hardware — embedded systems, mobile, edge AI, or small low-power devices.
Strong proficiency in C, C++, and/or Rust, with experience writing performance-critical code for constrained environments.
Hands-on experience with model optimization for on-device deployment, including quantization, pruning, knowledge distillation, or architecture-specific compilation.
Familiarity with edge inference runtimes (e.g., ONNX Runtime, TensorRT, TFLite, ExecuTorch) and/or vendor-specific NPU/DSP toolchains.
A strong understanding of hardware-software interaction — CPU/GPU/NPU/DSP architectures, memory hierarchies, fixed-point/integer arithmetic, and power management — and how they affect inference performance.
Experience working close to the metal: bare-metal or RTOS environments (e.g., FreeRTOS, Zephyr), embedded Linux, or microcontroller and edge SoC development.
Strong communication skills and a builder mindset — you can scope an ambiguous optimization problem, drive it to a measurable result, and explain the tradeoffs clearly.
It Would Be Great if You Had
Experience with real-time audio processing on embedded platforms — DSP pipelines, audio codec optimization, wake-word or always-on listening, or streaming inference on microcontrollers and edge SoCs.
Depth in ML optimization techniques — custom quantization schemes, mixed-precision inference, or neural architecture search for edge targets.
Background in hardware evaluation and benchmarking — systematically comparing accelerators, SoCs, or GPUs for specific workload profiles.
Experience shipping AI features in consumer products at scale, and the instinct for what "production quality" means on a battery-powered device.
Familiarity with model compilation and optimization toolchains and their tradeoffs across hardware targets.
Experience with secure, robust on-device deployment practices — code signing, encrypted model storage, and safe update mechanisms.
$150k - $250k
...powering the future of physical AI. Founded in 2017 and now... ...We are building on-device intelligence for a next-generation... ...end-to-end lifecycle of embedded ML systems, ensuring models behave predictably and... ...Computer Science, Electrical Engineering, or a related technical...SuggestedFull timeFor contractorsFor subcontractorCasual workWork at officeLocal areaRemote workDay shift$150k - $225k
...our overall quality of life. As a Senior AI / Embedded Engineer, you will be responsible for the full lifecycle... ...hardware. This includes data ingestion, model development, optimization, and deployment on embedded devices. This role is critical for building reliable...SuggestedFull timeWork at officeImmediate startVisa sponsorshipNight shift$138k - $197k
...the software infrastructure and architecture for on-device AI/ML accelerators.Implement, model, analyze, and rigorously test C-models for on-device... ...Minimum qualifications:Bachelor's degree in Electrical Engineering, Computer Engineering, Computer Science, a related field...SuggestedWorldwide$100k - $150k
...Embedded AI Engineer - Remote Bright Vision Technologies is a technology consulting and software development... ...design, optimize, and deploy machine learning models that run efficiently on resource-constrained edge devices, including mobile platforms, embedded systems,...SuggestedFull timeH1bLocal areaImmediate startRemote workVisa sponsorship- ...'re ALTEN Technology USA, an engineering company helping clients bring... ...exploration and life-saving medical devices to building autonomous... ...taking architectural decisions on Embedded Products with evaluating... ...Enterprise Architectural frameworks, Model-based Systems Engineering...SuggestedFor contractors
- ...of a multi-modality foundation model to drive the next generation... ...Model Optimization & Deployment Engineer, you will focus on bringing... ...deterministic execution on edge devices. In this role, you will:... ...maximize memory bandwidth on AI accelerators. Write production...Temporary workRelocation package
$197.8k - $271.98k
About Analog DevicesAnalog Devices, Inc. (NASDAQ: ADI) is a... ...analog, digital, AI, and software technologies... ...are seeking a Staff AI Engineer in Time-Series & Sensor Foundation Models to advance AI engineering... ...research in time series embedding and compression to enable...Permanent employmentFull timeWork at officeDay shift$190k - $260k
...developed an artificial intelligence (AI) powered technology stack... ...it. Every improvement to our models – from GigaFusionNet to large-... .... We are looking for engineers who make model training fast:... ...models that don't fit on a single device Maximize utilization of modern...Temporary workWork at officeVisa sponsorshipFlexible hours- ...Figure is an AI robotics company developing autonomous general-purpose humanoid robots. Our goal is to build embodied AI systems... ...that power humanoid autonomy. We are looking for a Helix AI Engineer, Modeling to design and advance the core model architectures and learning...Full timeWork at office
- ...Position Summary - Engineer III (R & D Engineering) The Engineer III leads the design,... ...sustaining engineering of vascular access devices and their associated manufacturing processes... ...completeness and accuracy. • Design, model (SolidWorks), and fabricate test and manufacturing...Full timeFor contractors
- ...patients worldwide.We’re a team of engineers, clinicians, and innovators... ...platforms. As a Senior AI/ML Research Engineer, you will... ...and fine-tune the foundation models—VFMs, VLMs, and VLA models—that... ...timeFunction: EngineeringExperience level: AssociateIndustry: Medical DeviceLocal areaWorldwideFlexible hours
- ...From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at... ...career.THE ROLE: Join our innovative team at AMD as an AI Models Software Engineer. We are seeking passionate individuals who are dedicated to...
- ...Fathom to eliminate the needless overhead of meetings. Our AI assistant captures, summarizes, and organizes the key moments... ...Sign up today (it’s free)! ROLE OVERVIEW We're hiring a Model Performance Engineer to own the speed, cost, and reliability of our model inference...Full timeRemote work
- ...Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This... ...computation. Cerebras works with the leading model labs, global enterprises, and cutting-... ...do this on a loop." You'll sit between engineering, product, and customer-facing teams....Full time
- ...Job Title: Embedded Software Engineer (AFC Device) About the Role STraffic America, headquartered in Vienna, Virginia, is a leading provider of Automatic Fare Collection (AFC) systems for major U.S. transit agencies. As the U.S. subsidiary of STraffic — a global...Full timeRemote work
$149 per hour
...day. A member of our recruitment team will provide more details.Vice President - Technical AI Foundation Model EngineerRole SummaryThe VP, Technical AI Foundation Model Engineer is responsible for designing, building, deploying, and optimizing enterprise-grade AI solutions...Full timeWork at officeLocal areaRemote work1 day per week$125k - $150k
Job DescriptionEverforth ECS is seeking an AI Model Engineer to work in a hybrid remote/onsite capacity, with minimum of 3 business days onsite at our Fairfax, VA corporate office and/or our Ashburn, VA customer site. Please note: This position is contingent upon contract...Contract workWork at officeRemote work$70k - $300k
...Field AI is transforming how robots interact with the real world. We are building risk... ...real-world results and rapidly improving models through real-field applications. Are... ...decision-making? We are seeking a Robotics AI Engineer – Field Foundation Models & Dynamics...Full timeRemote work$188k - $300k
...At Sonatus, we’re driving the transformation to AI-enabled software-defined vehicles. Traditional automotive... ...Sonatus is seeking a highly motivated Staff AI Engineer with expertise in data analytics and modeling to build AI-based software for next-generation AI-driven...Full timeWork at officeWorldwideFlexible hoursShift work- ...kWe are partnered with a leading medical device company that develops innovative... ...They are looking to hire a Senior Firmware Engineer to join their Spinal team, developing next... ...production release. The ideal candidate has deep embedded firmware expertise, enjoys making...
$110k - $220k
...OfficeThe Opportunity: Walmart’s Supply Chain AI Lab & Innovation Factory is building a... ...improve through rigorous evaluation and model post-training. This role exists to build... ...role. It is a deeply hands-on AI systems engineering position focused on designing, building,...Full timeContract workTemporary workPart time$139.1k - $231.9k
...innovative and technically accomplished AI/ML Engineering leader to accelerate the transformation... ...data engineering, and vaccine science. Embedded within Vaccines Research and supporting... ...enables advanced analytics, predictive modeling, generative AI applications, and...Permanent employmentFull timeH1bLocal areaVisa sponsorshipWork visaRelocation package2 days per week$108.4k - $227.5k
Job Title: Staff AI/ML Engineer (Large Language Model) Job Category: Science Time Type: Full time Minimum Clearance Required to Start: TS/SCI Employee Type: Regular Percentage of Travel Required: Up to 10% Type of Travel: Local * * * The Opportunity...Full timeContract workWork experience placementLocal areaFlexible hours- ...career opportunity now! Position Overview: The Principal AI/ML Engineer will support the development of AI/ML algorithms in a multitude... ...processing, reinforcement learning, and large language models. We offer generous relocation benefits for eligible candidates...Full timeTemporary workWork at officeLocal areaRemote workVisa sponsorshipRelocation packageFlexible hours
- ...of supercomputing, high-performance computing, cloud, and AI. Whether you’re designing next-gen processors, enabling... ...AMD AI Group is looking for a Senior Software Development Engineer to own the end-to-end model execution stack on AMD Instinct GPUs - spanning training...
- ...Embedded AI Engineer Remote | Contractor | Half-time to Full-time About SPACE44 AI is changing how organisations operate, but most... ...Business-fluent English (written and verbal). Working Model & Benefits Fully Remote Environment: Work from anywhere...Full timeFor contractorsRemote workFlexible hours
- ...Description Job Description Palona’s AI agents operate in real restaurant environments... ...requires more than selecting the newest model. It requires disciplined evaluation, high-... ...We are looking for an applied AI Modeling Engineer to improve the intelligence, accuracy,...Temporary workImmediate start
$197.3k - $225.1k
...Lead AI Engineer (Vision model customization, VLM) Overview At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer...Full timePart timeLocal area- ...everything we do, we're shaping the future of automation in dynamic markets.As an AI Embedded Engineer IV, you own the layer where artificial intelligence (AI) meets the machine. Models that run fine on a workstation have to run on an NVIDIA Jetson bolted to a vehicle,...Work at officeLocal area
- ...seeking a highly motivated and technically skilled Edge AI/Model Optimization Engineer to support the deployment, optimization, and sustainment... ...benchmarking, and operationalizing Large Language Models (LLMs), embedding models, and AI inference services for constrained...Local area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Embedded AI Engineer, On-Device Models. Be the first to apply!
- c++ embedded engineer United States
- embedded developer United States
- embedded linux engineer United States
- embedded electrical engineer United States
- embedded software engineer remote United States
- embedded firmware developer United States
- embedded software engineer United States
- embedded engineer United States
- graduate embedded software engineer United States
- embedded systems software engineer United States



