Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

On-Device AI Inference Engineer

$200k
Full-time

Hark

About Hark

Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.

We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines. While today's AI largely operates through chat boxes and decade-old devices, Hark is focused on what comes next: agentic systems that interact naturally with people and the real world.

To get there, we're developing multimodal models and next-generation AI hardware together - designed from the ground up as a single, unified interface for a new era of intelligent systems.

About the Role

You'll make Hark's models run fast on the hardware we ship. That means writing the kernels, building the runtime paths, and profiling transformer workloads on DSPs, NPUs, and other constrained targets until they hit the latency and power budgets our devices are built around. This is hands-on systems work close to the metal, on a small team where the code you write is what users feel as response time.

Responsibilities

  • Write and optimize the low-level kernels and runtime paths that transformer workloads execute through on target silicon.
  • Decide how multiple models share limited memory and power i.e. residency, scheduling, and swap behavior across concurrent workloads.
  • Profile models on real hardware, find the bottlenecks, and close the gap between theoretical and delivered performance.
  • Take models from full precision to INT8/INT4 and get them running within per-product size, latency, and power budgets.
  • Get transformer workloads executing efficiently on new accelerators as they come online, working alongside the hardware team.
  • Feed real deployment constraints back to the model teams so architecture decisions account for what the hardware can actually do.

Requirements

  • 4-8+ years writing performance-critical software, with hands-on optimization on GPUs, NPUs, DSPs, or similar accelerators.
  • Strong C/C++ and comfort with SIMD, custom kernels, memory layout, and the profiling tools that go with them.
  • You understand attention, KV-cache behavior, and where transformer inference actually spends its time and memory bandwidth.
  • You reason in compute, memory, and power budgets, and you've optimized against them rather than around them.
  • You've had a model you optimized run in a product on constrained hardware.

Bonus Qualifications

  • Experience with Hexagon DSP, Ambiq-class MCUs, or comparable embedded AI silicon.
  • Familiarity with ONNX Runtime, TVM, MLIR, TensorRT, or similar inference and compiler toolchains.
  • Background in speech or audio inference, where latency is perceptible to the user.

Compensation

The US base salary range for this full-time position is between $200,000 - $450,000 annually.

The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components/benefits depending on the specific role. This information will be shared if an employment offer is extended.

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the On-Device AI Inference Engineer in San Jose, CA vacancy
  •  ...community where each individual can thrive.Job DescriptionThe AI Inference Engineer plays a critical role in the AI lifecycle by bridging the...  ...environments—from GPU-rich data centers to resource-constrained edge devices—with a strong emphasis on maximizing throughput, minimizing... 
    Suggested
    Full time
    Local area
    Immediate start

    F5 Networks

    San Jose, CA
    1 day ago
  • $151.8k - $332.2k

    What you can expect We are looking for an AI Inference Engineer with a solid background in speech recognition and model inference. In this role, you will develop state-of-the-art automatic speech recognition system and ship it to various Zoom products. You will work on... 
    Suggested
    Full time
    Work at office
    Remote work

    Zoom

    San Jose, CA
    4 days ago
  • $300k

     ...humans and machines. While today's AI largely operates through chat boxes and decade-old devices, Hark is focused on what comes...  ..., and building the low-level inference stack that turns a trained...  .... Hire and lead a team of engineers on performance-critical software... 
    Suggested
    Full time

    Hark

    San Jose, CA
    3 days ago
  • $229.9k - $286.2k

    AI Engineer 5 (FM Hosting, LLM Inference) At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences... 
    Suggested
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    5 days ago
  • $229.9k - $262.4k

     ...Sr. Lead AI Engineer (FM Hosting, LLM Inference) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized... 
    Suggested
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    23 hours ago
  •  ...Embodied AI Engineer UnitX builds the world's leading physical AI systems to automate repetitive...  ...these models directly to edge compute devices on our global factory network,...  ...and optimizing ML models for real-time inference on robotic hardware (e.g., NVIDIA Jetson... 

    UnitX

    Milpitas, CA
    4 days ago
  • $274k - $300k

     ...Job Description Job Description Saviynt's AI-powered identity platform manages and governs human and non-human access...  .... For more information, please visit  AI Platform Engineer – Training & Inference Saviynt's AI-powered identity platform manages and governs... 

    Saviynt

    Milpitas, CA
    12 days ago
  • $136.3k - $231.7k

     ...ecosystem. Virtually every electronic device in the world is produced using...  ...expert teams of physicists, engineers, data scientists and problem-...  ...is seeking a motivated AI Engineer with a growth mindset...  ...accuracy; optimize models for inference throughput, including GPU-accelerated... 
    Minimum wage
    Full time
    Work experience placement
    Flexible hours

    KLA-Tencor

    Milpitas, CA
    23 hours ago
  • $152k - $241.5k

     ...into the unlimited potential of AI to define the next era of computing...  ...an AI & Deep Learning Compiler Engineer. NVIDIA is hiring software...  ...has been the backbone of NVIDIA’s inference engine, spanning across data centers, personal devices, automotive, and robotics. The compiler... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    23 hours ago
  • $229.9k - $262.4k

     ...Overview AI Engineer 5 (FM Hosting, LLM Inference) At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer... 
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    5 days ago
  • $117.7k - $221.4k

     ...practical, and cost efficient for embodied AI systems. We believe the next generation...  .... This operating model reflects how Cola engineers think: build durable intermediate artifacts...  ...the data processing, featurization, and inference foundations that power scalable world... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    2 days ago
  • $152k - $241.5k

     ...GPU deep learning ignited modern AI — the next era of computing —...  ...an AI & Deep Learning Compiler Engineer. NVIDIA is hiring software engineers...  ...been the backbone of NVIDIA’s inference engine, spanning across data centers, personal devices, automotive, and robotics. The compiler... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    23 hours ago
  •  ...next-generation computing experiences—from AI and data centers, to PCs, gaming and...  ...advance your career. THE ROLEWe are hiring AI Engineers to build recursive self-improvement...  ...JAX, TensorFlow, or distributed training/inference systems.Experience with reinforcement learning... 

    AMD

    Santa Clara, CA
    23 hours ago
  • $100k

     ...is leading the industry on cutting-edge AI technology, revolutionizing performance expectations...  ...Speed Interconnect / Signal Integrity Engineer to design and validate high-bandwidth...  ...technologies for next-generation AI inference and training clusters. This role is on-site... 
    Permanent employment

    Tenstorrent

    Santa Clara, CA
    4 days ago
  • $229.9k - $262.4k

     ...Overview Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning... 
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    16 days ago
  •  ...worry. Arlo's deep expertise in AI- and CV-powered analytics,...  ...About the role As a Staff AI Engineer for Applied AI, you'll be a technical...  ...agent runs. Keep production inference fast and economical - serving...  .... Nice to have Edge/on-device inference (quantization-aware... 
    Odd job
    Night shift

    Arlo Technologies, Inc.

    Milpitas, CA
    1 day ago
  •  ...Job Description In this AI/ML ASIC Performance Engineer position, you will develop AI Storage...  ...processing subsystems and SoC CPUs on the device side and Host CPUs. Author...  ...bandwidth Architect memory-efficient inference/training systems utilizing techniques... 
    Full time
    Temporary work
    Remote work
    Flexible hours
    Shift work
    Night shift

    Sandisk

    Milpitas, CA
    12 days ago
  •  ...next‑generation computing experiences—from AI and data centers, to PCs, gaming and...  ...ROLE We are hiring Forward Deployed AI Engineer to build the prototypes, tools, integrations...  ..., profilers, or distributed training/inference. Experience with hardware design, verification... 

    Advanced Micro Devices , Inc.

    Santa Clara, CA
    1 day ago
  • $160k - $192k

     ...AI Engineer Spectro Cloud is the Kubernetes and AI infrastructure platform behind some of the most demanding AI environments in the world. Our PaletteAI platform, AI Inference Launchpad, and Instinct Coder solutions let operators stand up production-grade AI clusters... 
    Work at office
    Immediate start
    Remote work
    Work from home
    Work visa

    Spectro Cloud

    San Jose, CA
    3 days ago
  •  ...AI Engineer Opportunity Hope you are doing well Number of Position: 2 Only W2 I Abhishek would like to share a job opportunity as AI...  ...libraries like TensorRT and CUDA . Understanding of generative AI and inference engines. Responsibilities: Preparing and fine-tuning models... 
    Work visa

    Syntricate Technologies

    Santa Clara, CA
    4 days ago
  •  ...Conduct advanced research in Generative AI, focusing on the latest advancements in LLMs. Develop and implement advanced techniques in multimodal LLMs, agentic AI, fine-tuning, distillation, inference optimization, test-time scaling, and reasoning models. Collaborate... 

    Tata Consultancy Services

    Santa Clara, CA
    4 days ago
  • $140k - $165k

     ...today's most advanced electronic devices and IT infrastructure,...  ...Join Us? Build foundational AI infrastructure that powers next...  ...We are seeking a hands-on AI Engineer to design, deploy, and maintain...  ...; optimize for edge/on-device inference. Implement Model Control Protocols... 

    SK hynix memory solutions America Inc.

    San Jose, CA
    1 day ago
  • $147k - $210k

     ...processing.Develop LiteRT, Google's on-device AI framework for first- and third-party, enabling...  ....Improve performance of on-device model inference via optimizations in on-device runtime...  ...with on-device ML.Google's software engineers develop the next-generation technologies... 

    Google

    Sunnyvale, CA
    1 day ago
  • $184k - $287.5k

    NVIDIA seeks a Senior Software Engineer specializing in Deep Learning Inference for our growing team. As a key contributor, you will help design, build, and...  ...accelerated software that powers today’s most sophisticated AI applications. Our team is responsible for developing... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $307k - $427k

     ...end-to-end technical architecture for on-device LLM integration and voice processing pipelines...  ...model performance for constrained Edge-AI environments, focusing on latency, power,...  ....Mentor senior technical leads and staff engineers, fostering a culture of technical... 

    Google

    San Jose, CA
    3 days ago
  •  ...next-generation computing experiences—from AI and data centers, to PCs, gaming and...  ...looking for a Forward Deployed Research Engineer to build, evaluate, and deploy cutting-edge...  ...ModelingModel or Agent EvaluationAI Training or Inference InfrastructureDeep understanding of... 

    AMD

    Santa Clara, CA
    1 day ago
  •  ...our overall quality of life. As an AI / Embedded Engineer, you will be responsible for the full...  ...optimization, and deployment on embedded devices. This role is critical for building...  ...distillation to reduce model size and inference latency ◦ Use frameworks including TensorFlow... 
    Full time
    Work at office
    Immediate start
    Visa sponsorship
    Night shift

    E-Space

    Saratoga, CA
    26 days ago
  • $152k - $241.5k

    We are looking for outstanding Senior High Performance AI Engineers to build the next generation of agentic AI systems for the CUDA ecosystem...  ...across NVIDIA's software and hardware stack, from models and inference through compilers, runtimes, libraries, kernels, and GPUs.... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

     ...recently, GPU deep learning ignited modern AI — the next era of computing — with the...  ...company”.NVIDIA is hiring a Senior AI Compiler Engineer. GPUs are driving rapid progress in deep...  ...-based AI compiler that powers NVIDIA’s inference engine end to end, with a focus on... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

    We are looking for a strong engineer to join the DRIVE Road Structure / Online Mapping / Context...  ...you interested in inventing human-level AI for navigation in the unconstrained world...  ...-precision techniques, and efficient GPU inference using NVIDIA software and hardware is... 
    Full time
    Remote work
    Shift work

    Nvidia

    Santa Clara, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to On-Device AI Inference Engineer. Be the first to apply!