Embedded ML Inference & Optimization Engineer
Applied Intuition Inc.
Applied Intuition Inc. is seeking a software engineer with deep expertise in optimizing ML models for production-grade embedded runtime environments. You will work across the ML framework stack and target a range of embedded compute platforms used in on- and off-road ADAS/AD stacks. You will collaborate with ML engineers and software developers, profiling model performance, implementing pruning/quantization, and driving efficient architectures for memory-constrained devices. #J-18808-Ljbffr Applied Intuition Inc.
$195.2k - $361.2k
...future of AI should belong to the people it servesRole SummaryMake models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments — GPU/iGPUs, Vulkan backends — not datacenter H100 environment...SuggestedFull timeInternshipLocal areaImmediate startShift work$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop...SuggestedFull time- NVIDIA is seeking a Senior DL Algorithms Engineer to optimize LLM/Omni models and enhance performance across its software stack. The ideal... ...and 3+ years of experience in deep learning, specifically in inference. This role involves profiling, analyzing bottlenecks, and...Suggested
- Intel Corporation is seeking a software engineer to make models fast on hardware people own, optimizing inference engines for edge environments. You will work with llama.cpp and vLLM, tuning KV cache, batching, and quantization while reducing CPU overhead and startup costs...SuggestedLocal area
- ...future of AI should belong to the people it serves Role Summary Make models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments - GPU/iGPUs, Vulkan backends - not datacenter H100...SuggestedLocal areaShift work
- ...industry-leading training and inference speeds and empowers machine... ...effortlessly run large-scale ML applications, without the hassle... ...‑LLM), GPU kernel‑level optimization toolchains (CUDA, Triton), and... ...Collaborate with Product and Engineering to identify where competitors...Contract workFor contractorsFor subcontractorShift work
$184k - $287.5k
...highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale... ...high-performance inference stacks, optimize GPU kernels and compilers, drive industry... ...pareto frontier for the field of ML Systems; survey recent publications...Full time- ...and data centers, to PCs, gaming and embedded systems. Grounded in a culture of... ...ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis... ...framework performance: Profile and optimize inference engines including vLLM, SGLang...
$120k - $275k
...and software to train and run the largest ML workloads for AGI. MatX is seeking silicon micro-architects and design engineers to join our team as we create best-in-class... ...dynamic and static power reduction.Drive power optimization across compute, memory, interconnect, PCIe,...Full timeWork experience placementLocal areaRemote workMonday to FridayFlexible hours$181.1k - $318.4k
Embedded Real Time Critical Control Firmware Engineer Apple is where individual imaginations gather together, committing... ...rigid real time deadlines. Use AI/ML as a tool for improved... ...control firmware which is highly optimized for cycles and memory. Deep understanding...Relocation$207k - $300k
...hardware, software, and system engineering teams to drive pre-silicon... ...Software (SW) co-development, optimize integration, and align cross-... ...years of experience working with embedded operating systems.3 years of... ...compute servers and ML headnodes. You will also contribute...Worldwide- Tilde Research is seeking a Kernel Engineer in Palo Alto to design, implement, and optimize high-performance GPU kernels for training and inference workloads. You will collaborate with ML researchers and engineers to push hardware limits, improve throughput, and reduce...
$130k - $200k
...high-level programming languages and AI/ML frameworks. This level of efficiency... ...About the Role We are seeking Software Optimization Engineers to join our growing team. Efficient’s Optimization... ...closely with Efficient’s compiler and embedded teams contributing feedback on compiler...Immediate start$170.6k - $261.3k
...Senior Machine Learning Engineer on the State Estimation... ...develop and improve the ML perception model that powers... ...efficient training and inference pipelines, including model optimization techniques (e.g.,... ...deploying ML models on embedded or resource‑constrained...Full timeLocal areaRemote workWork from homeRelocation packageFlexible hours- ...inventory management will optimize the global physical... ...Learning Software Engineers to build compute-constrained... ...teams to deploy CV/ML models into production... ...networks (both training and inference) ~ Knowledge of model... ...techniques for embedded systems, including knowledge...Full timeVisa sponsorship
- ...Volcano Ark MaaS platform invites talented engineers to advance large-model inference systems across China and internal products. You will optimize inference performance, reduce costs, and... ...Join a team that blends cutting-edge ML with practical systems, working on disaggregated...
- dorle-controls-llc is seeking experienced Embedded Software Engineers in Sunnyvale, California. Your role will focus on optimizing QNX OS performance for production automotive systems while enhancing features. Proficiency in QNX internals and significant experience with...
$197.9k - $270k
...continued success depends on experienced embedded software engineers like you joining the Roku OS Streaming... ...to all our users. This includes optimizing network interactions between our players... ...source developmentA familiarity with AI/ML and LLM technologies.Experience with...Work at officeLocal areaRemote workMonday to ThursdayFlexible hours$218.8k - $335.3k
...cutting‑edge robotics, optimization, and machine learning... ...for a Staff Software Engineer to provide technical leadership... ...analytical models and ML‑based forecasting,... .../accelerator‑based ML inference, model deployment, and... ...with real‑time or embedded platforms and constraints...Full timeLocal areaRemote workWork from homeFlexible hours$152k - $241.5k
...TensorRT team as a Senior Software Engineer, and be at the forefront of... ...high-performance AI inference solutions for automotive safety... ...to performance optimization and benchmarking efforts for... ...environmentBackground with systems programming, embedded systems, and/or compiler...Full time$184k - $287.5k
...We are seeking an AI Compiler Engineer with deep expertise in... ...implement end-to-end compiler optimization workflows, from feature engineering... ...software engineering and AI/ML experience, preferably in tools... ...solutions in production and embedded environments.Strong...Full timeRemote work$184k - $287.5k
...extraordinary Senior Perception Engineer to develop and... ...KPI building and optimization. This includes careful... ...capable of finding detailed ML bugs and iterating... ...DNN-based solutions to embedded platforms for real time... ...part of training or inference pipelines.Your base salary...Full timeWork experience placement$184k - $287.5k
...now looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior engineers who... ...are mindful of performance analysis and optimization to help us squeeze every last clock... ...Implement language and multimodal model inference as part of NVIDIA Inference Microservices...Full time$152k - $241.5k
...NVIDIA is seeking top-tier AI Compiler Engineers to drive innovation within our world-class... ...generation and computational graph optimizations for next-generation NVIDIA GPUs.Advance... ...compilation problems for AI workloads (both inference and training) and successfully...Full time$184.7k - $324.8k
...Senior Software Engineer, Middleware - Special Projects... ...and distributed inference platforms to large‑scale... ...with hardware, robotics, ML, design, and platform... ...Collaborate with hardware, embedded systems, and AI... ...Commitment to measuring and optimizing performance; skilled...Relocation- ...Apple seeks a senior engineer to design the core auction system powering its ads platform at global scale. You will drive optimal outcomes for advertisers, users, and the platform, applying advanced auction theory and machine learning in production environments. You...
- ...this data with language and embedding models, we power critical... ...years of experience in data engineering, including building and maintaining... ..., and designing schemas optimized for analytics and reporting... ...and image) and integrate ML model inference - including LLMs and...
$218.8k - $335.3k
...Computer Science, Robotics, Electrical Engineering, Mechanical Engineering, or a... ...with GPU or accelerator-based ML inference, model deployment, and optimization tools such as TensorRT or ONNX Runtime... ...familiarity with real-time or embedded platforms. Responsibilities:...Full time$240k
...industry-leading training and inference speeds; over 10 times faster... ...performance test strategies for AI/ML systems using Python, REST... ...: analyze results to ensure optimal performance of distributed AI... ...logs, monitoring tools, and engineering best practices.Document test...$224k - $356.5k
...Senior / Principal Deep Learning Engineer — Model Evaluation & AI... ...Work alongside model training, inference, and product divisions to provide... ...that inform release and optimization decisions.What we need to see... ...evaluation frameworks, benchmarks, or ML infrastructure used by other...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Embedded ML Inference & Optimization Engineer. Be the first to apply!
- embedded engineer Sunnyvale, CA
- embedded software engineer Sunnyvale, CA
- embedded developer Sunnyvale, CA
- embedded systems software engineer Sunnyvale, CA
- ai ml engineer Sunnyvale, CA
- senior ml engineer Sunnyvale, CA
- computer vision machine learning engineer Sunnyvale, CA
- machine learning engineer Sunnyvale, CA
- machine learning ai engineer Sunnyvale, CA
- machine learning software engineer Sunnyvale, CA



