Embedded ML Inference Optimization Engineer
Decisive Point
Decisive Point is seeking a Software Engineer in Sunnyvale, California, with expertise in optimizing machine learning models for embedded systems. This role involves performance optimization for embedded compute platforms, collaborating with ML engineers, and requires strong software development skills. The ideal candidate has a Bachelor’s degree and at least 3 years of experience in ML accelerators and deep learning frameworks such as PyTorch and JAX. The position offers a competitive salary and comprehensive benefits. #J-18808-Ljbffr Decisive Point
$195.2k - $361.2k
...future of AI should belong to the people it servesRole SummaryMake models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments — GPU/iGPUs, Vulkan backends — not datacenter H100 environment...SuggestedFull timeInternshipLocal areaImmediate startShift work$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop...SuggestedFull time- ...industry‑leading training and inference speeds and empowers machine... ...effortlessly run large‑scale ML applications, without the hassle... ...‑LLM), GPU kernel‑level optimization toolchains (CUDA, Triton), and... ...Collaborate with Product and Engineering to identify where competitors...SuggestedContract workShift work
- Intel Corporation is seeking a software engineer to make models fast on hardware people own, optimizing inference engines for edge environments. You will work with llama.cpp and vLLM, tuning KV cache, batching, and quantization while reducing CPU overhead and startup costs...SuggestedLocal area
- NVIDIA is seeking a Senior DL Algorithms Engineer to optimize LLM/Omni models and enhance performance across its software stack. The ideal... ...and 3+ years of experience in deep learning, specifically in inference. This role involves profiling, analyzing bottlenecks, and...Suggested
- ...future of AI should belong to the people it serves Role Summary Make models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments - GPU/iGPUs, Vulkan backends - not datacenter H100...Local areaShift work
$184k - $287.5k
...highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale... ...high-performance inference stacks, optimize GPU kernels and compilers, drive industry... ...pareto frontier for the field of ML Systems; survey recent publications...Full time- ...and data centers, to PCs, gaming and embedded systems. Grounded in a culture of... ...ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis... ...framework performance: Profile and optimize inference engines including vLLM, SGLang...
$120k - $275k
...and software to train and run the largest ML workloads for AGI. MatX is seeking silicon micro-architects and design engineers to join our team as we create best-in-class... ...dynamic and static power reduction.Drive power optimization across compute, memory, interconnect, PCIe,...Full timeWork experience placementLocal areaRemote workMonday to FridayFlexible hours$130k - $200k
...high-level programming languages and AI/ML frameworks. This level of efficiency... ...About the Role We are seeking Software Optimization Engineers to join our growing team. Efficient’s Optimization... ...closely with Efficient’s compiler and embedded teams contributing feedback on compiler...Immediate start$181.1k - $318.4k
Embedded Real Time Critical Control Firmware Engineer Apple is where individual imaginations gather together, committing... ...rigid real time deadlines. Use AI/ML as a tool for improved... ...control firmware which is highly optimized for cycles and memory. Deep understanding...Relocation$207k - $300k
...hardware, software, and system engineering teams to drive pre-silicon... ...Software (SW) co-development, optimize integration, and align cross-... ...years of experience working with embedded operating systems.3 years of... ...compute servers and ML headnodes. You will also contribute...Worldwide- Tilde Research is seeking a Kernel Engineer in Palo Alto to design, implement, and optimize high-performance GPU kernels for training and inference workloads. You will collaborate with ML researchers and engineers to push hardware limits, improve throughput, and reduce...
$170.6k - $261.3k
...Senior Machine Learning Engineer on the State Estimation... ...develop and improve the ML perception model that powers... ...efficient training and inference pipelines, including model optimization techniques (e.g.,... ...deploying ML models on embedded or resource‑constrained...Full timeLocal areaRemote workWork from homeRelocation packageFlexible hours- ...inventory management will optimize the global physical... ...Learning Software Engineers to build compute-constrained... ...teams to deploy CV/ML models into production... ...networks (both training and inference) ~ Knowledge of model... ...techniques for embedded systems, including knowledge...Full timeVisa sponsorship
- dorle-controls-llc is seeking experienced Embedded Software Engineers in Sunnyvale, California. Your role will focus on optimizing QNX OS performance for production automotive systems while enhancing features. Proficiency in QNX internals and significant experience with...
$152k - $241.5k
...TensorRT team as a Senior Software Engineer, and be at the forefront of... ...high-performance AI inference solutions for automotive safety... ...to performance optimization and benchmarking efforts for... ...environmentBackground with systems programming, embedded systems, and/or compiler...Full time$218.8k - $335.3k
...cutting‑edge robotics, optimization, and machine learning... ...for a Staff Software Engineer to provide technical leadership... ...analytical models and ML‑based forecasting,... .../accelerator‑based ML inference, model deployment, and... ...with real‑time or embedded platforms and constraints...Full timeLocal areaRemote workWork from homeFlexible hours$184k - $287.5k
...now looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior engineers who... ...are mindful of performance analysis and optimization to help us squeeze every last clock... ...Implement language and multimodal model inference as part of NVIDIA Inference Microservices...Full time$184k - $287.5k
...We are seeking an AI Compiler Engineer with deep expertise in... ...implement end-to-end compiler optimization workflows, from feature engineering... ...software engineering and AI/ML experience, preferably in tools... ...solutions in production and embedded environments.Strong...Full timeRemote work$184k - $287.5k
...extraordinary Senior Perception Engineer to develop and... ...KPI building and optimization. This includes careful... ...capable of finding detailed ML bugs and iterating... ...DNN-based solutions to embedded platforms for real time... ...part of training or inference pipelines.Your base salary...Full timeWork experience placement$193.3k - $261.5k
We are looking for a Senior Inference Engineer to own inference for real-time multimodalconversational AI. This is a full-stack inference role... ...choices before they arelocked in• Implement and optimize the inference path for large-scale multimodal models — attentionand...InternshipLocal areaFlexible hours$152k - $241.5k
...NVIDIA is seeking top-tier AI Compiler Engineers to drive innovation within our world-class... ...generation and computational graph optimizations for next-generation NVIDIA GPUs.Advance... ...compilation problems for AI workloads (both inference and training) and successfully...Full time- ...integration of hardware and software performance engineering. We seek a senior engineer who will shape core... ...elevating the team as they implement high-throughput inference systems. You will own the inference engine, optimize runtimes, and push the boundaries of latency and...
- ...Intel is seeking a seasoned software engineer to accelerate AI inference on edge hardware. You will optimize llama.cpp/vLLM, tune KV cache, batching and scheduling, and push quantization strategies to balance speed and quality. This role focuses on low-latency, privacy...
- ...this data with language and embedding models, we power critical... ...years of experience in data engineering, including building and maintaining... ..., and designing schemas optimized for analytics and reporting... ...and image) and integrate ML model inference - including LLMs and...
$184.7k - $324.8k
Senior Software Engineer, Middleware - Special Projects... ...and distributed inference platforms to large‑scale... ...with hardware, robotics, ML, design, and platform... ...Collaborate with hardware, embedded systems, and AI... ...Commitment to measuring and optimizing performance; skilled...Relocation- Apple seeks a senior engineer to design the core auction system powering its ads platform at global scale. You will drive optimal outcomes for advertisers, users, and the platform, applying advanced auction theory and machine learning in production environments. You will...
$224k - $356.5k
...Senior / Principal Deep Learning Engineer — Model Evaluation & AI... ...Work alongside model training, inference, and product divisions to provide... ...that inform release and optimization decisions.What we need to see... ...evaluation frameworks, benchmarks, or ML infrastructure used by other...Full time$240k
...industry-leading training and inference speeds; over 10 times faster... ...performance test strategies for AI/ML systems using Python, REST... ...: analyze results to ensure optimal performance of distributed AI... ...logs, monitoring tools, and engineering best practices.Document test...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Embedded ML Inference Optimization Engineer. Be the first to apply!
- embedded engineer Sunnyvale, CA
- embedded software engineer Sunnyvale, CA
- embedded developer Sunnyvale, CA
- embedded systems software engineer Sunnyvale, CA
- ai ml engineer Sunnyvale, CA
- senior ml engineer Sunnyvale, CA
- computer vision machine learning engineer Sunnyvale, CA
- machine learning engineer Sunnyvale, CA
- machine learning ai engineer Sunnyvale, CA
- machine learning software engineer Sunnyvale, CA


