Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Embedded ML Inference & Optimization Engineer

Applied Intuition Inc.

Applied Intuition Inc. is seeking a software engineer with deep expertise in optimizing ML models for production-grade embedded runtime environments. You will work across the ML framework stack and target a range of embedded compute platforms used in on- and off-road ADAS/AD stacks. You will collaborate with ML engineers and software developers, profiling model performance, implementing pruning/quantization, and driving efficient architectures for memory-constrained devices. #J-18808-Ljbffr Applied Intuition Inc.

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Embedded ML Inference & Optimization Engineer in Sunnyvale, CA vacancy
  • $195.2k - $361.2k

     ...future of AI should belong to the people it servesRole SummaryMake models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments — GPU/iGPUs, Vulkan backends — not datacenter H100 environment... 
    Suggested
    Full time
    Internship
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • NVIDIA is seeking a Senior DL Algorithms Engineer to optimize LLM/Omni models and enhance performance across its software stack. The ideal...  ...and 3+ years of experience in deep learning, specifically in inference. This role involves profiling, analyzing bottlenecks, and... 
    Suggested

    NVIDIA

    Santa Clara, CA
    5 days ago
  • Intel Corporation is seeking a software engineer to make models fast on hardware people own, optimizing inference engines for edge environments. You will work with llama.cpp and vLLM, tuning KV cache, batching, and quantization while reducing CPU overhead and startup costs... 
    Suggested
    Local area

    Intel Corporation

    Santa Clara, CA
    4 days ago
  •  ...future of AI should belong to the people it serves Role Summary Make models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments - GPU/iGPUs, Vulkan backends - not datacenter H100... 
    Suggested
    Local area
    Shift work

    PVH (Tommy Hilfiger/Calvin Klein)

    Santa Clara, CA
    6 days ago
  •  ...industry-leading training and inference speeds and empowers machine...  ...effortlessly run large-scale ML applications, without the hassle...  ...‑LLM), GPU kernel‑level optimization toolchains (CUDA, Triton), and...  ...Collaborate with Product and Engineering to identify where competitors... 
    Contract work
    For contractors
    For subcontractor
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    6 days ago
  • $184k - $287.5k

     ...highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale...  ...high-performance inference stacks, optimize GPU kernels and compilers, drive industry...  ...pareto frontier for the field of ML Systems; survey recent publications... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...and data centers, to PCs, gaming and embedded systems. Grounded in a culture of...  ...ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis...  ...framework performance: Profile and optimize inference engines including vLLM, SGLang... 

    AMD

    Santa Clara, CA
    2 days ago
  • $120k - $275k

     ...and software to train and run the largest ML workloads for AGI. MatX is seeking silicon micro-architects and design engineers to join our team as we create best-in-class...  ...dynamic and static power reduction.Drive power optimization across compute, memory, interconnect, PCIe,... 
    Full time
    Work experience placement
    Local area
    Remote work
    Monday to Friday
    Flexible hours

    MatX

    Mountain View, CA
    4 days ago
  • $181.1k - $318.4k

    Embedded Real Time Critical Control Firmware Engineer Apple is where individual imaginations gather together, committing...  ...rigid real time deadlines. Use AI/ML as a tool for improved...  ...control firmware which is highly optimized for cycles and memory. Deep understanding... 
    Relocation

    Apple Inc.

    Sunnyvale, CA
    5 days ago
  • $207k - $300k

     ...hardware, software, and system engineering teams to drive pre-silicon...  ...Software (SW) co-development, optimize integration, and align cross-...  ...years of experience working with embedded operating systems.3 years of...  ...compute servers and ML headnodes. You will also contribute... 
    Worldwide

    Google

    Sunnyvale, CA
    3 days ago
  • Tilde Research is seeking a Kernel Engineer in Palo Alto to design, implement, and optimize high-performance GPU kernels for training and inference workloads. You will collaborate with ML researchers and engineers to push hardware limits, improve throughput, and reduce... 

    Tilde Research

    Palo Alto, CA
    5 days ago
  • $130k - $200k

     ...high-level programming languages and AI/ML frameworks. This level of efficiency...  ...About the Role We are seeking Software Optimization Engineers to join our growing team. Efficient’s Optimization...  ...closely with Efficient’s compiler and embedded teams contributing feedback on compiler... 
    Immediate start

    Efficient Computer

    San Jose, CA
    6 days ago
  • $170.6k - $261.3k

     ...Senior Machine Learning Engineer on the State Estimation...  ...develop and improve the ML perception model that powers...  ...efficient training and inference pipelines, including model optimization techniques (e.g.,...  ...deploying ML models on embedded or resource‑constrained... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    21 hours ago
  •  ...inventory management will optimize the global physical...  ...Learning Software Engineers to build compute-constrained...  ...teams to deploy CV/ML models into production...  ...networks (both training and inference) ~ Knowledge of model...  ...techniques for embedded systems, including knowledge... 
    Full time
    Visa sponsorship

    Corvus Robotics

    Mountain View, CA
    1 day ago
  •  ...Volcano Ark MaaS platform invites talented engineers to advance large-model inference systems across China and internal products. You will optimize inference performance, reduce costs, and...  ...Join a team that blends cutting-edge ML with practical systems, working on disaggregated... 

    ByteDance

    San Jose, CA
    10 hours ago
  • dorle-controls-llc is seeking experienced Embedded Software Engineers in Sunnyvale, California. Your role will focus on optimizing QNX OS performance for production automotive systems while enhancing features. Proficiency in QNX internals and significant experience with... 

    dorle-controls-llc

    Sunnyvale, CA
    4 days ago
  • $197.9k - $270k

     ...continued success depends on experienced embedded software engineers like you joining the Roku OS Streaming...  ...to all our users. This includes optimizing network interactions between our players...  ...source developmentA familiarity with AI/ML and LLM technologies.Experience with... 
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    2 days ago
  • $218.8k - $335.3k

     ...cutting‑edge robotics, optimization, and machine learning...  ...for a Staff Software Engineer to provide technical leadership...  ...analytical models and ML‑based forecasting,...  .../accelerator‑based ML inference, model deployment, and...  ...with real‑time or embedded platforms and constraints... 
    Full time
    Local area
    Remote work
    Work from home
    Flexible hours

    General Motors

    Sunnyvale, CA
    6 hours ago
  • $152k - $241.5k

     ...TensorRT team as a Senior Software Engineer, and be at the forefront of...  ...high-performance AI inference solutions for automotive safety...  ...to performance optimization and benchmarking efforts for...  ...environmentBackground with systems programming, embedded systems, and/or compiler... 
    Full time

    Nvidia

    Santa Clara, CA
    6 hours ago
  • $184k - $287.5k

     ...We are seeking an AI Compiler Engineer with deep expertise in...  ...implement end-to-end compiler optimization workflows, from feature engineering...  ...software engineering and AI/ML experience, preferably in tools...  ...solutions in production and embedded environments.Strong... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...extraordinary Senior Perception Engineer to develop and...  ...KPI building and optimization. This includes careful...  ...capable of finding detailed ML bugs and iterating...  ...DNN-based solutions to embedded platforms for real time...  ...part of training or inference pipelines.Your base salary... 
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...now looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior engineers who...  ...are mindful of performance analysis and optimization to help us squeeze every last clock...  ...Implement language and multimodal model inference as part of NVIDIA Inference Microservices... 
    Full time

    Nvidia

    Santa Clara, CA
    6 hours ago
  • $152k - $241.5k

     ...NVIDIA is seeking top-tier AI Compiler Engineers to drive innovation within our world-class...  ...generation and computational graph optimizations for next-generation NVIDIA GPUs.Advance...  ...compilation problems for AI workloads (both inference and training) and successfully... 
    Full time

    Nvidia

    Santa Clara, CA
    6 hours ago
  • $184.7k - $324.8k

     ...Senior Software Engineer, Middleware - Special Projects...  ...and distributed inference platforms to large‑scale...  ...with hardware, robotics, ML, design, and platform...  ...Collaborate with hardware, embedded systems, and AI...  ...Commitment to measuring and optimizing performance; skilled... 
    Relocation

    Apple Inc.

    Cupertino, CA
    10 hours ago
  •  ...Apple seeks a senior engineer to design the core auction system powering its ads platform at global scale. You will drive optimal outcomes for advertisers, users, and the platform, applying advanced auction theory and machine learning in production environments. You... 

    Socket.dev

    Cupertino, CA
    10 hours ago
  •  ...this data with language and embedding models, we power critical...  ...years of experience in data engineering, including building and maintaining...  ..., and designing schemas optimized for analytics and reporting...  ...and image) and integrate ML model inference - including LLMs and... 

    Socket.dev

    Cupertino, CA
    3 days ago
  • $218.8k - $335.3k

     ...Computer Science, Robotics, Electrical Engineering, Mechanical Engineering, or a...  ...with GPU or accelerator-based ML inference, model deployment, and optimization tools such as TensorRT or ONNX Runtime...  ...familiarity with real-time or embedded platforms. Responsibilities:... 
    Full time

    General Motors

    Sunnyvale, CA
    8 days ago
  • $240k

     ...industry-leading training and inference speeds; over 10 times faster...  ...performance test strategies for AI/ML systems using Python, REST...  ...: analyze results to ensure optimal performance of distributed AI...  ...logs, monitoring tools, and engineering best practices.Document test... 

    Cerebras Systems

    Sunnyvale, CA
    3 days ago
  • $224k - $356.5k

     ...Senior / Principal Deep Learning Engineer — Model Evaluation & AI...  ...Work alongside model training, inference, and product divisions to provide...  ...that inform release and optimization decisions.What we need to see...  ...evaluation frameworks, benchmarks, or ML infrastructure used by other... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Embedded ML Inference & Optimization Engineer. Be the first to apply!