Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Runtime Optimization Engineer

$159.05k - $199.3k

Applied Intuition

About Applied Intuition Applied Intuition, Inc. is powering the future of physical AI. Founded in 2017 and now valued at $15 billion, the Silicon Valley company is creating the digital infrastructure needed to bring intelligence to every moving machine on the planet. Applied Intuition services the automotive, defense, trucking, construction, mining and agriculture industries in three core areas: tools and infrastructure, operating systems, and autonomy. Eighteen of the top 20 global automakers, as well as the United States military and its allies, trust the company’s solutions to deliver physical intelligence. Applied Intuition is headquartered in Sunnyvale, California, with offices in Washington, D.C.; San Diego; Ft. Walton Beach, Florida; Ann Arbor, Michigan; London; Stuttgart; Munich; Stockholm; Bangalore; Seoul; and Tokyo. Learn more at applied.co. We are an in-office company, and our expectation is that employees primarily work from their Applied Intuition office 5 days a week. However, we also recognize the importance of flexibility and trust our employees to manage their schedules responsibly. This may include occasional remote work, starting the day with morning meetings from home before heading to the office, or leaving earlier when needed to accommodate family commitments. About the role We are looking for a software engineer with deep experience in optimizing ML models and deploying them on production-grade embedded runtime environments. You’ll work across the entire ML framework stack (e.g. PyTorch, JAX, ONNX, TensorRT, CUDA, XLA, Triton). At Applied Intuition, you will: Drive ML performance optimization on multiple technologies for on-road and off-road ADAS / AD stacks targeting deployment on a variety of embedded compute platforms Develop compute usage strategies to optimize efficiency and latency of model inference for compute boards selected by our customers Work on model pruning and quantization, and support deployment on memory constrained platforms Collaborate closely with ML engineers and software developers on technical efforts to find and optimize efficient model architecture solutions Set up methodologies to profile the model performance on target embedded compute platforms and identify performance bottlenecks as part of stack integration We’re looking for someone who has: Bachelors in Electrical Engineering or Computer Science, OR B.Sc. in Computer Science, Mathematics, Physics or a related field 3+ years of experience with ML accelerators, GPU, CPU, SoC architecture and micro-architecture Strong software development skills with the focus on embedded programming Experience profiling and optimizing model performance on embedded compute platforms Experience in working with deep learning frameworks (e.g., PyTorch, JAX, ONNX, etc.) Nice to have: M.Sc or PhD in a ML related area Built an MLoptimization framework from scratch before Deployed ML solutions to embedded chips for real time robotics applications Compensation at Applied Intuition for eligible roles includes base salary, equity, and benefits. Base salary is a single component of the total compensation package, which may also include equity in the form of options and/or restricted stock units, comprehensive health, dental, vision, life and disability insurance coverage, 401k retirement benefits with employer match, learning and wellness stipends, and paid time off. Note that benefits are subject to change and may vary based on jurisdiction of employment. Applied Intuition pay ranges reflect the minimum and maximum intended target base salary for new hire salaries for the position. The actual base salary offered to a successful candidate will additionally be influenced by a variety of factors including experience, credentials & certifications, educational attainment, skill level requirements, interview performance, and the level and scope of the position. Please reference the job posting’s subtitle for where this position will be located. For pay transparency purposes, the base salary range for this full-time position in the location listed is: $159,053 - $199,295 USD annually. Don’t meet every single requirement? If you’re excited about this role but your past experience doesn’t align perfectly with every qualification in the job description, we encourage you to apply anyway. You may be just the right candidate for this or other roles. Applied Intuition is an equal opportunity employer and federal contractor or subcontractor. Consequently, the parties agree that, as applicable, they will abide by the requirements of 41 CFR 60-1.4(a), 41 CFR 60-300.5(a) and 41 CFR 60-741.5(a) and that these laws are incorporated herein by reference. These regulations prohibit discrimination against qualified individuals based on their status as protected veterans or individuals with disabilities, and prohibit discrimination against all individuals based on their race, color, religion, sex, sexual orientation, gender identity or national origin. These regulations require that covered prime contractors and subcontractors take affirmative action to employ and advance in employment individuals without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, protected veteran status or disability. The parties also agree that, as applicable, they will abide by the requirements of Executive Order 13496 (29 CFR Part 471, Appendix A to Subpart A), relating to the notice of employee rights under federal labor laws. #J-18808-Ljbffr

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the ML Runtime Optimization Engineer in Sunnyvale, CA vacancy
  • $166k - $244k

     ...the Role:We are looking for a Software Engineer, Edge Systems & Runtime to join our team, focusing on high-...  ...edge GPUs.Your primary focus will be optimizing our production C++ runtime infrastructure...  .../smart memory management.On-Device ML Deployment: 3+ years hands-on... 
    Suggested
    Full time

    X Company

    Mountain View, CA
    5 days ago
  • $204k - $259k

     ...autonomous driving technology company is looking for an experienced engineer to improve compute performance in machine learning systems. This hybrid role involves collaboration with a world-class ML team and requires strong expertise in ML software or systems. The ideal... 
    Suggested

    Waymo

    Mountain View, CA
    2 days ago
  • $160k - $275k

     ...and software to train and run the largest ML workloads for AGI. MatX is seeking silicon micro-architects and design engineers to join our team as we create best-in-class...  ...dynamic and static power reduction.Drive power optimization across compute, memory, interconnect, PCIe,... 
    Suggested
    Daily paid
    Full time
    Work experience placement
    Local area
    Remote work
    Monday to Friday
    Flexible hours

    MatX

    Mountain View, CA
    1 day ago
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance...  ...across kernel execution, compiler decisions, and runtime scheduling.Direct experience with LLM inference frameworks... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...We are seeking an AI Compiler Engineer with deep expertise in...  ...implement end-to-end compiler optimization workflows, from feature engineering...  ...software engineering and AI/ML experience, preferably in tools...  ...measurable outcomes such as runtime gains, compile-time... 
    Suggested
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • Apple seeks a senior engineer to design the core auction system powering its ads platform at global scale. You will drive optimal outcomes for advertisers, users, and the platform, applying advanced auction theory and machine learning in production environments. You will... 

    Socket

    Cupertino, CA
    1 day ago
  • $106k - $205k

    Title: Staff System Engineers Engineer, AI/MLAbout GlobalFoundries:GlobalFoundries...  ...for a seasoned Staff AI/ML Systems Engineer to lead...  ...how we study, model, and optimize AI/ML workloads for current and...  ...software teams (compilers, runtimes, ML frameworks), serving as the... 
    Full time
    Work at office
    Local area

    Globalfoundries

    Santa Clara, CA
    4 days ago
  •  ...supercomputer — feel like one seamless engine. Developers can write once,...  ...the Role We're looking for a Runtime Engineer to design and build...  ...you'll take the output of our optimizing compiler and make it execute —...  ...the evolving needs of ML engineers and drive improvements... 

    Lemurian Labs Inc.

    Santa Clara, CA
    1 day ago
  • $120k - $200k

     ...to train and run the largest ML workloads for AGI. We primarily...  .../validation, compiler, and runtime teams Contribute to architectural...  ...and implement performance optimizations Requirements Bachelor of...  ...equivalent degree Excellent software engineering skills, with a focus on... 
    Full time
    Work at office
    3 days per week

    Acceler8 Talent

    Mountain View, CA
    2 days ago
  • Lemurian Labs Inc. in Santa Clara, CA, is hiring a Runtime Engineer to design and build the multi-target runtime at the heart of our AI compiler...  ...boundaries. You’ll work with compiler and product teams to optimize execution across diverse targets. #J-18808-Ljbffr Lemurian... 

    Lemurian Labs Inc.

    Santa Clara, CA
    1 day ago
  •  ...automation with Moveworks’ Reasoning Engine and natural language...  ...DescriptionThe RoleWe're building the runtime infrastructure that powers...  ...in real time. This is not an ML role. This is a distributed...  .../per-bot scoping, batch read optimization, and hot-reload... 
    Work at office
    Remote work
    Flexible hours

    Moveworks

    Mountain View, CA
    7 days ago
  • $124k - $195.5k

     ...computing, and low-level hardware optimization has never been more critical....  ...exploratory tools and runtime systems to profile and accelerate...  ...Computer Science, Computer Engineering, Electrical Engineering, or related...  ...deep learning compilers and ML systems, including graph-... 
    Full time

    Nvidia

    Santa Clara, CA
    7 days ago
  • NVIDIA Corporation is seeking systems software engineers to design and ship production C/C++ features for the CUDA driver and runtime in a high-performance, multi-team environment in Santa Clara, CA. You will optimize critical paths, mentor others, and influence hardware... 

    NVIDIA Corporation

    Santa Clara, CA
    3 days ago
  • $90k - $180k

     ...RAPIDS (cuDF/cuML/cuGraph) and optimize end‑to‑end performance. Use...  ...contribute to design reviews and engineering best practices. Mentor peers...  ...services (preferably in AI/ML contexts). Proficiency in:...  ...Triton Inference Server, ONNX Runtime, or TensorRT for high‑throughput... 
    Full time
    Temporary work
    Part time

    Walmart

    Sunnyvale, CA
    4 days ago
  • $224k - $356.5k

     ...Performance Senior Software Engineer to join our energetic team. You...  ...be doing:Play a key role in optimizing system software for Nvidia automotive...  ...teams to track key boot & runtime performance benchmarks.Ensure...  ...software efficiency.AI/ML experience is highly desirable... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $100k

     ...and looking for contributors of all seniorities.As a Software Engineer on the Metal Runtime team at Tenstorrent, you’ll work on the low-level software that powers our AI accelerators. You’ll build and optimize high-performance runtime systems that execute directly on the... 
    Permanent employment

    Tenstorrent

    Santa Clara, CA
    4 days ago
  • $218.8k - $335.3k

     ...with cutting‑edge robotics, optimization, and machine learning to build...  ...looking for a Staff Software Engineer to provide technical leadership...  ...robustness, and predictable runtime behavior under tight latency...  ...mix of analytical models and ML‑based forecasting, including... 
    Full time
    Local area
    Remote work
    Work from home
    Flexible hours

    General Motors

    Sunnyvale, CA
    2 days ago
  •  ...DescriptionIt all started when engineer Fred Luddy wrote code that...  ...DescriptionThe RoleWe're building the runtime infrastructure that powers...  ...in real time. This is not an ML role. This is a distributed...  .../per-bot scoping, batch read optimization, and hot-reload... 
    Work at office
    Immediate start
    Remote work
    Flexible hours

    Moveworks

    Mountain View, CA
    4 days ago
  • $163k - $400k

     ...search ranking, two-sided marketplace matching, and dispatch optimization. This is the engine of a local marketplace: learning-to-rank and retrieval...  ...prove causal impact in a marketplace setting.Bring modern ML/LLM techniques into ranking and matching where they add real... 
    Full time
    Local area
    Work from home

    News Break

    Mountain View, CA
    5 hours ago
  • $152k - $241.5k

     ...reinvented itself over two decades, inventing the GPU in 1999 to reshape PC gaming and modern computer graphics. As an engineer in our EDA Workflow Optimization team, you will partner closely with our engineering teams worldwide. You will understand workflows covering the... 
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    4 days ago
  • $266.05k - $396k

     .... Job Summary Distinguished Engineer - AI Infrastructure We are seeking...  ...Engineer with unrivaled depth in AI/ML inferencing at scale and the...  ...inference engines (TensorRT, vLLM, ONNX Runtime, Triton), model optimization (quantization, pruning, distillation... 
    Part time
    Work at office
    Local area

    NetApp

    San Jose, CA
    1 hour ago
  • $160k - $250k

     ...compiler, and kernels so each layer benefits from the others. The runtime owns the host-side stack and the contracts that bind those...  ...and debuggers — perf counters, traces, and the Python surfaces ML engineers actually use — and hit measurable performance targets on... 
    Daily paid
    Full time
    Contract work
    Work experience placement
    Local area
    Remote work
    Monday to Friday
    Flexible hours

    MatX

    Mountain View, CA
    1 day ago
  • $150.4k - $277.6k

     ...Software and Services The Profiling Tools team is looking for engineers with a passion for optimization to lead new capabilities end-to-end: what gets built,...  ..., and watchOS. Core responsibilities include improving runtime data collection, analysis, and visualization in Apple's... 
    Worldwide
    Relocation

    Apple Inc.

    Cupertino, CA
    4 days ago
  • NVIDIA is seeking systems software engineers in Santa Clara, CA to design and ship production C/C++ features for the CUDA driver and runtime. You will optimize critical execution paths, memory management, and interconnect use while collaborating across teams to drive performance... 

    Socket.dev

    Santa Clara, CA
    2 days ago
  • $198k - $326k

     ...work is centered on trust and optimized for culture, connection,...  ...for a Senior Staff Software Engineer with deep expertise at the intersection...  ...how models interact with runtimes, compilers, and hardware, and...  ...closely with ML, infrastructure, and product... 
    For contractors
    Work at office
    Flexible hours

    Linkedin

    Mountain View, CA
    4 days ago
  • $184.7k - $324.8k

    Machine Learning/ Search Engineer - Services Special Projects Cupertino...  ...to help design, develop, and optimize large-scale search systems....  ...You will work closely with AI/ML Scientists and engineers at the...  ...level experience with inference runtimes/compilers (ONNX Runtime,... 
    Relocation

    Apple

    Cupertino, CA
    2 days ago
  •  ...Clara, CA, headquarters 3 days per week.The role: Senior toStaff Runtime Systems EngineerWhat You Will Do:d-Matrix is developing an AI...  ...in datacenters.This position is for runtime software engineering, working on the architecture, development, and validation of the... 
    3 days per week

    d-Matrix

    Santa Clara, CA
    5 hours ago
  • $163k - $236k

     ...architectural and microarchitectural power optimization techniques.Define best practices and...  ...qualifications:Bachelor's degree in Electrical Engineering, Computer Engineering, Computer Science,...  ..., you’ll work to shape the future of AI/ML hardware acceleration. You will have an... 
    Worldwide

    Google

    Sunnyvale, CA
    6 days ago
  • $207k - $300k

     ...the design and implementation of solutions in specialized ML areas, optimize ML infrastructure, and guide the development of model optimization...  ...Buffers.Preferred qualifications:Master’s degree or PhD in Engineering, Computer Science, or a related technical field.5 years of... 
    Remote work

    Google

    Sunnyvale, CA
    5 hours ago
  • $136k - $218.5k

    We’re looking for a Senior Power Architecture & Optimization Engineer to push the limits of energy efficiency using advanced analytics and AI, including...  ...and productionize power‑aware models and flows, including ML/RL‑based techniques for anomaly detection, dynamic power... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Runtime Optimization Engineer. Be the first to apply!