Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Embedded Software Engineer - Inference AI/ML

$195k - $261k

Allen Control Systems

Job Description

Job Description

Company Overview

Allen Control Systems (ACS) is a cutting-edge defense startup founded by two former Navy electrical engineers with a proven track record in robotics and software. We are developing an autonomous gun turret using advanced computer vision and control systems to precisely detect, track, and neutralize enemy drones.

With an engineering-first culture, ACS values technical excellence and innovation. Backed by our founders’ successful exits from two previous ventures acquired for a combined $180M in 2022, we are committed to ensuring that the groundbreaking technologies we develop will have a real-world impact.

About The Role

We are looking for a Senior Embedded Software Engineer - Inference AI/ML to own the end-to-end process of taking trained ML models and deploying them efficiently onto resource-constrained edge hardware. This role sits at the intersection of machine learning, embedded systems, and hardware engineering.

You will integrate, convert, and optimize models to run within strict constraints on latency, memory, power, and thermal budget, and build the supporting C++ infrastructure that hosts them on device. You will partner closely with the CV/ML Engineering team who build the models, the Embedded and Firmware teams who own the device, and the product team who define performance targets. Success means models that are not just accurate in the lab but fast, small, and dependable in the field.

What You’ll Do

  • Apply quantization, pruning, knowledge distillation, operator fusion, and graph optimization to shrink models and reduce inference cost while protecting accuracy; convert trained models into edge-deployable formats using ONNX and TensorRT.

  • Profile inference on target accelerators including GPUs, NPUs, DSPs, and FPGAs; measure latency, throughput, memory footprint, and power consumption, then drive the changes needed to hit performance targets.

  • Design, write, and maintain the C++ application code that hosts inference on device, including pre- and post-processing pipelines, data and memory management, threading, and interfaces to the rest of the embedded system; ensure the combined model and C++ stack meets real-time constraints and fits within device memory budget.

  • Build test harnesses to verify on-device accuracy against reference results and catch regressions from optimization or quantization; contribute to tooling for packaging, versioning, and delivering model updates to deployed devices.

  • Set best practices for edge deployment, review designs and code, and mentor other engineers on optimization and embedded ML techniques; work closely with research, firmware, and product teams to set realistic performance targets and feed hardware constraints back into model design.

What You’ll Need

  • 10+ years of professional embedded software or systems engineering experience, including at least 2 years focused on deploying ML models to embedded or edge devices; Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, Computer Engineering, or equivalent practical experience.

  • Very strong C++ proficiency; working knowledge of CUDA; hands-on experience with PyTorch and at least one edge inference runtime such as TensorFlow Lite, ONNX Runtime, or TensorRT.

  • Practical experience with model optimization techniques including post-training quantization, quantization-aware training, pruning, and distillation; demonstrated ability to profile and optimize for latency, memory, and power on constrained hardware.

  • Working knowledge of embedded or edge platforms such as NVIDIA Jetson, Qualcomm, ARM Cortex, or comparable NPUs and SoCs, and of Linux or an RTOS; solid grasp of computer architecture concepts relevant to inference including memory hierarchy, fixed-point arithmetic, and accelerator offload; domain experience in computer vision or sensor processing on device.

You’ll Stand Out

  • Hands-on experience deploying computer vision models for detection or tracking tasks on embedded or edge hardware.

  • Experience with NVIDIA Jetson specifically, including TensorRT optimization and deployment on Jetson platforms.

  • Background in defense, autonomous systems, or robotics where real-time reliability matters.

  • Experience building or contributing to model update and OTA delivery pipelines for deployed edge devices.

What We Offer

  • Competitive salary

  • ACS Equity Package

  • Health, Dental, Vision Insurance

  • Paid Time Off

Allen Control Systems is an Equal Opportunity Employer, providing equal employment opportunities to all employees and applicants for employment. Allen Control Systems prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws. #LI-AS1

 

Compensation Range: $195K - $261K

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior Embedded Software Engineer - Inference AI/ML in Mountain View, CA vacancy
  •  ...Systems, Inc. is looking for a Senior Performance Engineer to enhance the performance benchmarking...  ...pricing models for their AI chip. The ideal candidate will...  ...experience with open-source inference frameworks and an understanding of ML systems. This role critically combines... 
    Senior

    Cerebras Systems, Inc.

    Sunnyvale, CA
    3 days ago
  • Accellor is seeking a Technical Architect — AI Systems, Inference & Platform Internals to design, scale, and optimize internal AI systems powering...  ...reliability. The ideal candidate has 10-12 years of software/ML infra experience, deep knowledge of PyTorch/JAX, and hands-... 
    Senior

    Accellor

    Mountain View, CA
    14 hours ago
  • $182.5k - $260.5k

     ...networking for the cloud and AI era. We secure and...  ...platform, its Zero Trust Engine, and the powerful NewEdge...  ...are available at Senior Staff and above. Candidates...  ...Scientist, you own the inference and optimization layer that...  ...with 4+ years hands-on in ML/AI (model development,... 
    Senior

    Netskope

    Santa Clara, CA
    14 hours ago
  •  ...builds the world's largest AI chip, 56 times larger...  ...-leading training and inference speeds; over 10 times...  ...are paying attention. As Senior Product Marketing...  ...influencers, and popular software communities Feature Cerebras...  ...experience in AI, ML infrastructure, or developer... 
    Senior
    Shift work
    Night shift

    Cerebras Systems

    Sunnyvale, CA
    4 days ago
  •  ...the world's largest AI chip, 56 times...  ...leading training and inference speeds; over 10 times...  ...by the Wafer-Scale Engine (WSE). This team will...  ...and running software reliably and at scale...  ...environments.Mentor senior SREs, support critical...  ....Background in AI/ML inference systems,... 
    Suggested
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    2 days ago
  •  ...builds the world’s largest AI chip with wafer-scale...  ...industry-leading training and inference speeds and allows users to run large-scale ML applications with less...  .... About The Role Senior Director of Technical Product...  ...of product, engineering, sales, and marketing to... 
    Senior

    Cerebras

    Sunnyvale, CA
    14 hours ago
  •  ...headquartered in Santa Clara, CA, seeks a Principal System Software Engineer for AI Inference Execution. You will join the software team to productize...  ...stack, developing deployment software and collaborating with ML, compiler, and hardware experts. Required: strong... 
    Senior

    Jobleads-US

    Santa Clara, CA
    1 day ago
  • A leader in AI technology in Palo Alto is seeking a Senior AI Systems Performance Engineer to optimize the latest foundation models on their innovative platform. This role involves...  ...learning, performance optimization, and major ML frameworks. This position offers competitive... 
    Senior

    SambaNova

    Palo Alto, CA
    3 days ago
  •  ...Systems in Mountain View, CA, is seeking a Senior Embedded Machine Learning Engineer to own end-to-end deployment of trained ML models onto resource-constrained edge hardware...  ...build the C++ infrastructure that hosts inference on devices. You’ll work closely with CVML,... 
    Senior

    Allen Control Systems

    Mountain View, CA
    3 days ago
  •  ...the potential of generative AI to power the transformation of...  .... We are at the forefront of software and hardware innovation, pushing...  ...: Principal System Software Engineer, AI Inference ExecutionWhat you will do:The...  ...closely with other software (ML and compilers) and hardware... 
    3 days per week

    d-Matrix

    Santa Clara, CA
    4 days ago
  • $332k

     ...unlimited potential of AI to define the next era of...  ...world.We are looking for a Senior leader to orchestrate embedded NVIDIA AI software Go-To-Market strategy...  ...Implement a leadership Inference go-to-market strategy!Strategic...  ...pre-sales with an AI/ML focus.Passion for... 
    Senior
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...experiences—from AI and data...  ...PCs, gaming and embedded systems. Grounded...  ...AMD is seeking a Senior Product Manager...  ...open-source GPU software stack, with a...  ...large-scale model inference on AMD Instinct...  ...will influence engineering roadmaps,...  ...open-source AI/ML community: monitor... 
    Senior
    Remote work

    AMD

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

    We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency....  ...research that pushes the pareto frontier for the field of ML Systems; survey recent publications and find a way to... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $224k - $356.5k

     ...is NVIDIA’s workstation-class AI computer—built on GB300...  ...applications like NemoClaw, LLM inference via NIM, Hermes agents, and deep...  ...a deeply technical systems software engineer who will own AI stack readiness...  ...hands-on experience in AI/ML workload optimization, GPU performance... 
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    2 days ago
  • $170.6k - $261.3k

     ...About the team: The AV ML Infra team at GM...  ...unique demands of AI and ML innovation,...  ...productivity of ML engineers, and drive the adoption...  ...: AI Validation & Inference: Ensures robust...  ...Position Overview: As a Senior AI/ML Full-Stack...  ...and build end-to-end software products, owning... 
    Senior
    Full time
    Local area
    Work from home
    Flexible hours

    General Motors

    Sunnyvale, CA
    3 days ago
  • $193.3k - $261.5k

     ...Amazon Neuron, the software development kit used...  ...and Trainium ML accelerators. This...  ...enabling unparalleled ML inference and training performance...  ...boundary, our engineers build systematic infrastructure...  ...what's possible in AI acceleration.As...  ...mentorship. Our senior members enjoy one-... 
    Senior
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  • $229.9k - $262.4k

    Senior Lead AI Engineer (FM Hosting, LLM Inference) Overview: At Capital One, we are creating responsible and reliable...  ...real time, our applications of AI & ML are bringing humanity and...  ...develop, test, deploy, and support AI software components including foundation model... 
    Senior
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    14 hours ago
  •  ...builds the world’s largest AI chip, 56 times larger...  ...-leading training and inference speeds; over 10 times...  ...Cerebras' Wafer-Scale Engine (WSE). We build the compiler...  ...responsible for ML model compilation and optimization...  ...future hardware/software co-design through ML model... 
    Senior

    Cerebras

    Sunnyvale, CA
    5 days ago
  • Tensordyne seeks an experienced Sr./Director of Technical Product Management to own AI inference compute in datacenters—from silicon to software. You will report to the VP of Product Management, guiding product efforts across hardware, software, and infrastructure to shape... 
    Senior

    Tensordyne

    Sunnyvale, CA
    1 day ago
  •  ...Sunnyvale, California. In this role, you will oversee capacity planning and fleet strategy for the AI Inference Service organization, collaborating closely with teams across Engineering, Product, and Operations. Your responsibilities include managing daily deployment tracking,... 
    Senior

    Cerebras

    Sunnyvale, CA
    1 day ago
  • $166k - $220k

    A leading technology company in Mountain View is seeking an experienced backend or infrastructure engineer to join their ML/AI Environments team. The ideal candidate will have 5+ years of experience and strong programming skills in Python, Scala, or Java. You will be responsible... 
    Senior

    Jobleads-US

    Mountain View, CA
    14 hours ago
  • $168k - $258.75k

    Inference is the fastest growing and most competitive area in Generative AI today. It is where AI models impact our...  ...deployment techniques. As a Senior Product Manager for...  ...and optimization software (ex. vLLM, SGLang,...  ...Science, Computer Engineering, or similar experience... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $150k - $225k

     ...overall quality of life. As a Senior AI / Embedded Engineer, you will be responsible...  ..., low-power, real-time ML systems that operate at the...  ...lightweight model design, embedded software, and hybrid LLM integration...  ...to reduce model size and inference latency ◦ Use frameworks... 
    Senior
    Full time
    Work at office
    Immediate start
    Visa sponsorship
    Night shift

    E-Space

    Saratoga, CA
    14 hours ago
  •  ...Inc. is looking for a Sr. Member of Technical Staff to design software features that enhance system resiliency and high availability...  ...distributed environments. The role includes developing scalable AI inference services and deploying cloud-based workflows. Ideal candidates... 
    Senior

    Cerebras Systems, Inc.

    Sunnyvale, CA
    2 days ago
  • Accelloris is an AI-native services firm focused on operationalizing...  ...through advanced AI, data, and engineering capabilities. A Technical Architect — AI Systems, Inference & Platform Internals is sought...  ..., and agentic workflows. This senior role emphasizes GPU-level... 
    Senior

    Worky

    Mountain View, CA
    4 days ago
  •  ...the world's largest AI chip, 56 times...  ...leading training and inference speeds; over 10...  ...closely with Hardware Engineering, Inference...  ...operational risks to senior leadershipRequired...  ...skillsPreferred Experience AI/ML, HPC, or...  ...are serious about software make their own hardware... 
    Senior

    Cerebras Systems

    Sunnyvale, CA
    4 days ago
  •  ...data centers, delivering low-latency, high-throughput AI for multi-node GPU workloads. As a Senior Engineer, you will shape core infrastructure and...  ...decisions, lead performance optimizations, and own the inference engine to scale research and production workloads.... 
    Senior

    Sanas

    Palo Alto, CA
    2 days ago
  • $230k - $250k

    Cerebras Systems is seeking a Sr. Member of Technical Staff in Sunnyvale, CA. This role involves designing resilient software features for cloud-based AI inference, leveraging AWS tools and services. Candidates should have a Master’s degree in Computer Science and experience... 
    Senior

    Cerebras Systems

    Sunnyvale, CA
    3 days ago
  • $195k - $230k

     ...information powered by advanced AI, recommendation systems,...  ...RoleWe are looking for a Senior Machine Learning Engineer to help evolve our large-scale...  ...offline training online inference A/B experimentation metric...  ...with large-scale data and ML systems (e.g., Spark, distributed... 
    Senior
    Full time
    Local area
    Work from home

    News Break

    Mountain View, CA
    14 hours ago
  •  ...builds the world's largest AI chip, 56 times larger...  ...-leading training and inference speeds; over 10 times...  ...infrastructure, and the software stack that serves the world...  ...opportunity for engineers who enjoy infrastructure...  ...tools.Experience with ML inference infrastructure... 
    Senior
    Work at office

    Cerebras Systems

    Sunnyvale, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Embedded Software Engineer - Inference AI/ML. Be the first to apply!