Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Embedded Software Engineer - Inference AI/ML

$195k - $261k

Allen Control Systems

Job Description

Job Description

Company Overview
ACS (Allen Control Systems) is a defense technology company building precision robotic systems for the United States and its allies. Founded by two former U.S. Navy electrical engineers with deep experience in robotics and software, ACS brings together AI, computer vision, precision motion, and advanced hardware to solve complex defense challenges across land, air, and maritime environments.

Our flagship product, Bullfrog, is an autonomous precision weapon system that transforms existing weapons into highly accurate counter-drone systems — giving warfighters a scalable, cost-effective response to one of the fastest-growing threats on the modern battlefield. Bullfrog is deployed with U.S. forces, and ACS works with organizations throughout the U.S. military and national security community.

Following a $200 million Series B at a $2.2 billion valuation, ACS is rapidly expanding manufacturing, accelerating Bullfrog deployments, and developing the next generation of autonomous battlefield systems. This is an opportunity to join a proven, fast-moving team and help scale technology with direct, real-world impact on national security.

ACS is headquartered in Austin, Texas, with additional operations in Alexandria, Virginia; Mountain View, California; and Huntsville, Alabama. For more information, visit allencontrolsystems.com .

About The Role

We are looking for a Senior Embedded Software Engineer - Inference AI/ML to own the end-to-end process of taking trained ML models and deploying them efficiently onto resource-constrained edge hardware. This role sits at the intersection of machine learning, embedded systems, and hardware engineering.

You will integrate, convert, and optimize models to run within strict constraints on latency, memory, power, and thermal budget, and build the supporting C++ infrastructure that hosts them on device. You will partner closely with the CV/ML Engineering team who build the models, the Embedded and Firmware teams who own the device, and the product team who define performance targets. Success means models that are not just accurate in the lab but fast, small, and dependable in the field.

What You’ll Do

  • Apply quantization, pruning, knowledge distillation, operator fusion, and graph optimization to shrink models and reduce inference cost while protecting accuracy; convert trained models into edge-deployable formats using ONNX and TensorRT.

  • Profile inference on target accelerators including GPUs, NPUs, DSPs, and FPGAs; measure latency, throughput, memory footprint, and power consumption, then drive the changes needed to hit performance targets.

  • Design, write, and maintain the C++ application code that hosts inference on device, including pre- and post-processing pipelines, data and memory management, threading, and interfaces to the rest of the embedded system; ensure the combined model and C++ stack meets real-time constraints and fits within device memory budget.

  • Build test harnesses to verify on-device accuracy against reference results and catch regressions from optimization or quantization; contribute to tooling for packaging, versioning, and delivering model updates to deployed devices.

  • Set best practices for edge deployment, review designs and code, and mentor other engineers on optimization and embedded ML techniques; work closely with research, firmware, and product teams to set realistic performance targets and feed hardware constraints back into model design.

What You’ll Need

  • 10+ years of professional embedded software or systems engineering experience, including at least 2 years focused on deploying ML models to embedded or edge devices; Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, Computer Engineering, or equivalent practical experience.

  • Very strong C++ proficiency; working knowledge of CUDA; hands-on experience with PyTorch and at least one edge inference runtime such as TensorFlow Lite, ONNX Runtime, or TensorRT.

  • Practical experience with model optimization techniques including post-training quantization, quantization-aware training, pruning, and distillation; demonstrated ability to profile and optimize for latency, memory, and power on constrained hardware.

  • Working knowledge of embedded or edge platforms such as NVIDIA Jetson, Qualcomm, ARM Cortex, or comparable NPUs and SoCs, and of Linux or an RTOS; solid grasp of computer architecture concepts relevant to inference including memory hierarchy, fixed-point arithmetic, and accelerator offload; domain experience in computer vision or sensor processing on device.

You’ll Stand Out

  • Hands-on experience deploying computer vision models for detection or tracking tasks on embedded or edge hardware.

  • Experience with NVIDIA Jetson specifically, including TensorRT optimization and deployment on Jetson platforms.

  • Background in defense, autonomous systems, or robotics where real-time reliability matters.

  • Experience building or contributing to model update and OTA delivery pipelines for deployed edge devices.

What We Offer

  • Competitive salary

  • ACS Equity Package

  • Health, Dental, Vision Insurance

  • Paid Time Off

Allen Control Systems is an Equal Opportunity Employer, providing equal employment opportunities to all employees and applicants for employment. Allen Control Systems prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws. #LI-AS1

Compensation Range: $195K - $261K

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Senior Embedded Software Engineer - Inference AI/ML in Mountain View, CA vacancy
  •  ...Systems, Inc. is looking for a Senior Performance Engineer to enhance the performance benchmarking...  ...pricing models for their AI chip. The ideal candidate will...  ...experience with open-source inference frameworks and an understanding of ML systems. This role critically combines... 
    Senior

    Cerebras Systems, Inc.

    Sunnyvale, CA
    3 days ago
  • $124.5k - $272k

     ...networking for the cloud and AI era. We secure and...  ...platform, its Zero Trust Engine, and the powerful NewEdge...  ...are available at Senior Staff and above. Candidates...  ...Scientist, you own the inference and optimization layer that...  ...with 4+ years hands-on in ML/AI (model development,... 
    Senior

    Netskope

    Santa Clara, CA
    23 hours ago
  •  ...the world's largest AI chip, 56 times...  ...leading training and inference speeds; over 10 times...  ...by the Wafer-Scale Engine (WSE). This team will...  ...and running software reliably and at scale...  ...environments.Mentor senior SREs, support critical...  ....Background in AI/ML inference systems,... 
    Suggested
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    2 days ago
  •  ...headquartered in Santa Clara, CA, seeks a Principal System Software Engineer for AI Inference Execution. You will join the software team to productize...  ...stack, developing deployment software and collaborating with ML, compiler, and hardware experts. Required: strong... 
    Senior

    Jobleads-US

    Santa Clara, CA
    1 day ago
  • A leader in AI technology in Palo Alto is seeking a Senior AI Systems Performance Engineer to optimize the latest foundation models on their innovative platform. This role involves...  ...learning, performance optimization, and major ML frameworks. This position offers competitive... 
    Senior

    SambaNova

    Palo Alto, CA
    3 days ago
  •  ...Systems in Mountain View, CA, is seeking a Senior Embedded Machine Learning Engineer to own end-to-end deployment of trained ML models onto resource-constrained edge hardware...  ...build the C++ infrastructure that hosts inference on devices. You’ll work closely with CVML,... 
    Senior

    Allen Control Systems

    Mountain View, CA
    3 days ago
  •  ...builds the world's largest AI chip, 56 times larger...  ...-leading training and inference speeds; over 10 times...  ...Cerebras' Wafer-Scale Engine (WSE). We build the compiler...  ...responsible for ML model compilation and optimization...  ...future hardware/software co-design through ML model... 
    Senior

    Cerebras Systems

    Sunnyvale, CA
    2 days ago
  • $332k

     ...unlimited potential of AI to define the next era of...  ...world.We are looking for a Senior leader to orchestrate embedded NVIDIA AI software Go-To-Market strategy...  ...Implement a leadership Inference go-to-market strategy!Strategic...  ...pre-sales with an AI/ML focus.Passion for... 
    Senior
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...the potential of generative AI to power the transformation of...  .... We are at the forefront of software and hardware innovation, pushing...  ...: Principal System Software Engineer, AI Inference ExecutionWhat you will do:The...  ...closely with other software (ML and compilers) and hardware... 
    3 days per week

    d-Matrix

    Santa Clara, CA
    4 days ago
  •  ...experiences—from AI and data...  ...PCs, gaming and embedded systems. Grounded...  ...AMD is seeking a Senior Product Manager...  ...open-source GPU software stack, with a...  ...large-scale model inference on AMD Instinct...  ...will influence engineering roadmaps,...  ...open-source AI/ML community: monitor... 
    Senior
    Remote work

    AMD

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

    We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency....  ...research that pushes the pareto frontier for the field of ML Systems; survey recent publications and find a way to... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $224k - $356.5k

     ...is NVIDIA’s workstation-class AI computer—built on GB300...  ...applications like NemoClaw, LLM inference via NIM, Hermes agents, and deep...  ...a deeply technical systems software engineer who will own AI stack readiness...  ...hands-on experience in AI/ML workload optimization, GPU performance... 
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    2 days ago
  • $170.6k - $261.3k

     ...About the team: The AV ML Infra team at GM...  ...unique demands of AI and ML innovation,...  ...productivity of ML engineers, and drive the adoption...  ...: AI Validation & Inference: Ensures robust...  ...Position Overview: As a Senior AI/ML Full-Stack...  ...and build end-to-end software products, owning... 
    Senior
    Full time
    Local area
    Work from home
    Flexible hours

    General Motors

    Sunnyvale, CA
    3 days ago
  • $193.3k - $261.5k

     ...Amazon Neuron, the software development kit used...  ...and Trainium ML accelerators. This...  ...enabling unparalleled ML inference and training performance...  ...boundary, our engineers build systematic infrastructure...  ...what's possible in AI acceleration.As...  ...mentorship. Our senior members enjoy one-... 
    Senior
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  • $166k - $220k

    A leading technology company in Mountain View is seeking an experienced backend or infrastructure engineer to join their ML/AI Environments team. The ideal candidate will have 5+ years of experience and strong programming skills in Python, Scala, or Java. You will be responsible... 
    Senior

    Jobleads-US

    Mountain View, CA
    23 hours ago
  • $168k - $258.75k

    Inference is the fastest growing and most competitive area in Generative AI today. It is where AI models impact our...  ...deployment techniques. As a Senior Product Manager for...  ...and optimization software (ex. vLLM, SGLang,...  ...Science, Computer Engineering, or similar experience... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $150k - $225k

     ...overall quality of life. As a Senior AI / Embedded Engineer, you will be responsible...  ..., low-power, real-time ML systems that operate at the...  ...lightweight model design, embedded software, and hybrid LLM integration...  ...to reduce model size and inference latency ◦ Use frameworks... 
    Senior
    Full time
    Work at office
    Immediate start
    Visa sponsorship
    Night shift

    E-Space

    Saratoga, CA
    23 hours ago
  •  ...the world's largest AI chip, 56 times...  ...leading training and inference speeds; over 10...  ...closely with Hardware Engineering, Inference...  ...operational risks to senior leadershipRequired...  ...skillsPreferred Experience AI/ML, HPC, or...  ...are serious about software make their own hardware... 
    Senior

    Cerebras Systems

    Sunnyvale, CA
    4 days ago
  • $195k - $230k

     ...information powered by advanced AI, recommendation systems,...  ...RoleWe are looking for a Senior Machine Learning Engineer to help evolve our large-scale...  ...offline training online inference A/B experimentation metric...  ...with large-scale data and ML systems (e.g., Spark, distributed... 
    Senior
    Full time
    Local area
    Work from home

    News Break

    Mountain View, CA
    23 hours ago
  •  ...builds the world's largest AI chip, 56 times larger...  ...-leading training and inference speeds; over 10 times...  ...infrastructure, and the software stack that serves the world...  ...opportunity for engineers who enjoy infrastructure...  ...tools.Experience with ML inference infrastructure... 
    Senior
    Work at office

    Cerebras Systems

    Sunnyvale, CA
    2 days ago
  • $193.93k - $352.29k

     ...most immediate and profound opportunity for AI to drive positive change in the physical...  ...other leading investors.About the RoleThe ML Infrastructure team is responsible for building...  ...-road validation.Maintain an in-house ML inference platform to serve large language models... 
    Senior
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    1 day ago
  • $184k - $287.5k

     ...Architect with a performance engineering background who can help...  ...accelerate Physical AI workloads using NVIDIA'...  ...NVIDIA hardware and software. If you are driven by innovation...  ...of hands-on validated ML/DL performance...  ...scale model training and inference.Effective verbal/... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...Systems builds the world's largest AI chip, 56 times larger than GPUs. This...  ...deliver industry-leading training and inference speeds; over 10 times faster than...  ...Core Infrastructure team builds the software systems that power engineering workflows across Cerebras.Our infrastructure... 
    Senior

    Cerebras Systems

    Sunnyvale, CA
    3 days ago
  • $174k - $252k

     ...development, and optimization of software components critical for Large Language Model (LLM) inference serving on GDC. This includes...  ...(Kubernetes, Google Kubernetes Engine (GKE)), networking infrastructure...  ...development, including AI/ML applications.Preferred qualifications... 
    Senior

    Google

    Sunnyvale, CA
    2 days ago
  •  ...Senior Enterprise Sales Executive AI Inference Infrastructure Tensordyne (formerly Recogni) is building the next generation of AI inference infrastructure...  ...full enterprise sales cycle and work closely with engineering and leadership to bring a disruptive platform to... 
    Senior

    Tensordyne

    Sunnyvale, CA
    2 days ago
  •  ...the world's largest AI chip, 56 times...  ...leading training and inference speeds; over 10 times...  ...with optimization engineers to implement techniques...  ...above the level of Senior PM.5+ years of...  ...experience (e.g. SWE, ML researcher,...  ...are serious about software make their own hardware... 
    Senior
    Work experience placement
    Work at office
    Remote work
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    4 days ago
  • $262k - $365k

     ...models of in app AI (tiny Gemma/Juno Nano...  ...'s on-device ML infrastructure (e....  ...of on-device model inference via optimizations...  ...of experience in software development.7 years...  ...degree or PhD in Engineering, Computer Science,...  ...web browsers, or embedded devices).Google's... 
    Senior

    Google

    Sunnyvale, CA
    1 day ago
  • $163.2k - $220.8k

     ...growth. Wilson Sonsini is looking for a Senior AI Security Engineer to join the Security Operations team....  ..., and supply chain attacks on AI/ML dependenciesEngineer secure data pipelines...  ...model API keys, network isolation for AI inference endpoints, and identity-aware proxy... 
    Senior
    Full time
    Work experience placement
    Remote work
    Worldwide
    Shift work

    Wilson Sonsini Goodrich & Rosati

    Palo Alto, CA
    2 days ago
  • $193.3k - $261.5k

     ...about large-scale systems, AI innovation, and...  ...Architect and build scalable ML infrastructure powering...  ...excellence* Raise the engineering bar through technical...  ...internship professional software development experience-...  ...architecture, training/inference lifecycles, and optimization... 
    Senior
    Internship
    Local area
    Flexible hours

    Amazon

    Palo Alto, CA
    4 days ago
  • $194k

     ...world. About the TeamThe Search AI Product & Mobile team sits...  ...high velocity, work closely with ML and backend platform teams,...  ...looking for a Manager - Mobile Enginerring, who will lead the Search API...  ...integration of ML models (on-device inference, server-side ranking signal... 
    Temporary work

    Coupang

    Mountain View, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Embedded Software Engineer - Inference AI/ML. Be the first to apply!