Senior Embedded Software Engineer - Inference AI/ML
$195k - $261kAllen Control Systems
Job Description
Job Description
Company Overview
ACS (Allen Control Systems) is a defense technology company building precision robotic systems for the United States and its allies. Founded by two former U.S. Navy electrical engineers with deep experience in robotics and software, ACS brings together AI, computer vision, precision motion, and advanced hardware to solve complex defense challenges across land, air, and maritime environments.
Our flagship product, Bullfrog, is an autonomous precision weapon system that transforms existing weapons into highly accurate counter-drone systems — giving warfighters a scalable, cost-effective response to one of the fastest-growing threats on the modern battlefield. Bullfrog is deployed with U.S. forces, and ACS works with organizations throughout the U.S. military and national security community.
Following a $200 million Series B at a $2.2 billion valuation, ACS is rapidly expanding manufacturing, accelerating Bullfrog deployments, and developing the next generation of autonomous battlefield systems. This is an opportunity to join a proven, fast-moving team and help scale technology with direct, real-world impact on national security.
ACS is headquartered in Austin, Texas, with additional operations in Alexandria, Virginia; Mountain View, California; and Huntsville, Alabama. For more information, visit allencontrolsystems.com .
About The Role
We are looking for a Senior Embedded Software Engineer - Inference AI/ML to own the end-to-end process of taking trained ML models and deploying them efficiently onto resource-constrained edge hardware. This role sits at the intersection of machine learning, embedded systems, and hardware engineering.
You will integrate, convert, and optimize models to run within strict constraints on latency, memory, power, and thermal budget, and build the supporting C++ infrastructure that hosts them on device. You will partner closely with the CV/ML Engineering team who build the models, the Embedded and Firmware teams who own the device, and the product team who define performance targets. Success means models that are not just accurate in the lab but fast, small, and dependable in the field.
What You’ll Do
Apply quantization, pruning, knowledge distillation, operator fusion, and graph optimization to shrink models and reduce inference cost while protecting accuracy; convert trained models into edge-deployable formats using ONNX and TensorRT.
Profile inference on target accelerators including GPUs, NPUs, DSPs, and FPGAs; measure latency, throughput, memory footprint, and power consumption, then drive the changes needed to hit performance targets.
Design, write, and maintain the C++ application code that hosts inference on device, including pre- and post-processing pipelines, data and memory management, threading, and interfaces to the rest of the embedded system; ensure the combined model and C++ stack meets real-time constraints and fits within device memory budget.
Build test harnesses to verify on-device accuracy against reference results and catch regressions from optimization or quantization; contribute to tooling for packaging, versioning, and delivering model updates to deployed devices.
Set best practices for edge deployment, review designs and code, and mentor other engineers on optimization and embedded ML techniques; work closely with research, firmware, and product teams to set realistic performance targets and feed hardware constraints back into model design.
What You’ll Need
10+ years of professional embedded software or systems engineering experience, including at least 2 years focused on deploying ML models to embedded or edge devices; Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, Computer Engineering, or equivalent practical experience.
Very strong C++ proficiency; working knowledge of CUDA; hands-on experience with PyTorch and at least one edge inference runtime such as TensorFlow Lite, ONNX Runtime, or TensorRT.
Practical experience with model optimization techniques including post-training quantization, quantization-aware training, pruning, and distillation; demonstrated ability to profile and optimize for latency, memory, and power on constrained hardware.
Working knowledge of embedded or edge platforms such as NVIDIA Jetson, Qualcomm, ARM Cortex, or comparable NPUs and SoCs, and of Linux or an RTOS; solid grasp of computer architecture concepts relevant to inference including memory hierarchy, fixed-point arithmetic, and accelerator offload; domain experience in computer vision or sensor processing on device.
You’ll Stand Out
Hands-on experience deploying computer vision models for detection or tracking tasks on embedded or edge hardware.
Experience with NVIDIA Jetson specifically, including TensorRT optimization and deployment on Jetson platforms.
Background in defense, autonomous systems, or robotics where real-time reliability matters.
Experience building or contributing to model update and OTA delivery pipelines for deployed edge devices.
What We Offer
Competitive salary
ACS Equity Package
Health, Dental, Vision Insurance
Paid Time Off
Allen Control Systems is an Equal Opportunity Employer, providing equal employment opportunities to all employees and applicants for employment. Allen Control Systems prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws. #LI-AS1
Compensation Range: $195K - $261K
- ...Systems, Inc. is looking for a Senior Performance Engineer to enhance the performance benchmarking... ...pricing models for their AI chip. The ideal candidate will... ...experience with open-source inference frameworks and an understanding of ML systems. This role critically combines...Senior
$124.5k - $272k
...networking for the cloud and AI era. We secure and... ...platform, its Zero Trust Engine, and the powerful NewEdge... ...are available at Senior Staff and above. Candidates... ...Scientist, you own the inference and optimization layer that... ...with 4+ years hands-on in ML/AI (model development,...Senior- ...the world's largest AI chip, 56 times... ...leading training and inference speeds; over 10 times... ...by the Wafer-Scale Engine (WSE). This team will... ...and running software reliably and at scale... ...environments.Mentor senior SREs, support critical... ....Background in AI/ML inference systems,...SuggestedShift work
- ...headquartered in Santa Clara, CA, seeks a Principal System Software Engineer for AI Inference Execution. You will join the software team to productize... ...stack, developing deployment software and collaborating with ML, compiler, and hardware experts. Required: strong...Senior
- A leader in AI technology in Palo Alto is seeking a Senior AI Systems Performance Engineer to optimize the latest foundation models on their innovative platform. This role involves... ...learning, performance optimization, and major ML frameworks. This position offers competitive...Senior
- ...Systems in Mountain View, CA, is seeking a Senior Embedded Machine Learning Engineer to own end-to-end deployment of trained ML models onto resource-constrained edge hardware... ...build the C++ infrastructure that hosts inference on devices. You’ll work closely with CVML,...Senior
- ...builds the world's largest AI chip, 56 times larger... ...-leading training and inference speeds; over 10 times... ...Cerebras' Wafer-Scale Engine (WSE). We build the compiler... ...responsible for ML model compilation and optimization... ...future hardware/software co-design through ML model...Senior
$332k
...unlimited potential of AI to define the next era of... ...world.We are looking for a Senior leader to orchestrate embedded NVIDIA AI software Go-To-Market strategy... ...Implement a leadership Inference go-to-market strategy!Strategic... ...pre-sales with an AI/ML focus.Passion for...SeniorFull timeWorldwide- ...the potential of generative AI to power the transformation of... .... We are at the forefront of software and hardware innovation, pushing... ...: Principal System Software Engineer, AI Inference ExecutionWhat you will do:The... ...closely with other software (ML and compilers) and hardware...3 days per week
- ...experiences—from AI and data... ...PCs, gaming and embedded systems. Grounded... ...AMD is seeking a Senior Product Manager... ...open-source GPU software stack, with a... ...large-scale model inference on AMD Instinct... ...will influence engineering roadmaps,... ...open-source AI/ML community: monitor...SeniorRemote work
$184k - $287.5k
We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency.... ...research that pushes the pareto frontier for the field of ML Systems; survey recent publications and find a way to...SeniorFull time$224k - $356.5k
...is NVIDIA’s workstation-class AI computer—built on GB300... ...applications like NemoClaw, LLM inference via NIM, Hermes agents, and deep... ...a deeply technical systems software engineer who will own AI stack readiness... ...hands-on experience in AI/ML workload optimization, GPU performance...SeniorFull timeLocal area$170.6k - $261.3k
...About the team: The AV ML Infra team at GM... ...unique demands of AI and ML innovation,... ...productivity of ML engineers, and drive the adoption... ...: AI Validation & Inference: Ensures robust... ...Position Overview: As a Senior AI/ML Full-Stack... ...and build end-to-end software products, owning...SeniorFull timeLocal areaWork from homeFlexible hours$193.3k - $261.5k
...Amazon Neuron, the software development kit used... ...and Trainium ML accelerators. This... ...enabling unparalleled ML inference and training performance... ...boundary, our engineers build systematic infrastructure... ...what's possible in AI acceleration.As... ...mentorship. Our senior members enjoy one-...SeniorWork experience placementInternshipLocal areaFlexible hours$166k - $220k
A leading technology company in Mountain View is seeking an experienced backend or infrastructure engineer to join their ML/AI Environments team. The ideal candidate will have 5+ years of experience and strong programming skills in Python, Scala, or Java. You will be responsible...Senior$168k - $258.75k
Inference is the fastest growing and most competitive area in Generative AI today. It is where AI models impact our... ...deployment techniques. As a Senior Product Manager for... ...and optimization software (ex. vLLM, SGLang,... ...Science, Computer Engineering, or similar experience...SeniorFull time$150k - $225k
...overall quality of life. As a Senior AI / Embedded Engineer, you will be responsible... ..., low-power, real-time ML systems that operate at the... ...lightweight model design, embedded software, and hybrid LLM integration... ...to reduce model size and inference latency ◦ Use frameworks...SeniorFull timeWork at officeImmediate startVisa sponsorshipNight shift- ...the world's largest AI chip, 56 times... ...leading training and inference speeds; over 10... ...closely with Hardware Engineering, Inference... ...operational risks to senior leadershipRequired... ...skillsPreferred Experience AI/ML, HPC, or... ...are serious about software make their own hardware...Senior
$195k - $230k
...information powered by advanced AI, recommendation systems,... ...RoleWe are looking for a Senior Machine Learning Engineer to help evolve our large-scale... ...offline training online inference A/B experimentation metric... ...with large-scale data and ML systems (e.g., Spark, distributed...SeniorFull timeLocal areaWork from home- ...builds the world's largest AI chip, 56 times larger... ...-leading training and inference speeds; over 10 times... ...infrastructure, and the software stack that serves the world... ...opportunity for engineers who enjoy infrastructure... ...tools.Experience with ML inference infrastructure...SeniorWork at office
$193.93k - $352.29k
...most immediate and profound opportunity for AI to drive positive change in the physical... ...other leading investors.About the RoleThe ML Infrastructure team is responsible for building... ...-road validation.Maintain an in-house ML inference platform to serve large language models...SeniorImmediate startFlexible hours$184k - $287.5k
...Architect with a performance engineering background who can help... ...accelerate Physical AI workloads using NVIDIA'... ...NVIDIA hardware and software. If you are driven by innovation... ...of hands-on validated ML/DL performance... ...scale model training and inference.Effective verbal/...SeniorFull timeRemote work- ...Systems builds the world's largest AI chip, 56 times larger than GPUs. This... ...deliver industry-leading training and inference speeds; over 10 times faster than... ...Core Infrastructure team builds the software systems that power engineering workflows across Cerebras.Our infrastructure...Senior
$174k - $252k
...development, and optimization of software components critical for Large Language Model (LLM) inference serving on GDC. This includes... ...(Kubernetes, Google Kubernetes Engine (GKE)), networking infrastructure... ...development, including AI/ML applications.Preferred qualifications...Senior- ...Senior Enterprise Sales Executive AI Inference Infrastructure Tensordyne (formerly Recogni) is building the next generation of AI inference infrastructure... ...full enterprise sales cycle and work closely with engineering and leadership to bring a disruptive platform to...Senior
- ...the world's largest AI chip, 56 times... ...leading training and inference speeds; over 10 times... ...with optimization engineers to implement techniques... ...above the level of Senior PM.5+ years of... ...experience (e.g. SWE, ML researcher,... ...are serious about software make their own hardware...SeniorWork experience placementWork at officeRemote workShift work
$262k - $365k
...models of in app AI (tiny Gemma/Juno Nano... ...'s on-device ML infrastructure (e.... ...of on-device model inference via optimizations... ...of experience in software development.7 years... ...degree or PhD in Engineering, Computer Science,... ...web browsers, or embedded devices).Google's...Senior$163.2k - $220.8k
...growth. Wilson Sonsini is looking for a Senior AI Security Engineer to join the Security Operations team.... ..., and supply chain attacks on AI/ML dependenciesEngineer secure data pipelines... ...model API keys, network isolation for AI inference endpoints, and identity-aware proxy...SeniorFull timeWork experience placementRemote workWorldwideShift work$193.3k - $261.5k
...about large-scale systems, AI innovation, and... ...Architect and build scalable ML infrastructure powering... ...excellence* Raise the engineering bar through technical... ...internship professional software development experience-... ...architecture, training/inference lifecycles, and optimization...SeniorInternshipLocal areaFlexible hours$194k
...world. About the TeamThe Search AI Product & Mobile team sits... ...high velocity, work closely with ML and backend platform teams,... ...looking for a Manager - Mobile Enginerring, who will lead the Search API... ...integration of ML models (on-device inference, server-side ranking signal...Temporary work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Embedded Software Engineer - Inference AI/ML. Be the first to apply!
- embedded engineer Mountain View, CA
- embedded systems software engineer Mountain View, CA
- embedded developer Mountain View, CA
- embedded software engineer Mountain View, CA
- senior network engineer remote Mountain View, CA
- senior app developer Mountain View, CA
- senior manager legal Mountain View, CA
- sr project manager Mountain View, CA
- senior account executive Mountain View, CA
- senior manager strategic initiatives Mountain View, CA

