Senior Embedded Software Engineer - Inference AI/ML
$195k - $261kAllen Control Systems
Job Description
Job Description
Company Overview
Allen Control Systems (ACS) is a cutting-edge defense startup founded by two former Navy electrical engineers with a proven track record in robotics and software. We are developing an autonomous gun turret using advanced computer vision and control systems to precisely detect, track, and neutralize enemy drones.
With an engineering-first culture, ACS values technical excellence and innovation. Backed by our founders’ successful exits from two previous ventures acquired for a combined $180M in 2022, we are committed to ensuring that the groundbreaking technologies we develop will have a real-world impact.
About The Role
We are looking for a Senior Embedded Software Engineer - Inference AI/ML to own the end-to-end process of taking trained ML models and deploying them efficiently onto resource-constrained edge hardware. This role sits at the intersection of machine learning, embedded systems, and hardware engineering.
You will integrate, convert, and optimize models to run within strict constraints on latency, memory, power, and thermal budget, and build the supporting C++ infrastructure that hosts them on device. You will partner closely with the CV/ML Engineering team who build the models, the Embedded and Firmware teams who own the device, and the product team who define performance targets. Success means models that are not just accurate in the lab but fast, small, and dependable in the field.
What You’ll Do
Apply quantization, pruning, knowledge distillation, operator fusion, and graph optimization to shrink models and reduce inference cost while protecting accuracy; convert trained models into edge-deployable formats using ONNX and TensorRT.
Profile inference on target accelerators including GPUs, NPUs, DSPs, and FPGAs; measure latency, throughput, memory footprint, and power consumption, then drive the changes needed to hit performance targets.
Design, write, and maintain the C++ application code that hosts inference on device, including pre- and post-processing pipelines, data and memory management, threading, and interfaces to the rest of the embedded system; ensure the combined model and C++ stack meets real-time constraints and fits within device memory budget.
Build test harnesses to verify on-device accuracy against reference results and catch regressions from optimization or quantization; contribute to tooling for packaging, versioning, and delivering model updates to deployed devices.
Set best practices for edge deployment, review designs and code, and mentor other engineers on optimization and embedded ML techniques; work closely with research, firmware, and product teams to set realistic performance targets and feed hardware constraints back into model design.
What You’ll Need
10+ years of professional embedded software or systems engineering experience, including at least 2 years focused on deploying ML models to embedded or edge devices; Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, Computer Engineering, or equivalent practical experience.
Very strong C++ proficiency; working knowledge of CUDA; hands-on experience with PyTorch and at least one edge inference runtime such as TensorFlow Lite, ONNX Runtime, or TensorRT.
Practical experience with model optimization techniques including post-training quantization, quantization-aware training, pruning, and distillation; demonstrated ability to profile and optimize for latency, memory, and power on constrained hardware.
Working knowledge of embedded or edge platforms such as NVIDIA Jetson, Qualcomm, ARM Cortex, or comparable NPUs and SoCs, and of Linux or an RTOS; solid grasp of computer architecture concepts relevant to inference including memory hierarchy, fixed-point arithmetic, and accelerator offload; domain experience in computer vision or sensor processing on device.
You’ll Stand Out
Hands-on experience deploying computer vision models for detection or tracking tasks on embedded or edge hardware.
Experience with NVIDIA Jetson specifically, including TensorRT optimization and deployment on Jetson platforms.
Background in defense, autonomous systems, or robotics where real-time reliability matters.
Experience building or contributing to model update and OTA delivery pipelines for deployed edge devices.
What We Offer
Competitive salary
ACS Equity Package
Health, Dental, Vision Insurance
Paid Time Off
Allen Control Systems is an Equal Opportunity Employer, providing equal employment opportunities to all employees and applicants for employment. Allen Control Systems prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws. #LI-AS1
Compensation Range: $195K - $261K
- ...Systems, Inc. is looking for a Senior Performance Engineer to enhance the performance benchmarking... ...pricing models for their AI chip. The ideal candidate will... ...experience with open-source inference frameworks and an understanding of ML systems. This role critically combines...Senior
- Accellor is seeking a Technical Architect — AI Systems, Inference & Platform Internals to design, scale, and optimize internal AI systems powering... ...reliability. The ideal candidate has 10-12 years of software/ML infra experience, deep knowledge of PyTorch/JAX, and hands-...Senior
$182.5k - $260.5k
...networking for the cloud and AI era. We secure and... ...platform, its Zero Trust Engine, and the powerful NewEdge... ...are available at Senior Staff and above. Candidates... ...Scientist, you own the inference and optimization layer that... ...with 4+ years hands-on in ML/AI (model development,...Senior- ...builds the world's largest AI chip, 56 times larger... ...-leading training and inference speeds; over 10 times... ...are paying attention. As Senior Product Marketing... ...influencers, and popular software communities Feature Cerebras... ...experience in AI, ML infrastructure, or developer...SeniorShift workNight shift
- ...the world's largest AI chip, 56 times... ...leading training and inference speeds; over 10 times... ...by the Wafer-Scale Engine (WSE). This team will... ...and running software reliably and at scale... ...environments.Mentor senior SREs, support critical... ....Background in AI/ML inference systems,...SuggestedShift work
- ...builds the world’s largest AI chip with wafer-scale... ...industry-leading training and inference speeds and allows users to run large-scale ML applications with less... .... About The Role Senior Director of Technical Product... ...of product, engineering, sales, and marketing to...Senior
- ...headquartered in Santa Clara, CA, seeks a Principal System Software Engineer for AI Inference Execution. You will join the software team to productize... ...stack, developing deployment software and collaborating with ML, compiler, and hardware experts. Required: strong...Senior
- A leader in AI technology in Palo Alto is seeking a Senior AI Systems Performance Engineer to optimize the latest foundation models on their innovative platform. This role involves... ...learning, performance optimization, and major ML frameworks. This position offers competitive...Senior
- ...Systems in Mountain View, CA, is seeking a Senior Embedded Machine Learning Engineer to own end-to-end deployment of trained ML models onto resource-constrained edge hardware... ...build the C++ infrastructure that hosts inference on devices. You’ll work closely with CVML,...Senior
- ...the potential of generative AI to power the transformation of... .... We are at the forefront of software and hardware innovation, pushing... ...: Principal System Software Engineer, AI Inference ExecutionWhat you will do:The... ...closely with other software (ML and compilers) and hardware...3 days per week
$332k
...unlimited potential of AI to define the next era of... ...world.We are looking for a Senior leader to orchestrate embedded NVIDIA AI software Go-To-Market strategy... ...Implement a leadership Inference go-to-market strategy!Strategic... ...pre-sales with an AI/ML focus.Passion for...SeniorFull timeWorldwide- ...experiences—from AI and data... ...PCs, gaming and embedded systems. Grounded... ...AMD is seeking a Senior Product Manager... ...open-source GPU software stack, with a... ...large-scale model inference on AMD Instinct... ...will influence engineering roadmaps,... ...open-source AI/ML community: monitor...SeniorRemote work
$184k - $287.5k
We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency.... ...research that pushes the pareto frontier for the field of ML Systems; survey recent publications and find a way to...SeniorFull time$224k - $356.5k
...is NVIDIA’s workstation-class AI computer—built on GB300... ...applications like NemoClaw, LLM inference via NIM, Hermes agents, and deep... ...a deeply technical systems software engineer who will own AI stack readiness... ...hands-on experience in AI/ML workload optimization, GPU performance...SeniorFull timeLocal area$170.6k - $261.3k
...About the team: The AV ML Infra team at GM... ...unique demands of AI and ML innovation,... ...productivity of ML engineers, and drive the adoption... ...: AI Validation & Inference: Ensures robust... ...Position Overview: As a Senior AI/ML Full-Stack... ...and build end-to-end software products, owning...SeniorFull timeLocal areaWork from homeFlexible hours$193.3k - $261.5k
...Amazon Neuron, the software development kit used... ...and Trainium ML accelerators. This... ...enabling unparalleled ML inference and training performance... ...boundary, our engineers build systematic infrastructure... ...what's possible in AI acceleration.As... ...mentorship. Our senior members enjoy one-...SeniorWork experience placementInternshipLocal areaFlexible hours$229.9k - $262.4k
Senior Lead AI Engineer (FM Hosting, LLM Inference) Overview: At Capital One, we are creating responsible and reliable... ...real time, our applications of AI & ML are bringing humanity and... ...develop, test, deploy, and support AI software components including foundation model...SeniorFull timePart timeLocal area- ...builds the world’s largest AI chip, 56 times larger... ...-leading training and inference speeds; over 10 times... ...Cerebras' Wafer-Scale Engine (WSE). We build the compiler... ...responsible for ML model compilation and optimization... ...future hardware/software co-design through ML model...Senior
- Tensordyne seeks an experienced Sr./Director of Technical Product Management to own AI inference compute in datacenters—from silicon to software. You will report to the VP of Product Management, guiding product efforts across hardware, software, and infrastructure to shape...Senior
- ...Sunnyvale, California. In this role, you will oversee capacity planning and fleet strategy for the AI Inference Service organization, collaborating closely with teams across Engineering, Product, and Operations. Your responsibilities include managing daily deployment tracking,...Senior
$166k - $220k
A leading technology company in Mountain View is seeking an experienced backend or infrastructure engineer to join their ML/AI Environments team. The ideal candidate will have 5+ years of experience and strong programming skills in Python, Scala, or Java. You will be responsible...Senior$168k - $258.75k
Inference is the fastest growing and most competitive area in Generative AI today. It is where AI models impact our... ...deployment techniques. As a Senior Product Manager for... ...and optimization software (ex. vLLM, SGLang,... ...Science, Computer Engineering, or similar experience...SeniorFull time$150k - $225k
...overall quality of life. As a Senior AI / Embedded Engineer, you will be responsible... ..., low-power, real-time ML systems that operate at the... ...lightweight model design, embedded software, and hybrid LLM integration... ...to reduce model size and inference latency ◦ Use frameworks...SeniorFull timeWork at officeImmediate startVisa sponsorshipNight shift- ...Inc. is looking for a Sr. Member of Technical Staff to design software features that enhance system resiliency and high availability... ...distributed environments. The role includes developing scalable AI inference services and deploying cloud-based workflows. Ideal candidates...Senior
- Accelloris is an AI-native services firm focused on operationalizing... ...through advanced AI, data, and engineering capabilities. A Technical Architect — AI Systems, Inference & Platform Internals is sought... ..., and agentic workflows. This senior role emphasizes GPU-level...Senior
- ...the world's largest AI chip, 56 times... ...leading training and inference speeds; over 10... ...closely with Hardware Engineering, Inference... ...operational risks to senior leadershipRequired... ...skillsPreferred Experience AI/ML, HPC, or... ...are serious about software make their own hardware...Senior
- ...data centers, delivering low-latency, high-throughput AI for multi-node GPU workloads. As a Senior Engineer, you will shape core infrastructure and... ...decisions, lead performance optimizations, and own the inference engine to scale research and production workloads....Senior
$230k - $250k
Cerebras Systems is seeking a Sr. Member of Technical Staff in Sunnyvale, CA. This role involves designing resilient software features for cloud-based AI inference, leveraging AWS tools and services. Candidates should have a Master’s degree in Computer Science and experience...Senior$195k - $230k
...information powered by advanced AI, recommendation systems,... ...RoleWe are looking for a Senior Machine Learning Engineer to help evolve our large-scale... ...offline training online inference A/B experimentation metric... ...with large-scale data and ML systems (e.g., Spark, distributed...SeniorFull timeLocal areaWork from home- ...builds the world's largest AI chip, 56 times larger... ...-leading training and inference speeds; over 10 times... ...infrastructure, and the software stack that serves the world... ...opportunity for engineers who enjoy infrastructure... ...tools.Experience with ML inference infrastructure...SeniorWork at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Embedded Software Engineer - Inference AI/ML. Be the first to apply!
- embedded software engineer Mountain View, CA
- embedded systems software engineer Mountain View, CA
- embedded engineer Mountain View, CA
- embedded developer Mountain View, CA
- senior associate architect Mountain View, CA
- senior dynamics crm developer Mountain View, CA
- senior application security Mountain View, CA
- senior account director Mountain View, CA
- sr hr business partner Mountain View, CA
- senior plumbing designer Mountain View, CA


