Staff AI Inference and Acceleration Engineer
$180k - $275kFigure
We are looking for a Staff AI Inference & Acceleration Engineer to join the Platform Software team and own the on-board inference architecture for Figure’s humanoid robots. You will be the technical authority on how AI workloads are mapped, optimized, and executed across the robot’s compute hardware — driving down power consumption and cost while meeting the strict latency and reliability demands of a real-time autonomous system.
Responsibilities- Own the on-board inference architecture — mapping models to available accelerators (NPU, GPU, DSP, CPU) based on latency, power, and memory budgets.
- Partition inference workloads across heterogeneous compute resources, balancing real-time performance with power and thermal constraints.
- Define and maintain a system-level compute budget across all inference tasks running on the robot.
- Evaluate next-generation acceleration hardware and contribute to the definition of future compute platform requirements.
- Optimize inference toolchains end-to-end — from model export through runtime execution — for target hardware.
- Apply quantization (INT8, INT4, mixed-precision), pruning, operator fusion, and other compression techniques to reduce compute, memory, and power footprint.
- Profile inference pipelines to identify and eliminate bottlenecks in latency, memory bandwidth, and power consumption.
- Optimize kernel scheduling, memory layout, and data movement across the compute hierarchy.
- Partner closely with the AI/ML team to define model architecture constraints that are hardware-friendly from the outset.
- Work with the Platform Software team on runtime integration, scheduling, and power management.
- Engage with silicon vendors and research teams to track the accelerator landscape and influence hardware roadmaps.
- M.S. or Ph.D. in Computer Engineering, Electrical Engineering, Computer Science, or a related field — or equivalent industry experience.
- At least 8 years of industry experience in hardware acceleration, ML systems, or compute architecture.
- Deep understanding of AI/ML inference — model formats (ONNX, TFLite, etc.), inference runtimes, and deployment pipelines.
- Hands-on experience optimizing models for edge or embedded hardware using quantization, pruning, and operator-level tuning.
- Strong understanding of computer architecture — memory hierarchies, data movement, and heterogeneous compute.
- Experience profiling and benchmarking inference workloads across CPU, GPU, NPU, DSP.
- Familiarity with low-level toolchains and compilation frameworks (e.g. TVM, MLIR, TensorRT, Torch, SNPE/QNN, JAX, CUDA, ROCm).
- Solid software engineering skills in C++ and Python.
- Strong cross-functional communication skills — able to work effectively across hardware, software, and AI/ML teams.
- Knowledge of real-time operating constraints and their impact on inference scheduling.
- Track record of co-designing model architectures with ML teams to meet hardware constraints.
- RadixArk is seeking a Member of Technical Staff to accelerate LLM inference and training on modern GPUs. You will profile, optimize, and extend SGLang... ...contributing to the roadmap and ecosystem integrations that power scalable AI inference and training. #J-18808-Ljbffr RadixArkSuggested
- MatX is building custom silicon for LLM inference and training. You will develop the host-side interface library, manage device memory, DMA, streams and events, and extend the executable format to enable safe evolution of compiler-runtime contracts. You will design the...Suggested
$148.7k - $201.2k
...backbone of Generative AI at AWS? Do you want to... ...cloud for AI training and inference, delivering continuous... ...a Systems Development Engineer to develop automation software... ...infrastructure for our accelerated (AI/ML) server... ...employees, supervisors, and staff; adhere to standards of...SuggestedInternshipLocal areaWorldwideFlexible hours$183k - $247.6k
...to shape the future of AI? Join the team... ...cloud for AI training and inference — where multi-billion-... ...hardware, and network engineers, supply chain specialists... ...performance server and/or accelerator server and rack system... ..., supervisors, and staff; adhere to standards...SuggestedLocal areaFlexible hours- RadixArk is seeking a Member of Technical Staff, Developer Technology (DevTech) to accelerate LLM inference and training on modern GPU hardware. Our engines power trillions of tokens daily and... ...researchers and partners to scale AI workloads. You will profile, optimize,...Suggested
$184k - $287.5k
We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency... ..., and performance teams to push the frontier of accelerated computing for AI.What you’ll be doing:Contribute...Full time- ...Inc. is seeking a Member of Technical Staff (Hardware) to design and run AI hardware benchmarks for GPUs, TPUs... ...standards. The role requires 3+ years in AI accelerators, strong Python/data-analysis skills, and deep knowledge of inference economics. #J-18808-Ljbffr...
- ...building the world's most efficient software for inference and agent hosting. In this role, you'll be one of the first engineers on Sailboxes, contributing across the stack—... ...environment with generous equipment and setup to accelerate #J-18808-Ljbffr Sail Research Inc.Work at office
$122.6k - $185k
...backbone of Generative AI cloud at AWS? Do you... ...cloud for AI training and inference? Want to do industry... ...you. The AWS Hardware Engineering team creates server... ...future/new designs for AWS Accelerated server solutions for... ..., supervisors, and staff; adhere to standards of...Local areaFlexible hours- ...the compiler→runtime contract. You will design the custom-kernel ABI, and implement Python bindings to move tensors from Python to accelerator hardware. The role involves working with CUDA/ROCm-style accelerators, memory models, and high-performance computing stacks,...Contract work
- MakerMaker.AI is looking for a Senior Machine Learning Systems Engineer in San Francisco. In this role, you will build and operate production inference systems, optimizing for performance and reliability... ...have strong knowledge in GPU-accelerated inference. Excellent...
$142.8k - $274.8k
...Hardware, and Infrastructure Engineering (SCHIE) is the team behind Microsoft... ...organization is developing AI-native silicon and hyperscale... ...-leading AI training and inference. The Platform Systems... ...team is seeking a Principal AI Accelerator Tools Development Engineer to...Ongoing contractWork at officeLocal areaWorldwide3 days per week- d‑Matrix is seeking a Staff Analog Design Engineer to lead analog‑mixed signal IC design for our AI inference accelerators. You will own circuit design from schematic through silicon bring‑up, guiding layout engineers on 7nm and below and ensuring low‑power performance...
$148.7k - $201.2k
...available to Generative AI customers? Do you want... ...scale?AWS Hardware Engineering is looking for a Systems... ...- the servers, accelerators, and storage platforms... ...systems for AI training, inference, and compute workloads... ...employees, supervisors, and staff; adhere to standards...InternshipLocal areaWorldwideFlexible hours- An innovative company is seeking a talented software engineer to join their dynamic Inference team. This role involves designing and implementing infrastructure... ...researchers and product teams to push the boundaries of AI technology, ensuring reliable production services. If...
- ...is seeking a Member of Technical Staff to own coverage of the serverless inference landscape. You will benchmark endpoints... ...and keep us ahead in frontier AI benchmarking. You will analyze metrics... ...frontiers, collaborating with engineers and industry leaders. #J-18808-Ljbffr...
$200k - $360k
...professional to own MLIR dialect design and lowering passes for their AI accelerator. You'll work closely with chip-design and software teams.... ...knowledge of tensor compilation, and 5+ years of compiler engineering experience. The role includes a competitive compensation...$175k - $250k
...bit about us We're a well-funded AI infrastructure startup... ...Job Details We're looking for an engineer to help build and maintain a high-performance inference library designed to support modern... ..., ROCm, Triton, or similar GPU/accelerator programming technologies Experience...Local area- A leading AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for designing high-performance GPU kernels and using advanced techniques to boost computation efficiency. Ideal...
- Platform Recruitment is seeking a SIPI Characterization Engineer for an AI hardware client in Mountain View, CA. The role focuses on high-speed SerDes and PDN characterization across AI accelerator systems, from design through production. The candidate will build automated...
- Jobot is seeking an engineer to build and maintain a high-performance inference library for modern AI models across diverse compute architectures. The role targets someone who understands how LLM inference works under the hood and aims to squeeze maximum performance from...
$175k - $250k
Global Inference Library Engineer Experience: Senior Level Salary: $175,000 - $250,000 per year Job Details... ...library designed to support modern AI models across a variety of compute... ...with CUDA, ROCm, Triton, or similar GPU/accelerator programming technologies Experience...$260k - $320k
DensityAI is seeking an expert to develop compute kernels for a specialized AI accelerator in Mountain View, California. Your role will focus on writing performance-critical kernels while collaborating with architecture and compiler teams to enhance silicon design. Qualifications...- ...you. Lucid Motors is seeking a Staff QA and Test Infrastructure Development Engineer to help build the next generation... ...Bonus Qualifications – Agentic AI Tools & Intelligent Automation... ...leveraging agentic AI systems to accelerate development, automate complex workflows...Full timeImmediate start
- ...is seeking a Senior hardware architect to design and implement accelerators for codesigned systems. You will own subsystems from requirements... ...with ML model design and numerics. Google DeepMind values safety and ethics in AI development. #J-18808-Ljbffr Google DeepMind
- Amazon Data Services, Inc. is seeking a Systems Development Engineer to design automation software, diagnostic tooling, and fleet health infrastructure for AI/ML accelerator servers. You will own systems, tackle complex problems, and deliver scalable solutions across compute...
- Causal Labs in San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting against historical physical observations. You will design techniques to improve latency and throughput, optimize...
$150.98k - $218.62k
...Intelligent Edge. ADI combines analog, digital, AI, and software technologies into solutions... .... Learn more at and on LinkedIn and X.Staff System Architecture and Design... ...DCE)San Jose, CAAbout the RoleAs a System Engineer, you will own the system development of high...Permanent employmentFull timeWork at officeShift workDay shift- Cerebras Systems builds the world’s largest AI chip, 56 times larger than GPUs. This... ...deliver industry-leading training and inference speeds; over 10 times faster than GPU-based... ...Master’s/PhD in Computer or Electrical Engineering + 3 years industry experience,...
- Google Sunnyvale is hiring a Chip Package SI/PI Engineer to advance TPU hardware acceleration. You will verify complex digital designs, focusing on TPU architecture and AI/ML system integration, shaping next‑gen silicon solutions used by millions of users. As part of the...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff AI Inference and Acceleration Engineer. Be the first to apply!


