Staff AI Inference and Acceleration Engineer
$180k - $275kFull-time
Figure
Figure is an AI robotics company developing autonomous general-purpose humanoid robots. The goal of the company is to ship humanoid robots with human level intelligence. Its robots are engineered to perform a variety of tasks in the home and commercial markets. Figure is headquartered in San Jose, CA.
We are looking for a Staff AI Inference & Acceleration Engineer to join the Platform Software team and own the on-board inference architecture for Figure’s humanoid robots. You will be the technical authority on how AI workloads are mapped, optimized, and executed across the robot’s compute hardware — driving down power consumption and cost while meeting the strict latency and reliability demands of a real-time autonomous system.
Responsibilities- Own the on-board inference architecture — mapping models to available accelerators (NPU, GPU, DSP, CPU) based on latency, power, and memory budgets.
- Partition inference workloads across heterogeneous compute resources, balancing real-time performance with power and thermal constraints.
- Define and maintain a system-level compute budget across all inference tasks running on the robot.
- Evaluate next-generation acceleration hardware and contribute to the definition of future compute platform requirements.
- Optimize inference toolchains end-to-end — from model export through runtime execution — for target hardware.
- Apply quantization (INT8, INT4, mixed-precision), pruning, operator fusion, and other compression techniques to reduce compute, memory, and power footprint.
- Profile inference pipelines to identify and eliminate bottlenecks in latency, memory bandwidth, and power consumption.
- Optimize kernel scheduling, memory layout, and data movement across the compute hierarchy.
- Partner closely with the AI/ML team to define model architecture constraints that are hardware-friendly from the outset.
- Work with the Platform Software team on runtime integration, scheduling, and power management.
- Engage with silicon vendors and research teams to track the accelerator landscape and influence hardware roadmaps.
- M.S. or Ph.D. in Computer Engineering, Electrical Engineering, Computer Science, or a related field — or equivalent industry experience.
- At least 8 years of industry experience in hardware acceleration, ML systems, or compute architecture.
- Deep understanding of AI/ML inference — model formats (ONNX, TFLite, etc.), inference runtimes, and deployment pipelines.
- Hands-on experience optimizing models for edge or embedded hardware using quantization, pruning, and operator-level tuning.
- Strong understanding of computer architecture — memory hierarchies, data movement, and heterogeneous compute.
- Experience profiling and benchmarking inference workloads across CPU, GPU, NPU, DSP.
- Familiarity with low-level toolchains and compilation frameworks (e.g. TVM, MLIR, TensorRT, Torch, SNPE/QNN, JAX, CUDA, ROCm).
- Solid software engineering skills in C++ and Python.
- Strong cross-functional communication skills — able to work effectively across hardware, software, and AI/ML teams.
- Knowledge of real-time operating constraints and their impact on inference scheduling.
- Track record of co-designing model architectures with ML teams to meet hardware constraints.
Vacancy posted 6 days ago
Similar jobs that could be interesting for youBased on the Staff AI Inference and Acceleration Engineer in California vacancy
- RadixArk is seeking a Member of Technical Staff to accelerate LLM inference and training on modern GPUs. You will profile, optimize, and extend SGLang... ...contributing to the roadmap and ecosystem integrations that power scalable AI inference and training. #J-18808-Ljbffr RadixArkSuggested
$152k - $241.5k
...deep learning ignited modern AI — the next era of computing —... ...seeking top-tier AI Compiler Engineers to drive innovation within our... ...for AI workloads (both inference and training) and successfully... ...on CPU, GPU, and/or custom AI accelerator architectures.LLM Knowledge:...SuggestedFull time- RadixArk is seeking a Member of Technical Staff, Developer Technology (DevTech) to accelerate LLM inference and training on modern GPU hardware. Our engines power trillions of tokens daily and... ...researchers and partners to scale AI workloads. You will profile, optimize,...Suggested
$184k - $287.5k
We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency... ..., and performance teams to push the frontier of accelerated computing for AI.What you’ll be doing:Contribute...SuggestedFull time- ...Inc. is seeking a Member of Technical Staff (Hardware) to design and run AI hardware benchmarks for GPUs, TPUs... ...standards. The role requires 3+ years in AI accelerators, strong Python/data-analysis skills, and deep knowledge of inference economics. #J-18808-Ljbffr...Suggested
- ...and sustainability, design and engineering, ambition and integrity. In... ...mobility.Lucid Motors is seeking a Staff QA and Test Infrastructure... ...Qualifications – Agentic AI Tools & Intelligent AutomationLucid... ...agentic AI systems to accelerate development, automate complex...
$122.6k - $185k
...backbone of Generative AI cloud at AWS? Do you... ...cloud for AI training and inference? Want to do industry... ...you. The AWS Hardware Engineering team creates server... ...future/new designs for AWS Accelerated server solutions for... ..., supervisors, and staff; adhere to standards of...Local areaFlexible hours$142.8k - $274.8k
...Hardware, and Infrastructure Engineering (SCHIE) is the team behind Microsoft... ...organization is developing AI-native silicon and hyperscale... ...-leading AI training and inference. The Platform Systems... ...team is seeking a Principal AI Accelerator Tools Development Engineer to...Ongoing contractWork at officeLocal areaWorldwide3 days per week- MakerMaker.AI is looking for a Senior Machine Learning Systems Engineer in San Francisco. In this role, you will build and operate production inference systems, optimizing for performance and reliability... ...have strong knowledge in GPU-accelerated inference. Excellent...
- ...is seeking a Member of Technical Staff to own coverage of the serverless inference landscape. You will benchmark endpoints... ...and keep us ahead in frontier AI benchmarking. You will analyze metrics... ...frontiers, collaborating with engineers and industry leaders. #J-18808-Ljbffr...
- Crusoe in San Francisco is seeking a Staff Technical Program Manager to lead the Managed Inference platform team. You will ensure end-to-end program delivery for LLM workloads, driving innovation in AI infrastructure powered by clean energy. The ideal candidate has extensive...
- An innovative company is seeking a talented software engineer to join their dynamic Inference team. This role involves designing and implementing infrastructure... ...researchers and product teams to push the boundaries of AI technology, ensuring reliable production services. If...
- Anyscale is seeking a Distributed LLM Inference Engineer in San Francisco, California. This pivotal role involves pushing the boundaries of performance for ML inference at scale. You'll work closely with product teams to deliver end-to-end solutions while leveraging open...
$200k - $360k
...professional to own MLIR dialect design and lowering passes for their AI accelerator. You'll work closely with chip-design and software teams.... ...knowledge of tensor compilation, and 5+ years of compiler engineering experience. The role includes a competitive compensation...- ...seeks an experienced Deep Learning Software Engineer, TensorRT Performance, to analyze and enhance the performance of NVIDIA’s inference ecosystem, including TensorRT, TensorRT‑... ...scheduling, and collaborate with teams across AI, automotive, robotics and vision domains to...
$260k - $320k
DensityAI is seeking an expert to develop compute kernels for a specialized AI accelerator in Mountain View, California. Your role will focus on writing performance-critical kernels while collaborating with architecture and compiler teams to enhance silicon design. Qualifications...- Acceler8 Talent is hiring a Kernel Engineer to develop and optimize compute kernels for a custom AI accelerator in Mountain View, CA. You will collaborate with architecture, compiler, simulation, and systems teams to define the kernel programming model and validate performance...
- A leading AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for designing high-performance GPU kernels and using advanced techniques to boost computation efficiency. Ideal...
- NVIDIA in Santa Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design and optimize kernels, build... ...for LLM serving engines, and contribute to accelerators and runtimes that power large language models and AI...
$262k - $365k
...rigorous code review, ensuring accuracy and engineering best practices.Serve as a Senior TLM to... ..., "speed-of-light" delivery of modern Accelerator appliances into Google Data Centers.... ...software teams by leveraging the latest AI tools to transform methodologies and accelerate...Worldwide- NVIDIA Corporation in Santa Clara, CA is seeking outstanding AI systems engineers to develop groundbreaking inference technologies for the hardware-accelerated stack. You will create libraries, code generators, and GPU kernel innovations for LLM workloads. Join a team...
- NVIDIA is seeking outstanding AI systems engineers in Santa Clara to advance the inference software stack. You will build libraries, code generators, and GPU kernels... ...for LLM serving engines and JIT compilers to accelerate large language models. The role emphasizes Python...
$150k - $275k
A cutting-edge tech company in San Jose is seeking a Supercomputing Engineer to ensure the reliability of its inference servers. This role involves designing and executing test suites, analyzing performance, and collaborating with engineering teams. Ideal candidates will...- Google in Sunnyvale, CA is seeking a Senior Staff Co-Design Engineer on the TPU Chip Architecture team to shape AI/ML hardware acceleration. You will drive TPU architecture and verification efforts for next-generation training and serving workloads, collaborating with research...
- ...is seeking a Senior hardware architect to design and implement accelerators for codesigned systems. You will own subsystems from requirements... ...with ML model design and numerics. Google DeepMind values safety and ethics in AI development. #J-18808-Ljbffr Google DeepMind
$150.98k - $218.62k
...Intelligent Edge. ADI combines analog, digital, AI, and software technologies into solutions... .... Learn more at and on LinkedIn and X.Staff System Architecture and Design... ...DCE)San Jose, CAAbout the RoleAs a System Engineer, you will own the system development of high...Permanent employmentFull timeWork at officeShift workDay shift- Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance. This...
- Causal Labs in San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting against historical physical observations. You will design techniques to improve latency and throughput, optimize...
- Vast.ai is seeking a systems engineer to scale AI inference and optimize GPU performance at our San Francisco or Los Angeles offices. You will leverage your HPC background to push the bleeding edge of AI, working with CUDA/C++ and a modern tech stack. Ideal candidates...Full time
$152k - $241.5k
...learning and eager to work on cutting-edge AI technology for safety-critical applications?... ...NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of technology, enabling high-performance AI inference solutions for automotive safety and other specialized...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff AI Inference and Acceleration Engineer. Be the first to apply!
