Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff AI Inference and Acceleration Engineer

$180k - $275k
Full-time

Figure

Figure is an AI robotics company developing autonomous general-purpose humanoid robots. The goal of the company is to ship humanoid robots with human level intelligence. Its robots are engineered to perform a variety of tasks in the home and commercial markets. Figure is headquartered in San Jose, CA.

We are looking for a Staff AI Inference & Acceleration Engineer to join the Platform Software team and own the on-board inference architecture for Figure’s humanoid robots. You will be the technical authority on how AI workloads are mapped, optimized, and executed across the robot’s compute hardware — driving down power consumption and cost while meeting the strict latency and reliability demands of a real-time autonomous system.

Responsibilities

  • Own the on-board inference architecture — mapping models to available accelerators (NPU, GPU, DSP, CPU) based on latency, power, and memory budgets.
  • Partition inference workloads across heterogeneous compute resources, balancing real-time performance with power and thermal constraints.
  • Define and maintain a system-level compute budget across all inference tasks running on the robot.
  • Evaluate next-generation acceleration hardware and contribute to the definition of future compute platform requirements.
  • Optimize inference toolchains end-to-end — from model export through runtime execution — for target hardware.
  • Apply quantization (INT8, INT4, mixed-precision), pruning, operator fusion, and other compression techniques to reduce compute, memory, and power footprint.
  • Profile inference pipelines to identify and eliminate bottlenecks in latency, memory bandwidth, and power consumption.
  • Optimize kernel scheduling, memory layout, and data movement across the compute hierarchy.
  • Partner closely with the AI/ML team to define model architecture constraints that are hardware-friendly from the outset.
  • Work with the Platform Software team on runtime integration, scheduling, and power management.
  • Engage with silicon vendors and research teams to track the accelerator landscape and influence hardware roadmaps.

Requirements

  • M.S. or Ph.D. in Computer Engineering, Electrical Engineering, Computer Science, or a related field — or equivalent industry experience.
  • At least 8 years of industry experience in hardware acceleration, ML systems, or compute architecture.
  • Deep understanding of AI/ML inference — model formats (ONNX, TFLite, etc.), inference runtimes, and deployment pipelines.
  • Hands-on experience optimizing models for edge or embedded hardware using quantization, pruning, and operator-level tuning.
  • Strong understanding of computer architecture — memory hierarchies, data movement, and heterogeneous compute.
  • Experience profiling and benchmarking inference workloads across CPU, GPU, NPU, DSP.
  • Familiarity with low-level toolchains and compilation frameworks (e.g. TVM, MLIR, TensorRT, Torch, SNPE/QNN, JAX, CUDA, ROCm).
  • Solid software engineering skills in C++ and Python.
  • Strong cross-functional communication skills — able to work effectively across hardware, software, and AI/ML teams.

Bonus Qualifications

  • Knowledge of real-time operating constraints and their impact on inference scheduling.
  • Track record of co-designing model architectures with ML teams to meet hardware constraints.

The US base salary range for this full-time position is between $180,000 - $275,000 annually.

The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components/benefits depending on the specific role. This information will be shared if an employment offer is extended.
Vacancy posted 6 days ago
Similar jobs that could be interesting for youBased on the Staff AI Inference and Acceleration Engineer in California vacancy
  • RadixArk is seeking a Member of Technical Staff to accelerate LLM inference and training on modern GPUs. You will profile, optimize, and extend SGLang...  ...contributing to the roadmap and ecosystem integrations that power scalable AI inference and training. #J-18808-Ljbffr RadixArk
    Suggested

    RadixArk

    Palo Alto, CA
    2 days ago
  • $152k - $241.5k

     ...deep learning ignited modern AI — the next era of computing —...  ...seeking top-tier AI Compiler Engineers to drive innovation within our...  ...for AI workloads (both inference and training) and successfully...  ...on CPU, GPU, and/or custom AI accelerator architectures.LLM Knowledge:... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • RadixArk is seeking a Member of Technical Staff, Developer Technology (DevTech) to accelerate LLM inference and training on modern GPU hardware. Our engines power trillions of tokens daily and...  ...researchers and partners to scale AI workloads. You will profile, optimize,... 
    Suggested

    RadixArk

    Palo Alto, CA
    1 day ago
  • $184k - $287.5k

    We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency...  ..., and performance teams to push the frontier of accelerated computing for AI.What you’ll be doing:Contribute... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...Inc. is seeking a Member of Technical Staff (Hardware) to design and run AI hardware benchmarks for GPUs, TPUs...  ...standards. The role requires 3+ years in AI accelerators, strong Python/data-analysis skills, and deep knowledge of inference economics. #J-18808-Ljbffr... 
    Suggested

    Artificial Analysis, Inc.

    San Francisco, CA
    4 days ago
  •  ...and sustainability, design and engineering, ambition and integrity. In...  ...mobility.Lucid Motors is seeking a Staff QA and Test Infrastructure...  ...Qualifications – Agentic AI Tools & Intelligent AutomationLucid...  ...agentic AI systems to accelerate development, automate complex... 

    Lucid Motors

    Newark, CA
    10 hours ago
  • $122.6k - $185k

     ...backbone of Generative AI cloud at AWS? Do you...  ...cloud for AI training and inference? Want to do industry...  ...you. The AWS Hardware Engineering team creates server...  ...future/new designs for AWS Accelerated server solutions for...  ..., supervisors, and staff; adhere to standards of... 
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $142.8k - $274.8k

     ...Hardware, and Infrastructure Engineering (SCHIE) is the team behind Microsoft...  ...organization is developing AI-native silicon and hyperscale...  ...-leading AI training and inference. The Platform Systems...  ...team is seeking a Principal AI Accelerator Tools Development Engineer to... 
    Ongoing contract
    Work at office
    Local area
    Worldwide
    3 days per week

    Microsoft

    Mountain View, CA
    10 hours ago
  • MakerMaker.AI is looking for a Senior Machine Learning Systems Engineer in San Francisco. In this role, you will build and operate production inference systems, optimizing for performance and reliability...  ...have strong knowledge in GPU-accelerated inference. Excellent... 

    MakerMaker.AI

    San Francisco, CA
    2 days ago
  •  ...is seeking a Member of Technical Staff to own coverage of the serverless inference landscape. You will benchmark endpoints...  ...and keep us ahead in frontier AI benchmarking. You will analyze metrics...  ...frontiers, collaborating with engineers and industry leaders. #J-18808-Ljbffr... 

    Artificial Analysis, Inc.

    San Francisco, CA
    4 days ago
  • Crusoe in San Francisco is seeking a Staff Technical Program Manager to lead the Managed Inference platform team. You will ensure end-to-end program delivery for LLM workloads, driving innovation in AI infrastructure powered by clean energy. The ideal candidate has extensive... 

    Crusoe

    San Francisco, CA
    4 days ago
  • An innovative company is seeking a talented software engineer to join their dynamic Inference team. This role involves designing and implementing infrastructure...  ...researchers and product teams to push the boundaries of AI technology, ensuring reliable production services. If... 

    Jobleads-US

    San Francisco, CA
    5 days ago
  • Anyscale is seeking a Distributed LLM Inference Engineer in San Francisco, California. This pivotal role involves pushing the boundaries of performance for ML inference at scale. You'll work closely with product teams to deliver end-to-end solutions while leveraging open... 

    Anyscale

    San Francisco, CA
    5 days ago
  • $200k - $360k

     ...professional to own MLIR dialect design and lowering passes for their AI accelerator. You'll work closely with chip-design and software teams....  ...knowledge of tensor compilation, and 5+ years of compiler engineering experience. The role includes a competitive compensation... 

    DensityAI

    Mountain View, CA
    2 days ago
  •  ...seeks an experienced Deep Learning Software Engineer, TensorRT Performance, to analyze and enhance the performance of NVIDIA’s inference ecosystem, including TensorRT, TensorRT‑...  ...scheduling, and collaborate with teams across AI, automotive, robotics and vision domains to... 

    NVIDIA Corporation

    Santa Clara, CA
    5 days ago
  • $260k - $320k

    DensityAI is seeking an expert to develop compute kernels for a specialized AI accelerator in Mountain View, California. Your role will focus on writing performance-critical kernels while collaborating with architecture and compiler teams to enhance silicon design. Qualifications... 

    DensityAI

    Mountain View, CA
    2 days ago
  • Acceler8 Talent is hiring a Kernel Engineer to develop and optimize compute kernels for a custom AI accelerator in Mountain View, CA. You will collaborate with architecture, compiler, simulation, and systems teams to define the kernel programming model and validate performance... 

    Acceler8 Talent

    Mountain View, CA
    2 days ago
  • A leading AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for designing high-performance GPU kernels and using advanced techniques to boost computation efficiency. Ideal... 

    Baseten

    San Francisco, CA
    5 days ago
  • NVIDIA in Santa Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design and optimize kernels, build...  ...for LLM serving engines, and contribute to accelerators and runtimes that power large language models and AI... 

    NVIDIA

    Santa Clara, CA
    3 days ago
  • $262k - $365k

     ...rigorous code review, ensuring accuracy and engineering best practices.Serve as a Senior TLM to...  ..., "speed-of-light" delivery of modern Accelerator appliances into Google Data Centers....  ...software teams by leveraging the latest AI tools to transform methodologies and accelerate... 
    Worldwide

    Google

    Sunnyvale, CA
    4 days ago
  • NVIDIA Corporation in Santa Clara, CA is seeking outstanding AI systems engineers to develop groundbreaking inference technologies for the hardware-accelerated stack. You will create libraries, code generators, and GPU kernel innovations for LLM workloads. Join a team... 

    NVIDIA Corporation

    Santa Clara, CA
    5 days ago
  • NVIDIA is seeking outstanding AI systems engineers in Santa Clara to advance the inference software stack. You will build libraries, code generators, and GPU kernels...  ...for LLM serving engines and JIT compilers to accelerate large language models. The role emphasizes Python... 

    Segment (Twilio)

    Santa Clara, CA
    5 days ago
  • $150k - $275k

    A cutting-edge tech company in San Jose is seeking a Supercomputing Engineer to ensure the reliability of its inference servers. This role involves designing and executing test suites, analyzing performance, and collaborating with engineering teams. Ideal candidates will... 

    Jobleads-US

    San Jose, CA
    5 days ago
  • Google in Sunnyvale, CA is seeking a Senior Staff Co-Design Engineer on the TPU Chip Architecture team to shape AI/ML hardware acceleration. You will drive TPU architecture and verification efforts for next-generation training and serving workloads, collaborating with research... 

    Socket.dev

    Sunnyvale, CA
    3 days ago
  •  ...is seeking a Senior hardware architect to design and implement accelerators for codesigned systems. You will own subsystems from requirements...  ...with ML model design and numerics. Google DeepMind values safety and ethics in AI development. #J-18808-Ljbffr Google DeepMind

    Google DeepMind

    Mountain View, CA
    3 days ago
  • $150.98k - $218.62k

     ...Intelligent Edge. ADI combines analog, digital, AI, and software technologies into solutions...  .... Learn more at and on LinkedIn and X.Staff System Architecture and Design...  ...DCE)San Jose, CAAbout the RoleAs a System Engineer, you will own the system development of high... 
    Permanent employment
    Full time
    Work at office
    Shift work
    Day shift

    Analog Devices

    San Jose, CA
    4 days ago
  • Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance. This... 

    Vast.ai Inc.

    San Francisco, CA
    3 days ago
  • Causal Labs in San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting against historical physical observations. You will design techniques to improve latency and throughput, optimize... 

    Causal Labs

    San Francisco, CA
    2 days ago
  • Vast.ai is seeking a systems engineer to scale AI inference and optimize GPU performance at our San Francisco or Los Angeles offices. You will leverage your HPC background to push the bleeding edge of AI, working with CUDA/C++ and a modern tech stack. Ideal candidates... 
    Full time

    Vast.ai

    San Francisco, CA
    3 days ago
  • $152k - $241.5k

     ...learning and eager to work on cutting-edge AI technology for safety-critical applications?...  ...NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of technology, enabling high-performance AI inference solutions for automotive safety and other specialized... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff AI Inference and Acceleration Engineer. Be the first to apply!