Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior AI Inference Engineer - Model Optimization & Deployment

Zoox Inc.

Job Description

Job Description

The Perception team is pioneering the development of a multi-modality foundation model to drive the next generation of autonomous system intelligence.


As a Model Optimization & Deployment Engineer, you will focus on bringing highly efficient, production-ready large-scale models to our on-vehicle stack. We are looking for experts with hands-on experience in compressing, accelerating, and deploying complex models (LLMs, VLMs, or FMs) for power- and thermal-constrained vehicle SOCs. You will optimize the ML models, write custom CUDA kernels, and build highly concurrent inference code to ensure real-time, deterministic execution on edge devices.

In this role, you will:
  • Optimize large-scale models (Multi-Modal Sensor Fusion models, LLMs, VLMs) using advanced quantization (PTQ, QAT), pruning, mixed-precision inference frameworks, and parameter-efficient fine-tuning (LoRA, QLoRA).
  • Architect and implement model conversion and compilation pipelines using TensorRT for edge deployment.
  • Perform rigorous parity checking, accuracy recovery, and latency benchmarking between PyTorch frameworks and compiled edge binaries.
  • Develop and optimize custom ML OPs and TensorRT Plugins with efficient CUDA kernels to minimize latency and maximize memory bandwidth on AI accelerators.
  • Write production-level, low latency, and memory-safe C++ and CUDA code for real-time inference on vehicle systems.
Qualifications:
  • Deep expertise in model quantization (PTQ, QAT) and mixed-precision inference frameworks (INT8, FP8, FP4, BF16/FP16).
  • Proven experience optimizing large-scale models (Multi-Modal Sensor Fusion models, LLMs, VLMs/VLAs) utilizing Efficient Attention mechanisms (e.g., FlashAttention, Linear Attention), KV-cache optimization (e.g., PagedAttention) and Speculative Decoding.
  • Extensive experience with model conversion/compilation pipelines (e.g., ONNX, TensorRT, torch.compile) and performing rigorous latency benchmark and model quality parity valuation.
  • Proficiency in low-level programming for AI accelerators, specifically developing and optimizing custom ML OPs and TensorRT Plugins with efficient CUDA kernel implementations.
  • Production-level C++ (14/17/20) and Python programming skills, with experience developing concurrent, memory-safe, real-time inference code for edge devices.
Bonus Qualifications:
  • Familiarity with SOTA autonomous driving perception algorithms (temporal 3D object detection, BEV, 3D Occupancy Networks) and multi-modal sensor processing (Vision, LiDAR, Radar).
  • Experience with distributed training pipelines and model/tensor parallelism (PyTorch Distributed, Ray, DeepSpeed, Megatron-LM) and runtime efficiency optimization for GPU clusters.
  • Experience with end-to-end autonomous driving paradigms (VLM/VLA models, Foundation models) and edge deployment technologies (e.g., TensorRT-LLM).

Base Salary Range

 

There are three major components to compensation for this position: salary, Amazon Restricted Stock Units (RSUs), and Zoox Stock Appreciation Rights. A sign-on bonus may be offered as part of the compensation package. The listed range applies only to the base salary. Compensation will vary based on geographic location and level. Leveling, as well as positioning within a level, is determined by a range of factors, including, but not limited to, a candidate's relevant years of experience, domain knowledge, and interview performance. The salary range listed in this posting is representative of the range of levels Zoox is considering for this position.

 

Zoox also offers a comprehensive package of benefits, including paid time off (e.g. sick leave, vacation, bereavement), unpaid time off, Zoox Stock Appreciation Rights, Amazon RSUs, health insurance, long-term care insurance, long-term and short-term disability insurance, and life insurance.

About Zoox

Zoox is developing the first ground-up, fully autonomous vehicle fleet and the supporting ecosystem required to bring this technology to market. Sitting at the intersection of robotics, machine learning, and design, Zoox aims to provide the next generation of mobility-as-a-service in urban environments. We’re looking for top talent that shares our passion and wants to be part of a fast-moving and highly execution-oriented team.

Follow us on LinkedIn

Accommodations

If you need an accommodation to participate in the application or interview process please reach out to View email address on us.fitly.work or your assigned recruiter.

A Final Note:

You do not need to match every listed expectation to apply for this position. Here at Zoox, we know that diverse perspectives foster the innovation we need to be successful, and we are committed to building a team that encompasses a variety of backgrounds, experiences, and skills.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Senior AI Inference Engineer - Model Optimization & Deployment in Washington DC vacancy
  •  ...Title: Senior Embodied AI Engineer About Us: UnitX builds...  ...inception, UnitX has deployed 1,000+ mission-critical...  ...of our foundation models. Pioneer Adaptive...  ...evaluated, diagnosed, and optimized for rapid deployment...  ...for real-time inference on robotic hardware... 
    Senior
    Full time

    Unitx

    Washington DC
    1 day ago
  •  ...government applications. Our engineers and space scientists...  ...are seeking a Senior Full‑Stack AI Application Engineer...  ...machine learning inference for real‑time analysis...  ...‑based workflows optimized for challenging physical...  ...enterprise mobile deployment or device management... 
    Senior
    Work at office
    Local area

    AST SpaceMobile

    Lanham, MD
    2 days ago
  • $188k - $275k

     ...Essential Cloud for AI™. Built for...  ...the team: The Inference team is responsible...  ...high-performance model serving capabilities...  ...for an Applied AI Engineer to help us understand...  ...driving targeted optimizations for both platform-...  ...customer-specific deployment strategies. Contribute... 
    Suggested
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Washington DC
    11 days ago
  •  ...Role : Senior AI Engineer Privacy Location : Bellevue WA...  ...apply large language models (LLMs), retrieval-augmented...  ...Build and optimize data pipelines using...  ...training, fine-tuning, and inference. Apply prompt engineering...  ...Cloud & MLOps Deploy and manage AI... 
    Senior

    Sumeru Solutions

    Washington DC
    1 day ago
  •  ...Description 540 is seeking a Senior AI/ML Engineer to support a mission-...  ...teams to develop, deploy, monitor, and scale models supporting complex defense...  ...reliable batch or real-time inference Establish model...  ...engineering, and data lineage Optimize AI/ML services and... 
    Senior
    Temporary work
    Work at office
    Local area
    Flexible hours

    540

    Arlington, VA
    11 days ago
  • $229.9k - $262.4k

    Senior Lead AI Engineer (FM Hosting, LLM Inference) Overview: At Capital One...  ..., and we build and deploy proprietary solutions that...  ...millions of customers. Our AI models and platforms empower...  ...state-of-the-art LLM optimization techniques to improve the... 
    Senior
    Full time
    Part time
    Local area

    Capital One Financial Corporation

    McLean, VA
    1 day ago
  •  ...seeking an experienced Senior AI Platform Engineer to join our Product...  ..., deliver, and optimize production-grade LLM...  ...policy enforcement, model lifecycle management...  ...CI/CD pipelines and deployment automationDefine and...  ...LiteLLM, and autoscaling inference on Kubernetes (e.g.,... 
    Senior
    Full time
    For contractors
    Work at office
    Immediate start
    Remote work

    Expression

    Washington DC
    a month ago
  • $113k - $188k

     ...position contributes to AI-enabled, data-...  ...delivers machine learning models and AI solutionsTests, deploys, and maintains AI...  ..., and feature engineering to extract valuable...  ...management and model optimization.Guides strategic AI...  ...capabilities to senior stakeholders.Active... 
    Senior
    Full time

    Maxar Technologies

    Arlington, VA
    3 days ago
  • $182.4k - $250.6k

     ...? Join the Embodied AI team at General Motors...  ...is developing and deploying machine learning solutions...  ...scenarios.As a Senior AI/ML Future Sensing Engineer in the Embodied AI...  ...improving ML and perception models that support safe...  ..., and performance optimization for perception and... 
    Senior
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Washington DC
    4 days ago
  • $158k - $197.5k

     ...practice. About the RoleThe Senior AI Defense Engineer is a technical leader...  ...What You Will Be DoingThreat Modeling & Risk Assessment - Guide...  ...engineering, training, evaluation, deployment, monitoring. Strong...  ...inversion and membership inference, overreliance/automation bias... 
    Senior
    Work experience placement
    Shift work

    Wilmer Hale

    Washington DC
    4 days ago
  •  ...partnering with Expression to find a Senior AI Software Engineer . See details below: About The Role Expression...  ...data processing, distributed inference, and resilient edge computing in...  ...retrieval architectures. Experience deploying applications using Docker and cloud‑native... 
    Senior
    For contractors
    Work at office
    Immediate start

    Hatchit Co

    Washington DC
    1 day ago
  • $165k - $180k

     ...seeking a highly skilled and motivated Sr. AI Data Engineer with a proven track record in...  ...data pipelines, automated testing, and model deployment strategies. Open-Source Integration:...  ...quality datasets for model training and inference. Required Qualifications: BA or BS... 
    Senior
    Work experience placement
    H1b
    Work at office
    Local area

    Karsun Solutions

    Washington DC
    4 days ago
  •  ...Description Job Description Senior AI Engineer Hybrid - Washington D.C....  ...of credit scoring models, generation of unique insights...  ...development, evaluation, and deployment of machine learning (ML)...  ...including back-testing, rejection inference, and performance analyses... 
    Senior
    Flexible hours

    VantageScore

    Washington DC
    more than 2 months ago
  •  ...Summary We are seeking a highly experienced Senior AI Solutions Engineer to design, develop, and deploy production-grade AI and Generative AI solutions....  ..., APIs, and enterprise systems. Evaluate and optimize model performance, latency, accuracy, and cost. Backend... 
    Senior

    Qaurs Techno Systems LLC

    Washington DC
    29 days ago
  •  ...Environment: Remote   The Data & AI Senior Engineer advances Amtrak’s mission to...  ...expertise to architect and optimize pipelines, integrations, and...  ..., API-first integrations, model lifecycle management,...  ...feature engineering, model deployment, and governance-as-code controls... 
    Senior
    Hourly pay
    Permanent employment
    Temporary work
    Work experience placement
    Interim role
    Local area
    Remote work
    Relocation
    Flexible hours

    Amtrak

    Washington DC
    more than 2 months ago
  • $176.76k - $232k

     ...The Enterprise Data & AI team is a strategic and...  ...As a Senior AI/ML Engineer, you will lead the delivery...  ...problems. You will build, deploy, scale and maintain AI...  ...challenges from setting up model training and fine-tuning...  ...design for serving AI/ML inference solutions in... 
    Senior
    Permanent employment
    Contract work
    Part time
    Work visa

    lululemon

    Washington DC
    more than 2 months ago
  • $140k - $180k

     ...business services, strategic business models and design-led user experiences. Their...  ...way their clients do business.  As a Senior AI Engineer, you will understand how AI is...  ...workshops to comprehensive enterprise-level deployments. Responsibilities Implementing AI... 
    Senior
    Full time

    Reply

    Washington DC
    1 day ago
  • $145k - $200k

     ....The Role We are a software engineering team with expertise in enabling ML models in production. We deploy AI models to run in variety of...  ...across the full stack, from inference engines, GPU scheduling to deployment...  ...our community enables us to optimize our opportunities to grow... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package

    Palantir Technologies

    Washington DC
    8 hours ago
  • $140k - $190k

     ...a highly skilled Senior AI Developer to design, build, and optimize advanced AI solutions...  ...strong hands-on engineering capabilities, deep...  ...data ingestion and model experimentation to...  ..., and deployment, and will collaborate...  ...logic.Build scalable inference services, optimize... 
    Senior
    Local area

    Steampunk

    McLean, VA
    8 hours ago
  • $141.5k - $236k

     ...career, and customer-oriented AI Engineer for a new initiative. This...  ...effort supports the rapid design, deployment, operation, and sustainment...  ...AI systems. Manage and optimize Kubernetes-based AI deployments, focusing on GPU inference optimization. Conduct core... 
    Hourly pay
    Contract work
    Temporary work
    Work experience placement
    Work at office
    Local area
    Remote work

    ManTech International Corporation

    Washington DC
    5 days ago
  • $128k - $252.5k

     ...services covering valuation modeling, cost optimization, restructuring,...  ...data scientists to AI strategists, machine...  ...specialists, and data engineers. SFL Scientific, a Deloitte...  ...is looking to add a Senior AI Engineer to their...  ...infrastructure and deployment services for our... 
    Senior
    Local area
    Visa sponsorship

    Deloitte

    McLean, VA
    2 days ago
  • $220k - $350k

     ...enabling human life on Mars.SR. AI ENGINEER, SPECIAL PROGRAMS - TOP...  ...own strategy, execution, and deployment of mission-critical AI solutions...  ...make current AI and future models maximally useful for...  ...RESPONSIBILITIES:Design, build, and optimize integrations between AI... 
    Senior
    Permanent employment
    Temporary work
    For contractors
    Shift work
    Weekend work

    SpaceX

    Washington DC
    3 days ago
  • $190k - $230k

     ...Platform unites agentic AI solutions with the...  ...and accountability. Senior Software Engineer, AI Platform Location...  ...horizon tasks, LLM-based models into the HackerOne...  ...Familiarity with model deployment workflows, evaluation...  ..., or LLM optimization techniques Experience... 
    Senior
    Apprenticeship
    Work at office
    Local area
    Remote work
    Flexible hours
    Shift work
    1 day per week

    hackerone

    Washington DC
    4 days ago
  •  ...Team Join the engineering teams that bring...  ...seek to learn from deployment and distribute the benefits of AI, while ensuring...  ...capabilities to optimizing how we serve inference in unique, high-...  ...on their latest models, developing...  ...Staff . We use Senior Staff externally... 
    Senior
    Full time

    OpenAI

    Washington DC
    1 day ago
  •  ...build and ship language-model-powered systems that...  ...-training models to deploying reliable inference and retrieval systems...  ...with product and engineering to translate real operational...  ...into high-performing AI features, operating...  .... Develop and optimize model serving infrastructure... 
    Full time
    Work at office
    Flexible hours

    Twenty Inc.

    Washington DC
    1 day ago
  •  ...Description:DataRobot delivers AI that maximizes impact and...  ...Organization:The Pre-Sales AI Solutions Engineer for the Federal Sector plays a...  ...alignment.Solution Design & Deployment Readiness: Define key...  ...requirements, recommend deployment models that meet FedRAMP, DoD IL, and... 
    Senior
    Full time
    Local area
    Remote work
    Worldwide
    Flexible hours
    Shift work

    DataRobot

    Washington DC
    4 days ago
  •  ...highly skilled AI Systems Engineer who will be part...  ...solutions.The Senior AI Systems Engineer...  ...operational deployment. Responsibilities...  ...interfaces, or model driven service frameworks...  ...with optimization techniques such...  ...validation layers, or inference acceleration.... 

    Tria Federal

    Arlington, VA
    2 days ago
  •  ...Job Title: Senior SAS/Python Model Validation and Modernization...  ...technical operations and deployments. Position...  ...identify opportunities for optimization and improvement....  ...Statistics, Mathematics, Engineering, Information...  ...technologies, particularly AI  Strategic... 
    Senior
    Full time
    Work at office
    Remote work

    BLN24

    McLean, VA
    more than 2 months ago
  •  ...is seeking an experienced AI/ML Engineer to design, optimize, and evaluate machine...  ...prioritization, distributed inference, and decision support while...  ..., and production deployment of AI systems. Security...  ...deploy open-weight foundation models appropriate for resource-... 
    Full time
    For contractors
    Work at office
    Immediate start
    Remote work

    Expression

    Washington DC
    a month ago
  •  ...We build AI agents that actually work in enterprise...  ..., not demos. We need engineer's who can own the entire...  ...integrations that are model-agnostic and built to last. You'll be deployed on client engagements as...  ..., and cost/latency optimization across long-running agent... 
    Full time
    Temporary work

    Trilagen

    Bethesda, MD
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior AI Inference Engineer - Model Optimization & Deployment. Be the first to apply!