Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization

$169.78k - $351k
Full-time

DiDi Labs

About the Company

DiDi's autonomous driving unit was established in 2016 with the mission of developing Level 4 autonomous driving (AD) technology to make transportation safer and more efficient. In August 2019, the unit became an independent company, DiDi Autonomous Driving, dedicated to advanced AD R&D, product application, and business expansion. We believe integrating AD technology into a shared-mobility fleet will generate immense social value. By leveraging DiDi's specialized technology, operational expertise, and integrated ecosystem, we are positioned to build and operate a highly efficient, user-oriented autonomous fleet.

About The Role

We are seeking an experienced and mission-driven Senior/Sr. Staff AI Infrastructure Engineer , Inference & Optimization to lead the performance tuning, deployment, and resource scheduling of cutting-edge AI models across on-vehicle and cloud infrastructure. In this role, you will design high-efficiency inference pipelines, build system-level stability frameworks, and optimize hardware execution to ensure ultra-low latency and rock-solid operational reliability. You will act as a technical leader in AI infrastructure, accelerating model iteration and bridging the gap between frontier deep learning algorithms and real-time autonomous systems.

Responsibilities

  • Own the deployment, optimization, and resource scheduling of vehicle-side AI models, ensuring high efficiency, low latency, and robust execution within embedded constraints.

  • Lead vehicle-side system stability initiatives, conducting independent root-cause analysis and driving resolution for complex, system-level performance bottlenecks and runtime anomalies.

  • Architect and scale service-oriented deployment environments for Large Language Models (LLMs) and foundational models to support offline simulation, automated annotation, and rapid model validation.

  • Track and evaluate cutting-edge industry methodologies, continuously integrating advanced optimization toolchains, quantization techniques, and execution engines.

  • Establish system-level profiling and telemetry frameworks using CUDA tools to monitor, analyze, and maximize hardware utilization across target GPU architectures.

  • Collaborate cross-functionally with Autonomous Driving Perception/Prediction, Cloud Infrastructure, and Safety teams to enable rapid algorithm iteration and scalable vehicle deployment.

Qualifications

  • Master’s or higher degree in Computer Science, Software Engineering, Systems Engineering, or a closely related technical field.

  • 3-8+ years of industry experience in high-performance computing, AI infrastructure, model optimization, or embedded deployment.

  • Strong proficiency in C++ and Python, with solid expertise in parallel programming (CUDA, OpenMP) and low-level system profiling tools.

  • Deep familiarity with mainstream inference engines (e.g., TensorRT, ONNX Runtime) and specialized LLM inference/serving frameworks (e.g., vLLM, SGLang, TensorRT-LLM).

  • Practical understanding of modern GPU hardware architectures (e.g., NVIDIA Hopper, Thor) and memory bandwidth management.

  • Demonstrated ability to diagnose complex software-hardware integration issues and drive scalable, production-grade solutions.

Preferred Qualifications

  • Hands-on experience optimizing and deploying AI models on the NVIDIA Thor platform, including hardware resource scheduling and acceleration.

  • Proven track record of serving large foundation models (e.g., LLaMA, Qwen, GPT) in production or high-throughput cloud pipelines using frameworks like vLLM, SGLang, TGI, or LightLLM.

  • Background in deep learning training frameworks (PyTorch) and practical experience with model quantization (INT8/FP8/AWQ), kernel fusion, or graph compilation.

  • Experience deploying real-time, high-availability AI workloads in autonomous vehicles, robotics, or edge devices.

The base salary range for this full-time position is $169,783 - $351,000 annually in addition to bonus, equity and benefits. Our salary ranges are determined by role, level, and location. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training.

I acknowledge that prior to submitting this application, I have read and accepted the Privacy Notice for California Residents which is available on
Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization in San Jose, CA vacancy
  • $203.45k - $344.3k

     ...innovation, integrating advanced AI and autonomous driving...  ....As a core member of our AI Infrastructure team, you will be responsible...  ...and driving iterative model optimization.Data Support for Production...  ...Computer Science, Software Engineering, Artificial Intelligence, or... 
    Senior
    Full time
    Overseas

    XPENG Motors

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...NVIDIA's DGX Cloud AI Efficiency Team...  ...contributing to the infrastructure that powers our...  ...tools for optimizing efficiency and resiliency...  ..., post-training, inference. Our objective is...  ...software engineer to join our team....  ...AI systems.As a senior DGX Cloud AI Infrastructure... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $229.9k - $262.4k

     ...Senior Lead AI Engineer (GenAI Platform Services) Overview...  ...investments in technology infrastructure and world‑class...  ...large language model inference, similarity search,...  ...state‑of‑the‑art LLM optimization techniques to improve...  ...,900 - $262,400 for Sr. Lead AI Engineer... 
    Senior
    Local area

    Comfort Systems USA

    San Jose, CA
    1 day ago
  • $229.9k - $262.4k

    Senior Lead AI Engineer (Gen AI Platform Services, Agentic AI) Overview...  ...in technology infrastructure and world-class...  ...large language model inference, similarity search,...  ...state-of-the-art LLM optimization techniques to improve...  ...,900 - $262,400 for Sr. Lead AI Engineer... 
    Senior
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    4 days ago
  • $152k - $241.5k

     ...deep learning ignited modern AI — the next era of computing —...  ...AI & Deep Learning Compiler Engineer. NVIDIA is hiring software engineers...  ...the backbone of NVIDIA’s inference engine, spanning across data...  ...and developing compiler optimization algorithms.Collaborating with... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $229.9k - $262.4k

     ...Overview Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform) Overview: At Capital One, we are creating responsible and...  ...customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience... 
    Senior
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    7 days ago
  • $124.5k - $272k

     ...networking for the cloud and AI era. We secure and...  ...One platform, its Zero Trust Engine, and the powerful NewEdge network...  ...Positions are available at Senior Staff and above. Candidates are...  ...Learning Scientist, you own the inference and optimization layer that makes AI in... 
    Senior

    Netskope

    Santa Clara, CA
    3 days ago
  • $183.6k - $297k

     ...and Inclusion. We weave AI into the fabric of everything...  ...AI into cybersecurity infrastructure — building intelligent...  ...As a Principal Software Engineer, you will own the...  ...or similar techniques to optimize LLM performance and reduce inference costSolid skills in multi... 
    Senior
    Full time
    Work at office

    Palo Alto Networks

    Santa Clara, CA
    2 days ago
  • $203.45k - $344.3k

     ...innovation, integrating advanced AI and autonomous driving...  ...perspective, making it the core infrastructure that supports weekly model...  ...and closed-loop engineering system.Job ResponsbilitiesResponsible...  ...ResponsbilitiesResponsible for the design and optimization of the vehicle-cloud... 
    Senior
    Full time
    Temporary work
    Work experience placement

    XPENG Motors

    Santa Clara, CA
    3 days ago
  •  ...computing experiences—from AI and data centers, to...  ...AI / ML Platform Engineers to build the platform...  ...role focuses on the infrastructure and platform systems...  ...distributed training and inference, experiment tracking,...  ...across kernel optimization, RTL/PPA optimization... 
    Senior

    AMD

    Santa Clara, CA
    3 days ago
  •  ...technology company for AI and Bitcoin mining infrastructure. Bitdeer is...  ...We are seeking a Senior AI Storage Infrastructure Engineer to build the critical...  ...AI model training and inference are profoundly I/O intensive...  ...distributed training. Optimize IOPS, throughput, and... 
    Senior
    Remote job
    Full time
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    14 days ago
  •  ...computing experiences—from AI and data centers, to...  ...career. TTHE ROLE: Senior level engineer who will be...  ...strategy, architecture, optimization and tooling to...  ...training and Distributed Inference Performance on AMD GPU...  ..., tools and infrastructure for performance estimation... 
    Senior

    AMD

    San Jose, CA
    1 day ago
  • $229.9k - $262.4k

     ...Sr. Lead AI Engineer (FM Hosting, LLM Inference) Overview: At Capital One, we are creating responsible and...  .... Our investments in technology infrastructure and world-class talent — along with...  ...introduce state-of-the-art LLM optimization techniques to improve the... 
    Senior
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    2 days ago
  • $193.93k - $352.29k

     ...and profound opportunity for AI to drive positive change in the...  ....About the RoleThe ML Infrastructure team is responsible for building...  ...to building to deploying the optimized models on Nuro’s fleet of self...  ...validation.Maintain an in-house ML inference platform to serve large... 
    Senior
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    4 days ago
  •  ...technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed...  ...and motivated Cloud Senior DevOps Engineer to join our AI Cloud team...  ..., driving automation, optimizing cloud-native...  ....g., vLLM, TGI, Triton Inference Server), and experience... 
    Senior
    Remote job
    Full time
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    14 days ago
  • $274k - $300k

     ...Description Job Description Saviynt's AI-powered identity platform manages...  ..., please visit  AI Platform Engineer – Training & Inference Saviynt's AI-powered identity platform...  ...cloud LLMs • Build RL training infrastructure: define Flyte workflows for RL... 

    Saviynt

    Milpitas, CA
    4 days ago
  • SOMERSET STAFFING seeks a network professional to operate and optimize a company’s internal data communications systems in Milpitas, CA. The role covers LAN/WAN management, network design, and configuration, with responsibilities that include troubleshooting, vendor coordination... 
    Senior

    SOMERSET STAFFING

    Milpitas, CA
    4 days ago
  •  ...that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems....  ...advance your career. THE ROLEWe are seeking a Principal GenAI Inference Optimization Engineer to join our Models and Applications team. This role... 

    AMD

    San Jose, CA
    1 day ago
  •  ...computing experiences—from AI and data centers, to...  ...a DevOps / Platform Engineer to join our team...  ...large-scale GPU compute infrastructure that powers AI and ML...  ...effectively and work optimally with their peers within...  ...development pods, CI pipelines, inference services, benchmarking... 

    AMD

    San Jose, CA
    4 days ago
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference...  ...-measured kernel benchmarking infrastructure, model-level performance...  ...agentic kernel optimization: applying AI-driven analysis to diagnose performance... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $224k - $356.5k

    We are now looking for a Senior System Software Engineer to work on Dynamo. NVIDIA is...  ...GPUs to power a revolution in AI, enabling breakthroughs in...  ...building Generative AI inference platform to make design and...  ...across available resources; optimize throughput under latency constraints... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

     ...deep learning ignited modern AI — the next era of...  ...company”.NVIDIA is hiring a Senior AI Compiler Engineer. GPUs are driving rapid progress...  ...that powers NVIDIA’s inference engine end to end, with a...  ...graph representations and optimizations for future GPU architectures... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    14 hours ago
  • $175k - $265k

     ...potential of generative AI to power the transformation...  ...days per week.The role: Sr. Staff, ML Researcher - LLM Algorithmic...  ...that will be used to optimize large language model inference on DNN accelerators we...  ..., ML researchers, and ML engineers who create and apply advanced... 
    Senior
    3 days per week

    d-Matrix

    Santa Clara, CA
    3 days ago
  • $190k - $260k

     ...an artificial intelligence (AI) powered technology stack purpose...  ...world models - depends on infrastructure that turns thousands of hours...  ...throughput. We are looking for engineers who make model training fast:...  ...keep GPUs saturatedBuild and optimize distributed training... 
    Senior
    Temporary work
    Work at office
    Visa sponsorship

    Kodiak Robotics

    Mountain View, CA
    2 days ago
  • $152k - $241.5k

    We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join...  ...to pinpoint and address infrastructure and application deficiencies...  ...resolving, training & inference performance end to endDebugging and optimization experience with NSight... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

    We are looking for outstanding Senior High Performance AI Engineers to build the next generation of agentic...  ...systems that can reason about, generate, optimize, and operate across NVIDIA's...  ...and hardware stack, from models and inference through compilers, runtimes, libraries... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

     ...into the unlimited potential of AI to define the next era of...  ...AI & Deep Learning Compiler Engineer. NVIDIA is hiring software engineers...  ...the backbone of NVIDIA’s inference engine, spanning across data...  ...networks and developing compiler optimization algorithms.Strong programming... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $255.65k - $299k

     ...available for this positionWhat you can expect:As a Senior AI Software Engineer, you will collaborate to design, implement, and optimize AI algorithms and software applications. You will ensure AI training, inference, deployment, and operation are functional, reliable,... 
    Senior
    Full time
    Work at office
    Remote work

    Zoom

    San Jose, CA
    1 day ago
  • $184k - $287.5k

     ...tapping into the unlimited potential of AI to define the next era of computing....  ...Co-Design Group (SCG) is seeking Senior AI Platform Engineers. They will set the technical direction...  ...platforms at the intersection of ML infrastructure and large-scale systems, this is your... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

    We are looking for a software engineer with a strong background in...  ...performance at the intersection of AI, high-performance computing,...  ...experts to analyze, optimize, and scale complex AI and HPC...  ...from the crowd: Experience with inference optimization techniques and deploying... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization. Be the first to apply!