Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior AI Infrastructure Engineer

$180k - $240k

Gatik AI

Senior AI Infrastructure EngineerSanta Clara, CAAbout the RoleGatik, the leader in autonomous middle-mile logistics, is revolutionizing the B2B supply chain with its autonomous transportation-as-a-service (ATaaS) solution and prioritizing safe, consistent deliveries while streamlining freight movement by reducing congestion. The company focuses on short-haul, B2B logistics for Fortune 500 retailers and in 2021 launched the world's first fully driverless commercial transportation service with Walmart. Gatik's Class 3-7 autonomous trucks are commercially deployed across major markets, including Texas, Arkansas, and Ontario, Canada, driving innovation in freight transportation.The company's proprietary Level 4 autonomous technology, Gatik Carrier™, is custom-built to transport freight safely and efficiently between pick-up and drop-off locations on the middle mile. With robust capabilities in both highway and urban environments, Gatik Carrier™ serves as an all-encompassing solution that integrates advanced software and hardware powering the fleet, facilitating effortless integration into customers' logistics operations.We are seeking a Senior AI Infrastructure Engineer to design, build, and scale the high-performance AI platform powering our autonomous driving models. While researchers focus on developing perception, planning, and world models, you will be responsible for the underlying infrastructure that enables distributed training, experiment tracking, and seamless model deployment. You will bridge the gap between research and production, ensuring our AI stack is scalable, resilient, and highly efficient.This role is onsite 5 days a week at our Santa Clara, CA office!What You'll DoDistributed Training & ML Systems SupportScale Research Workloads: Enable researchers to scale complex models (VLA, World Models) across multi-node setups using PyTorch Distributed, and Ray Train.Performance Optimization: Architect and optimize multi-GPU setups, ensuring efficient model parallelism and data parallelism techniques across H100/A100 clusters.Networking & Hardware Tuning: Optimize low-level communication (e.g., NCCL tuning, InfiniBand, or RoCE v2) to minimize latency for 3D Gaussian Splatting (3DGS) and large-scale training.Intelligent Resource Scheduling: Optimize hardware utilization and cost-efficiency through Kubernetes-native GPU scheduling (NVIDIA GPU Operator, KubeFlow).Inference Performance Engineering: Deploy and scale optimized model artifacts using TensorRT, ONNX Runtime, and Triton Inference Server, fine-tuning pipelines for both real-time and batch processingAgentic Infrastructure & AutomationSelf-Healing AI Infrastructure: Architect and deploy Autonomous AI Agents (LangGraph, CrewAI, or AutoGen) to monitor GPU cluster health, enabling automated real-time triage of hardware failures and NCCL timeouts.Agentic DevOps & CI/CD: Develop agent-driven automation, such as Agentic PR Reviewers for infrastructure code and AI agents that proactively suggest model-specific Kubernetes resource optimizations.Agentic Data Curation: Support researchers in building "Data Machines" where AI agents autonomously curate, label, and verify high-priority edge cases from raw data.Model Management & Lifecycle (MLOps)Automated Lifecycle Management: Design and maintain ML infrastructure leveraging MLFlow, Argo Workflows, and Kubernetes to automate the end-to-end model lifecycle.Experiment & Model Tracking: Integrate feature stores and experiment tracking systems to provide a robust system of record for every model iteration.Deployment Strategies: Implement robust serving mechanisms, including A/B testing, shadow deployments, and rollback mechanismsCloud-Native Foundations & Data IntegrationInfrastructure as Code: Drive the "Everything as Code" philosophy using Terraform and Helm.Data Pipelines: Collaborate with data teams to scale ETL pipelines using Apache Airflow, Kafka, and Spark for large-scale dataset management.Integrated Data Factories: Collaborate with data engineering teams to scale high-bandwidth ETL pipelines using Apache Airflow, Kafka, and Spark, ensuring seamless data flow from raw sensor logs to optimized storage in S3, GCS, or Delta LakeMonitoring & ObservabilitySystem Metrics: Define and track key ML system metrics, including training convergence, latency, throughput, and drift detection.Infrastructure Health: Maintain deep visibility into platform health using Prometheus, Grafana, OpenTelemetry, and ELK Stack.Deep Stack Observability: Develop comprehensive monitoring using Prometheus, Grafana, and OpenTelemetry to track low-level infrastructure health alongside high-level ML metrics like training convergence and throughput.AI-Specific Metrics & Drift: Define and monitor critical ML system KPIs, including model latency, inference throughput, and feature drift detectionWhat We're Looking ForExperience: 5+ years in ML infrastructure, MLOps, or DevOps supporting high-scale compute environments.ML Expertise: Deep understanding of multi-GPU training strategies (FSDP, DeepSpeed, Ray Train) and high-performance networking (NCCL, InfiniBand).Infrastructure Automation: Mastery of Kubernetes, Terraform, and Helm, with a focus on GPU-native orchestration.AI Agent Frameworks: Proven experience building or supporting Agentic Workflows for infrastructure or data automation (e.g., using LLMs to drive DevOps tasks).Platform Mastery: Expertise in MLFlow, Argo Workflows, and Kubernetes.Containerization: Strong experience with Docker, Kubernetes, and Helm.Data & CI/CD: Proficiency in Apache Airflow, Kafka, Spark, and GitOps automation.Core Skills: Proficiency in Python and Bash; experience with Go or Rust is a plusBonus QualificationsAdvanced AI Protocols: Familiarity with the Model Context Protocol (MCP) to standardize how AI agents interact with internal databases and orchestration APIs.Hybrid & Physical AI: Experience in hybrid cloud and on-prem GPU cluster management for Physical AI workloads (e.g., 3DGS, World Models).Agentic Observability: Experience utilizing LLMs for semantic monitoring and log analysis to detect complex distributed system failures that traditional threshold-based alerts miss.Salary Ranges - $180,000- $240,000More About GatikFounded in 2017 by experts in autonomous vehicle technology, Gatik has rapidly expanded its presence to Mountain View, Dallas-Fort Worth, Arkansas, and Toronto. As the first and only company to achieve fully driverless middle-mile commercial deliveries, Gatik holds a unique and defensible position in the AV industry, with a clear trajectory toward sustainable growth and profitability.We have delivered complete, proprietary AV technology - an integration of software and hardware - to enable earlier successes for our clients in constrained Level 4 autonomy. By choosing the middle mile – with defined point-to-point delivery, we have simplified some of the more complex AV challenges, enabling us to achieve full autonomy ahead of competitors. Given extensive knowledge of Gatik's well-defined, fixed route ODDs and hybrid architecture, we are able to hyper-optimize our models with exponentially less data, establish gate-keeping mechanisms to maintain explainability, and ensure continued safety of the system for unmanned operations.Visit us at Gatik for more company information and Careers at Gatik for more open roles.

Vacancy posted 14 hours ago
Similar jobs that could be interesting for youBased on the Senior AI Infrastructure Engineer in Santa Clara, CA vacancy
  • $286.2k - $326.7k

     ...Senior Director, AI Engineering -Agentic AI PlatformAt Capital One, we are creating responsible and reliable AI systems, changing banking for...  ...personalized customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience... 
    Senior
    Full time
    Part time
    Remote work

    Capital One

    San Jose, CA
    2 days ago
  • $229.9k - $262.4k

    Senior Lead AI Engineer (GenAI Platform Services) Overview At Capital One, we are creating responsible and reliable AI systems, changing banking...  ...customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience... 
    Senior
    Full time
    Part time
    Local area

    Capital One Financial Corp

    San Jose, CA
    3 days ago
  •  ...Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing...  ...visit Position Overview We are seeking a Senior AI Storage Infrastructure Engineer to build the critical data-delivery fabric of our... 
    Senior
    Remote job
    Full time
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    11 days ago
  • $229.9k - $262.4k

     ...Sr. Lead AI Engineer (Gen AI Platform Services) Overview At Capital One, we are creating responsible and reliable AI systems,...  ...personalized customer experiences. Our investments in technology infrastructure and world‑class talent — along with our deep experience in... 
    Senior
    Local area

    Capital One National Association

    San Jose, CA
    14 hours ago
  •  ...Nubank is seeking a Staff Software Engineer to lead the technical architecture for a private AI banking platform. You’ll own from Flutter mobile apps to distributed backends, enabling real-time AI-driven personalization at scale across web and mobile. You will collaborate... 
    Senior

    Jobleads-US

    Palo Alto, CA
    5 days ago
  • $184k - $287.5k

    Joining NVIDIA's DGX Cloud AI Efficiency Team means contributing to the infrastructure that powers our innovative AI research...  ...an AI infrastructure software engineer to join our team. You'll be instrumental...  ...availability of AI systems.As a senior DGX Cloud AI Infrastructure... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...JPMorganChase is seeking a Senior Director of Software Engineering to lead Data and AI platforms within the Commercial and Investment Bank Operations group. You will oversee multiple departments, set technical strategy, and drive enterprise-wide AI-enabled engineering... 
    Senior

    Jobleads-US

    Palo Alto, CA
    5 days ago
  • $184k - $287.5k

     ...tapping into the unlimited potential of AI to define the next era of computing....  ...Co-Design Group (SCG) is seeking Senior AI Platform Engineers. They will set the technical direction...  ...platforms at the intersection of ML infrastructure and large-scale systems, this is your... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing comprehensive...  ...We are seeking a highly skilled and motivated Cloud Senior DevOps Engineer to join our AI Cloud team. In this high-impact role,... 
    Senior
    Remote job
    Full time
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    11 days ago
  • $178k - $321k

     ...Audit function has an early but working AI-native capability: a multi-agent platform...  ...setting, and the governed data and AI infrastructure everything else depends on. We hire on demonstrated...  ...just implement it. This is a two-person engineering team: you deploy, debug, and hotfix your... 
    Senior

    OKX

    San Jose, CA
    6 days ago
  • $200k - $400k

     ...intelligent machines at scale. At Scout AI, we’re developing Fury, the first...  .... The Role We're looking for a Senior or Staff AI Engineer to join the Fury Orchestration Team with...  ...PyTorch, and modern machine learning infrastructure ~ Experience training, finetuning,... 
    Senior
    Full time
    Relocation package

    Scout Ai

    Sunnyvale, CA
    20 hours ago
  •  ...patients worldwide.We’re a team of engineers, clinicians, and innovators...  ...is building a new AI-powered platform to support how...  ...automate. We're looking for a senior full-stack engineer to help build...  ...the First 90 DaysCore platform infrastructure — authentication, the... 
    Senior
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    Sunnyvale, CA
    2 days ago
  • $145.6k - $246.4k

     ...forefront of innovation, integrating advanced AI and autonomous driving technologies into...  ...connectivity.As a core member of our AI Infrastructure team, you will be responsible for...  ...or higher in Computer Science, Software Engineering, Artificial Intelligence, or related fields... 
    Senior
    Full time
    Overseas

    XPENG Motors

    Santa Clara, CA
    4 days ago
  • $250.8k - $286.2k

    Senior Lead AI Engineer (MLX Emerging AI Patterns) Overview: At Capital One, we are creating responsible and reliable AI systems, changing...  ...customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in... 
    Senior
    Full time
    Part time
    Local area

    Capital One Financial Corporation

    San Jose, CA
    20 hours ago
  • $192.1k - $249.6k

     ...Senior AI Inference Infrastructure Software EngineerNIO is a pioneer and a leading company in the premium smart electric vehicle market. Founded in...  ...looking for a senior AI Inference Infrastructure Software Engineer with strong hands-on experience building, optimizing,... 
    Full time
    Temporary work
    Immediate start
    Flexible hours

    NIO

    San Jose, CA
    20 hours ago
  • $170.5k - $315.49k

     ...AI Infrastructure EngineerWe are looking for a performance-obsessed AI Infrastructure Engineer to push LLM inference to its absolute limits on Intel's next-generation GPU architectures. In this role, you will dive deep into the inference stack and redefine peak performance... 
    Local area
    Immediate start

    Intel

    Santa Clara, CA
    20 hours ago
  •  ...next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded...  ...ROLE:  We are seeking a DevOps / Platform Engineer to join our team building and operating large-scale GPU compute infrastructure that powers AI and ML workloads. THE... 

    Advanced Micro Devices , Inc.

    San Jose, CA
    2 days ago
  •  ...AI Infrastructure EngineerWe're looking for an AI Infrastructure Engineer to build and operate the serving infrastructure. You will work hands-on with vLLM and SGLang on Kubernetes. This is a foundational infrastructure role on a small, high-autonomy team.Essential Duties... 
    Work at office
    Immediate start
    Relocation package

    Netpreme

    Santa Clara, CA
    2 days ago
  • $187k - $215k

     ...the intelligence layer for global trade and logistics. We turn fragmented, messy data into instant, reliable intelligence, enabling AI to reason and act in high-stakes, real-world environments. We are a small, sharp team operating at the frontier, and we are looking for... 
    Senior
    Work at office
    Flexible hours

    KlearNow.AI

    San Jose, CA
    4 days ago
  • $123.24k - $200k

     ...Overview Of Role As a Sr./Principal AI Engineer within TSMC's Artificial Intelligence for...  ...Architect and implement the underlying MLOps infrastructure, including model serving, automated...  ...concepts to both engineering teams and senior management audiences. Education: B.S.... 
    Senior
    Work at office

    TSMC

    San Jose, CA
    1 day ago
  •  ...Overview We are seeking expertise in Generative AI and Python to work directly with...  ...’ll serve as a trusted advisor, hands-on engineer, and delivery lead — driving real-world AI...  ...in a startup or high-growth environment. Seniority level Mid-Senior level Employment type Full... 
    Senior
    Full time

    Infinite Computer Solutions

    Sunnyvale, CA
    14 hours ago
  • $98k - $182k

     ...Want To Make An Impact On The World Of Technology We are looking for a talented Software Engineer with experience in Machine Learning. You will work at the intersection of AI, high-performance software engineering, and electronic design automation (EDA) to build... 
    Senior

    Cadence Inc

    San Jose, CA
    3 days ago
  • $255.65k - $299k

     ...Senior AI Software EngineerImmigration sponsorship is not available for this position.What you can expect:As a Senior AI Software Engineer, you will collaborate to design, implement, and optimize AI algorithms and software applications. You will ensure AI training, inference... 
    Senior
    Work at office
    Remote work

    Zoom Video Communications

    San Jose, CA
    14 hours ago
  •  ...Job Title: Senior AI/ML Engineer Work Location with ZIP: Sunnyvale, CA 94085 (Hybrid) Technical Hiring Criteria (Must Haves) Top 3 Required skills: Machine Learning, AI, Python Years of experience in each of the must-have skills: 7 Years Job description... 
    Senior

    eTeam

    Sunnyvale, CA
    20 hours ago
  • $147k - $237.5k

    Palo Alto Networks, Inc. in Santa Clara, California, is looking for a skilled technical professional to develop and support cloud-based services. The role includes responsibilities such as requirements analysis, technical design, and collaboration with QA teams. Ideal ...
    Senior

    Jobleads-US

    Santa Clara, CA
    5 days ago
  •  ...Senior AI Data Science Engineer Advancing embodied AI in robotic surgery requires high-quality data across the robotic platform, the surgical...  ...data matters, drive the quality bar, and build the infrastructure that delivers research-ready datasets. This role calls... 
    Senior

    Intuitive

    Sunnyvale, CA
    12 hours ago
  •  ...Palo Alto Networks, Inc. is seeking a Senior Principal Software Engineer to lead development of tools, platforms, and infrastructure enabling engineering excellence at scale. This role combines deep software engineering expertise with strategic thinking and cross-team... 
    Senior

    Jobleads-US

    Santa Clara, CA
    1 day ago
  •  ...Senior Infrastructure EngineerWe are seeking a Senior Infrastructure Engineer with a strong focus on development and automation to build scalable, reliable, and efficient infrastructure solutions. The ideal candidate will have deep expertise in Java and Python, infrastructure... 
    Senior

    Omni Inclusive

    San Jose, CA
    14 hours ago
  • $170.6k - $261.3k

     ...Model Numerics Engineer The Compression and Parity team in GM's Autonomous Vehicle organization makes aggressive model optimization safe enough to ship repeatedly. We compress and quantize models headed for the car, and we own the analytical machinery that proves the... 
    Senior
    Flexible hours

    General Motors

    Sunnyvale, CA
    1 day ago
  •  ...Palo Alto Networks, Inc. is seeking a Senior Principal Backend Engineer for the Cortex group. You will lead backend development for Cortex XSOAR, XDR, and XSIAM, shaping data models, APIs, and scalable services in collaboration with cross-functional teams. The role... 
    Senior

    Jobleads-US

    Santa Clara, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior AI Infrastructure Engineer. Be the first to apply!