Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior AI Infrastructure Engineer

$180k - $240k

Gatik AI

Who we are

Gatik, the leader in autonomous middle-mile logistics, is revolutionizing the B2B supply chain with its autonomous transportation-as-a-service (ATaaS) solution and prioritizing safe, consistent deliveries while streamlining freight movement by reducing congestion. The company focuses on short-haul, B2B logistics for Fortune 500 retailers and in 2021 launched the world's first fully driverless commercial transportation service with Walmart. Gatik's Class 3-7 autonomous trucks are commercially deployed across major markets, including Texas, Arkansas, and Ontario, Canada, driving innovation in freight transportation.


The company's proprietary Level 4 autonomous technology, Gatik Carrier™, is custom-built to transport freight safely and efficiently between pick-up and drop-off locations on the middle mile. With robust capabilities in both highway and urban environments, Gatik Carrier™ serves as an all-encompassing solution that integrates advanced software and hardware powering the fleet, facilitating effortless integration into customers' logistics operations.
About the role

We are seeking a Senior AI Infrastructure Engineer to design, build, and scale the high-performance AI platform powering our autonomous driving models. While researchers focus on developing perception, planning, and world models, you will be responsible for the underlying infrastructure that enables distributed training, experiment tracking, and seamless model deployment. You will bridge the gap between research and production, ensuring our AI stack is scalable, resilient, and highly efficient

This role is onsite 5 days a week at our Santa Clara, CA office!

What you'll do

  • Distributed Training & ML Systems Support
    • Scale Research Workloads: Enable researchers to scale complex models (VLA, World Models) across multi-node setups using PyTorch Distributed, and Ray Train.
    • Performance Optimization: Architect and optimize multi-GPU setups, ensuring efficient model parallelism and data parallelism techniques across H100/A100 clusters.
    • Networking & Hardware Tuning: Optimize low-level communication (e.g., NCCL tuning, InfiniBand, or RoCE v2) to minimize latency for 3D Gaussian Splatting (3DGS) and large-scale training.
    • Intelligent Resource Scheduling: Optimize hardware utilization and cost-efficiency through Kubernetes-native GPU scheduling (NVIDIA GPU Operator, KubeFlow).
    • Inference Performance Engineering: Deploy and scale optimized model artifacts using TensorRT, ONNX Runtime, and Triton Inference Server, fine-tuning pipelines for both real-time and batch processing
  • Agentic Infrastructure & Automation
    • Self-Healing AI Infrastructure: Architect and deploy Autonomous AI Agents (LangGraph, CrewAI, or AutoGen) to monitor GPU cluster health, enabling automated real-time triage of hardware failures and NCCL timeouts.
    • Agentic DevOps & CI/CD: Develop agent-driven automation, such as Agentic PR Reviewers for infrastructure code and AI agents that proactively suggest model-specific Kubernetes resource optimizations.
    • Agentic Data Curation: Support researchers in building "Data Machines" where AI agents autonomously curate, label, and verify high-priority edge cases from raw data.
  • Model Management & Lifecycle (MLOps)
    • Automated Lifecycle Management: Design and maintain ML infrastructure leveraging MLFlow, Argo Workflows, and Kubernetes to automate the end-to-end model lifecycle.
    • Experiment & Model Tracking: Integrate feature stores and experiment tracking systems to provide a robust system of record for every model iteration.
    • Deployment Strategies: Implement robust serving mechanisms, including A/B testing, shadow deployments, and rollback mechanisms
  • Cloud-Native Foundations & Data Integration
    • Infrastructure as Code: Drive the "Everything as Code" philosophy using Terraform and Helm.
    • Data Pipelines: Collaborate with data teams to scale ETL pipelines using Apache Airflow, Kafka, and Spark for large-scale dataset management. •
    • Integrated Data Factories: Collaborate with data engineering teams to scale high-bandwidth ETL pipelines using Apache Airflow, Kafka, and Spark, ensuring seamless data flow from raw sensor logs to optimized storage in S3, GCS, or Delta Lake
  • Monitoring & Observability
    • System Metrics: Define and track key ML system metrics, including training convergence, latency, throughput, and drift detection.
    • Infrastructure Health: Maintain deep visibility into platform health using Prometheus, Grafana, OpenTelemetry, and ELK Stack.
    • Deep Stack Observability: Develop comprehensive monitoring using Prometheus, Grafana, and OpenTelemetry to track low-level infrastructure health alongside high-level ML metrics like training convergence and throughput.
    • AI-Specific Metrics & Drift: Define and monitor critical ML system KPIs, including model latency, inference throughput, and feature drift detection
What we're looking for
  • Experience: 5+ years in ML infrastructure, MLOps, or DevOps supporting high-scale compute environments.
  • ML Expertise: Deep understanding of multi-GPU training strategies (FSDP, DeepSpeed, Ray Train) and high-performance networking (NCCL, InfiniBand).
  • Infrastructure Automation: Mastery of Kubernetes, Terraform, and Helm, with a focus on GPU-native orchestration.
  • AI Agent Frameworks: Proven experience building or supporting Agentic Workflows for infrastructure or data automation (e.g., using LLMs to drive DevOps tasks).
  • Platform Mastery: Expertise in MLFlow, Argo Workflows, and Kubernetes.
  • Containerization: Strong experience with Docker, Kubernetes, and Helm.
  • Data & CI/CD: Proficiency in Apache Airflow, Kafka, Spark, and GitOps automation.
  • Core Skills: Proficiency in Python and Bash; experience with Go or Rust is a plus
Bonus Qualifications
  • Advanced AI Protocols: Familiarity with the Model Context Protocol (MCP) to standardize how AI agents interact with internal databases and orchestration APIs.
  • Hybrid & Physical AI: Experience in hybrid cloud and on-prem GPU cluster management for Physical AI workloads (e.g., 3DGS, World Models).
  • Agentic Observability: Experience utilizing LLMs for semantic monitoring and log analysis to detect complex distributed system failures that traditional threshold-based alerts miss.
Salary Ranges - $180,000- $240,000
More about Gatik

Founded in 2017 by experts in autonomous vehicle technology, Gatik has rapidly expanded its presence to Mountain View, Dallas-Fort Worth, Arkansas, and Toronto. As the first and only company to achieve fully driverless middle-mile commercial deliveries, Gatik holds a unique and defensible position in the AV industry, with a clear trajectory toward sustainable growth and profitability.

We have delivered complete, proprietary AV technology - an integration of software and hardware - to enable earlier successes for our clients in constrained Level 4 autonomy. By choosing the middle mile - with defined point-to-point delivery, we have simplified some of the more complex AV challenges, enabling us to achieve full autonomy ahead of competitors. Given extensive knowledge of Gatik's well-defined, fixed route ODDs and hybrid architecture, we are able to hyper-optimize our models with exponentially less data, establish gate-keeping mechanisms to maintain explainability, and ensure continued safety of the system for unmanned operations.

Visit us at Gatik for more company information and Careers at Gatik for more open roles.
Notable News
  • Bloomberg: Autonomous Trucking Firm Gatik Inks Contracts Worth $600 Million
  • Forbes: Hundreds' Of Gatik Robot Delivery Trucks Headed For U.S. Roads
  • Forbes:Gatik And Loblaw Announce Largest Commercial Deployment Of AV Trucks
  • Forbes: Forget robotaxis. Upstart Gatik sees middle-mile deliveries as the path to profitable AVs
  • Tech Brew: Gatik AI exec unpacks the regulations that could shape the AV industry
  • Business Wire: Gatik Paves the Way for Safe Driverless Operations ('Freight-Only') at Scale with Industry-First Third-Party Safety Assessment Framework
  • Auto Futures: Autonomous Trucking Group Gatik Secures Investment From NIPPON EXPRESS HOLDINGS
  • Automotive News: Gatik foresees hundreds of self-driving trucks on road soon, and that's just the beginning
  • Forbes: Isuzu And Gatik Go All In To Scale Up Driverless Freight Services
  • Bloomberg: Autonomous Vehicle Startup Takes Off by Picking Off Easier Routes
  • Reuters: Driverless vehicles on limited routes bump along despite US robotaxi scrutiny
Taking care of our team

At Gatik, we connect people of extraordinary talent and experience to an opportunity to create a more resilient supply chain and contribute to our environment's sustainability. We are diverse in our backgrounds and perspectives yet united by a bold vision and shared commitment to our values. Our culture emphasizes the importance of collaboration, respect and agility.

We at Gatik strive to create a diverse and inclusive environment where everyone feels they have opportunities to succeed and grow because we know that together we can do great things. We are committed to an inclusive and diverse team. We do not discriminate based on race, color, ethnicity, ancestry, national origin, religion, sex, gender, gender identity, gender expression, sexual orientation, age, disability, veteran status, genetic information, marital status or any legally protected status.
Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Senior AI Infrastructure Engineer in Santa Clara, CA vacancy
  • $172.5k - $306.63k

     ...organizations to create exceptional content effortlessly. The AI for Engineering team builds a scalable, production‑grade AI platform that...  ...engines, tools, and data streams into adaptive AI systems. Mentor senior engineers in modern AI system design, LLM orchestration... 
    Senior
    Local area

    Frame USA

    San Jose, CA
    3 days ago
  • $229.9k - $262.4k

     ...we are creating responsible and reliable AI systems, changing banking for good. For...  ...experiences. Our investments in technology infrastructure and world‑class talent, along with our...  ...world‑class applied science and engineering teams to deliver industry‑leading capabilities... 
    Senior
    Full time
    Part time
    Local area
    Visa sponsorship

    Capital One

    San Jose, CA
    2 days ago
  •  ...A tech company specializing in AI infrastructure is seeking a Software Engineer to build a scalable compute platform for its generative video models. The ideal candidate will have over 5 years of experience in MLOps or AI infrastructure management, along with strong Python... 
    Senior

    HeyGen

    Palo Alto, CA
    4 days ago
  • $180k - $200k

     ...Twin Health is the only company applying AI Digital Twin technology exclusively...  ...seeking a dynamic and innovative ML Ops Engineer. The ideal candidate is self-driven, versatile...  .... Be a subject matter expert on ML infrastructure, providing guidance to both internal teams... 
    Senior
    Remote work
    Flexible hours

    Twin-Health

    Mountain View, CA
    2 days ago
  • $190k - $260k

     ...has developed an artificial intelligence (AI) powered technology stack purpose-built...  ...large-scale world models – depends on infrastructure that turns thousands of hours of multimodal...  ...training throughput. We are looking for engineers who make model training fast: streaming... 
    Senior
    Temporary work
    Work at office
    Visa sponsorship
    Flexible hours

    Kodiak

    Mountain View, CA
    19 days ago
  • $178k - $321k

     ...Audit function has an early but working AI-native capability: a multi-agent platform...  ...setting, and the governed data and AI infrastructure everything else depends on. We hire on...  ...just implement it. This is a two-person engineering team: you deploy, debug, and hotfix your... 
    Senior
    Full time

    OKX

    San Jose, CA
    4 days ago
  • $229.9k - $262.4k

     ...Overview Senior Lead AI Engineer (Gen AI Platform Services) Overview: At Capital One, we are creating responsible and reliable AI...  ...personalized customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience... 
    Senior
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    4 days ago
  • $229.9k - $262.4k

     ...Overview Senior Lead AI Engineer (Gen AI Platform Services, Agentic AI) Overview: At Capital One, we are creating responsible...  ...personalized customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience... 
    Senior
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    more than 2 months ago
  • $250.8k - $286.2k

    Senior Lead AI Engineer (MLX Emerging AI Patterns) Overview: At Capital One, we are creating responsible and reliable...  ...customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in... 
    Senior
    Full time
    Part time
    Local area

    Capital One Financial Corporation

    San Jose, CA
    8 hours ago
  • $94.5k - $212.5k

     ...That's why we continuously invest in innovative ideas, such as AI-enabled insights and technology-powered solutions, to...  ...future of our industry. About the Role As an AI Infrastructure Engineer, you are a deep technical contributor with subsystem ownership... 
    Local area
    Worldwide

    Crowe

    San Jose, CA
    5 days ago
  • $227.5k - $300k

     ...At Sonatus, we’re driving the transformation to AI-enabled software-defined vehicles. Traditional automotive software methods...  ...future of mobility. Role Summary: We are seeking a Senior Staff AI Engineer with a combination of architectural expertise and production... 
    Senior
    Full time
    Work at office
    Worldwide
    Flexible hours
    Shift work

    Sonatus

    Sunnyvale, CA
    8 hours ago
  • $123.24k - $200k

     ...Senior / Principal AI Engineer for Business Intelligence Overview of Role As a Sr./Principal AI Engineer within TSMC's Artificial Intelligence...  ...for Scale: Architect and implement the underlying MLOps infrastructure, including model serving, automated testing, and... 
    Senior
    Work at office

    TSMC

    San Jose, CA
    2 days ago
  • $98k - $182k

     ...Want To Make An Impact On The World Of Technology We are looking for a talented Software Engineer with experience in Machine Learning. You will work at the intersection of AI, high-performance software engineering, and electronic design automation (EDA) to build... 
    Senior

    Cadence Inc

    San Jose, CA
    2 days ago
  •  ...who is passionate about crafting, implementing, and operating AI solutions that have a direct and measurable impact on Apple Sales and its customers. Description We’re looking for a Senior AI Engineer with strong software development skills and a passion for applying... 
    Senior
    Work experience placement

    Apple

    Cupertino, CA
    14 hours ago
  •  ...SOLN Job Level: Level 10 Direct/Indirect Indicator: Indirect Summary We are seeking a highly motivated and technically proficient AI Engineer to join our growing Data & Analytics team. In this role, you will be a key liaison between business stakeholders and the... 
    Senior
    Work experience placement
    Work at office
    Local area
    Remote work

    Celestica

    San Jose, CA
    3 days ago
  • $168.4k - $262.9k

     ...The Senior AI (Artificial Intelligence) Security Engineer will play a crucial role in safeguarding eBay’s AI projects against emerging threats. We are looking...  ...Analyze potential impacts on our AI projects and infrastructure. Security Advising: Serve as an authority on AI... 
    Senior
    Immediate start

    eBay Inc.

    San Jose, CA
    2 days ago
  •  ...Overview We are seeking expertise in Generative AI and Python to work directly with...  ...’ll serve as a trusted advisor, hands-on engineer, and delivery lead — driving real-world AI...  ...in a startup or high-growth environment. Seniority level Mid-Senior level Employment type Full... 
    Senior
    Full time

    Infinite Computer Solutions

    Sunnyvale, CA
    3 days ago
  •  ...Intuitive Surgical, Inc. is seeking a Senior Software Engineer to design, build, and operate foundational cloud infrastructure and developer enablement capabilities. This role will contribute to modern AWS infrastructure, CI/CD foundations, security instrumentation, and... 
    Senior

    Intuitive Surgical

    Sunnyvale, CA
    1 day ago
  •  ...AI Agent Engineer At F5, we strive to bring a better digital world to life. Our teams empower organizations across the globe to create, secure, and run applications that enhance how we experience our evolving digital world. We are passionate about cybersecurity, from... 
    Senior
    Work at office

    F5

    San Jose, CA
    1 day ago
  • $114.6k - $234.6k

     ...Senior AI Agent Engineer Oracle Health is seeking a Senior AI Agent Engineer to build production AI agents and workflow automation capabilities...  ...About Us Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry... 
    Senior
    Temporary work
    Flexible hours

    Oracle

    Santa Clara, CA
    3 days ago
  • $8 per hour

     ...SERV is building the AI commerce layer for home and commercial services. Buying a...  ...build a clear and transparent commerce infrastructure and we're expanding on this lead. Role You...  ...will be SERV's first dedicated Product Engineer for AI products. We have validated the Product... 
    Senior
    Price work

    SERV

    Campbell, CA
    2 days ago
  • Responsibilities As a Software and AI Engineer, IT Systems, you’ll design and develop AI agents and first‑party applications that bring intelligence, automation, and great user experience to CoreWeave’s internal ecosystem You’ll work across the stack ranging from agentic... 
    Senior

    CoreWeave

    Sunnyvale, CA
    5 days ago
  • $293.6k - $335.1k

     ...Distinguished AI Engineer (Agentic AI Platform) At Capital One, we are creating responsible...  .... Our investments in technology infrastructure and world-class talent — along with our...  ...hours, mentoring Staff, Principal and Senior engineers, authoring technical design documents... 
    Full time
    Part time
    Work at office
    Local area

    Capital One

    San Jose, CA
    2 days ago
  •  ...small smart high-end electric cars with the FIREFLY brand. About the Position   We are looking for a senior AI Inference Infrastructure Software Engineer with strong hands-on experience building, optimizing, and deploying high-performance, scalable inference systems... 
    Full time
    Temporary work
    Immediate start
    Flexible hours

    NIO USA, INC

    San Jose, CA
    9 days ago
  •  ...Senior Infrastructure Engineer We are seeking a Senior Infrastructure Engineer with a strong focus on development and automation to build scalable, reliable, and efficient infrastructure solutions. The ideal candidate will have deep expertise in Java and Python, infrastructure... 
    Senior

    Omni Inclusive

    San Jose, CA
    5 days ago
  •  ...SmithRx in California is seeking a Senior DevOps Engineer responsible for building and managing cloud-based infrastructures. The role entails developing CI/CD pipelines, monitoring systems, and ensuring scalability and security within the DevOps practices. With a collaborative... 
    Senior
    Remote work

    SmithRx

    San Jose, CA
    2 days ago
  • $140k - $165k

     ...most advanced electronic devices and IT infrastructure, enabling enhanced performance and user...  ...Why Join Us? Build foundational AI infrastructure that powers next-gen enterprise...  ...the Role: We are seeking a hands-on AI Engineer to design, deploy, and maintain on-prem... 
    Senior

    SK hynix memory solutions America Inc.

    San Jose, CA
    4 days ago
  •  ...we are creating responsible and reliable AI systems, changing banking for good. For...  ...experiences. Our investments in technology infrastructure and world‑class talent — along with our...  ...build world‑class applied science and engineering teams to deliver our industry‑leading capabilities... 
    Senior
    Full time
    Local area

    SwiftCruit

    San Jose, CA
    4 days ago
  • $126.4k - $192.8k

     ...from legacy energy sources toward a lower carbon future. As Senior Infrastructure Engineer, you will be the technical foundation that lets our...  ...SRE practices that move the team from reactive to proactive. AI‑driven operations – develop and deploy AIOps tooling, agentic... 
    Senior
    Work at office

    QuantumScape Corporation

    San Jose, CA
    2 days ago
  • $109.7k - $234.6k

     ...Oracle is hiring a Principal Software Engineer in Santa Clara, California, to lead software design and development for its Cloud Infrastructure. Ideal candidates must have 8+ years of experience in software engineering, proficient in languages such as Java, C, and C++.... 
    Senior

    Oracle

    Santa Clara, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior AI Infrastructure Engineer. Be the first to apply!