Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior AI Infrastructure Engineer

$180k - $240k

Gatik AI

Who we are

Gatik, the leader in autonomous middle-mile logistics, is revolutionizing the B2B supply chain with its autonomous transportation-as-a-service (ATaaS) solution and prioritizing safe, consistent deliveries while streamlining freight movement by reducing congestion. The company focuses on short-haul, B2B logistics for Fortune 500 retailers and in 2021 launched the world's first fully driverless commercial transportation service with Walmart. Gatik's Class 3-7 autonomous trucks are commercially deployed across major markets, including Texas, Arkansas, and Ontario, Canada, driving innovation in freight transportation.


The company's proprietary Level 4 autonomous technology, Gatik Carrier™, is custom-built to transport freight safely and efficiently between pick-up and drop-off locations on the middle mile. With robust capabilities in both highway and urban environments, Gatik Carrier™ serves as an all-encompassing solution that integrates advanced software and hardware powering the fleet, facilitating effortless integration into customers' logistics operations.
About the role

We are seeking a Senior AI Infrastructure Engineer to design, build, and scale the high-performance AI platform powering our autonomous driving models. While researchers focus on developing perception, planning, and world models, you will be responsible for the underlying infrastructure that enables distributed training, experiment tracking, and seamless model deployment. You will bridge the gap between research and production, ensuring our AI stack is scalable, resilient, and highly efficient

This role is onsite 5 days a week at our Santa Clara, CA office!

What you'll do

  • Distributed Training & ML Systems Support
    • Scale Research Workloads: Enable researchers to scale complex models (VLA, World Models) across multi-node setups using PyTorch Distributed, and Ray Train.
    • Performance Optimization: Architect and optimize multi-GPU setups, ensuring efficient model parallelism and data parallelism techniques across H100/A100 clusters.
    • Networking & Hardware Tuning: Optimize low-level communication (e.g., NCCL tuning, InfiniBand, or RoCE v2) to minimize latency for 3D Gaussian Splatting (3DGS) and large-scale training.
    • Intelligent Resource Scheduling: Optimize hardware utilization and cost-efficiency through Kubernetes-native GPU scheduling (NVIDIA GPU Operator, KubeFlow).
    • Inference Performance Engineering: Deploy and scale optimized model artifacts using TensorRT, ONNX Runtime, and Triton Inference Server, fine-tuning pipelines for both real-time and batch processing
  • Agentic Infrastructure & Automation
    • Self-Healing AI Infrastructure: Architect and deploy Autonomous AI Agents (LangGraph, CrewAI, or AutoGen) to monitor GPU cluster health, enabling automated real-time triage of hardware failures and NCCL timeouts.
    • Agentic DevOps & CI/CD: Develop agent-driven automation, such as Agentic PR Reviewers for infrastructure code and AI agents that proactively suggest model-specific Kubernetes resource optimizations.
    • Agentic Data Curation: Support researchers in building "Data Machines" where AI agents autonomously curate, label, and verify high-priority edge cases from raw data.
  • Model Management & Lifecycle (MLOps)
    • Automated Lifecycle Management: Design and maintain ML infrastructure leveraging MLFlow, Argo Workflows, and Kubernetes to automate the end-to-end model lifecycle.
    • Experiment & Model Tracking: Integrate feature stores and experiment tracking systems to provide a robust system of record for every model iteration.
    • Deployment Strategies: Implement robust serving mechanisms, including A/B testing, shadow deployments, and rollback mechanisms
  • Cloud-Native Foundations & Data Integration
    • Infrastructure as Code: Drive the "Everything as Code" philosophy using Terraform and Helm.
    • Data Pipelines: Collaborate with data teams to scale ETL pipelines using Apache Airflow, Kafka, and Spark for large-scale dataset management. •
    • Integrated Data Factories: Collaborate with data engineering teams to scale high-bandwidth ETL pipelines using Apache Airflow, Kafka, and Spark, ensuring seamless data flow from raw sensor logs to optimized storage in S3, GCS, or Delta Lake
  • Monitoring & Observability
    • System Metrics: Define and track key ML system metrics, including training convergence, latency, throughput, and drift detection.
    • Infrastructure Health: Maintain deep visibility into platform health using Prometheus, Grafana, OpenTelemetry, and ELK Stack.
    • Deep Stack Observability: Develop comprehensive monitoring using Prometheus, Grafana, and OpenTelemetry to track low-level infrastructure health alongside high-level ML metrics like training convergence and throughput.
    • AI-Specific Metrics & Drift: Define and monitor critical ML system KPIs, including model latency, inference throughput, and feature drift detection
What we're looking for
  • Experience: 5+ years in ML infrastructure, MLOps, or DevOps supporting high-scale compute environments.
  • ML Expertise: Deep understanding of multi-GPU training strategies (FSDP, DeepSpeed, Ray Train) and high-performance networking (NCCL, InfiniBand).
  • Infrastructure Automation: Mastery of Kubernetes, Terraform, and Helm, with a focus on GPU-native orchestration.
  • AI Agent Frameworks: Proven experience building or supporting Agentic Workflows for infrastructure or data automation (e.g., using LLMs to drive DevOps tasks).
  • Platform Mastery: Expertise in MLFlow, Argo Workflows, and Kubernetes.
  • Containerization: Strong experience with Docker, Kubernetes, and Helm.
  • Data & CI/CD: Proficiency in Apache Airflow, Kafka, Spark, and GitOps automation.
  • Core Skills: Proficiency in Python and Bash; experience with Go or Rust is a plus
Bonus Qualifications
  • Advanced AI Protocols: Familiarity with the Model Context Protocol (MCP) to standardize how AI agents interact with internal databases and orchestration APIs.
  • Hybrid & Physical AI: Experience in hybrid cloud and on-prem GPU cluster management for Physical AI workloads (e.g., 3DGS, World Models).
  • Agentic Observability: Experience utilizing LLMs for semantic monitoring and log analysis to detect complex distributed system failures that traditional threshold-based alerts miss.
Salary Ranges - $180,000- $240,000
More about Gatik

Founded in 2017 by experts in autonomous vehicle technology, Gatik has rapidly expanded its presence to Mountain View, Dallas-Fort Worth, Arkansas, and Toronto. As the first and only company to achieve fully driverless middle-mile commercial deliveries, Gatik holds a unique and defensible position in the AV industry, with a clear trajectory toward sustainable growth and profitability.

We have delivered complete, proprietary AV technology - an integration of software and hardware - to enable earlier successes for our clients in constrained Level 4 autonomy. By choosing the middle mile - with defined point-to-point delivery, we have simplified some of the more complex AV challenges, enabling us to achieve full autonomy ahead of competitors. Given extensive knowledge of Gatik's well-defined, fixed route ODDs and hybrid architecture, we are able to hyper-optimize our models with exponentially less data, establish gate-keeping mechanisms to maintain explainability, and ensure continued safety of the system for unmanned operations.

Visit us at Gatik for more company information and Careers at Gatik for more open roles.
Notable News
  • Bloomberg: Autonomous Trucking Firm Gatik Inks Contracts Worth $600 Million
  • Forbes: Hundreds' Of Gatik Robot Delivery Trucks Headed For U.S. Roads
  • Forbes:Gatik And Loblaw Announce Largest Commercial Deployment Of AV Trucks
  • Forbes: Forget robotaxis. Upstart Gatik sees middle-mile deliveries as the path to profitable AVs
  • Tech Brew: Gatik AI exec unpacks the regulations that could shape the AV industry
  • Business Wire: Gatik Paves the Way for Safe Driverless Operations ('Freight-Only') at Scale with Industry-First Third-Party Safety Assessment Framework
  • Auto Futures: Autonomous Trucking Group Gatik Secures Investment From NIPPON EXPRESS HOLDINGS
  • Automotive News: Gatik foresees hundreds of self-driving trucks on road soon, and that's just the beginning
  • Forbes: Isuzu And Gatik Go All In To Scale Up Driverless Freight Services
  • Bloomberg: Autonomous Vehicle Startup Takes Off by Picking Off Easier Routes
  • Reuters: Driverless vehicles on limited routes bump along despite US robotaxi scrutiny
Taking care of our team

At Gatik, we connect people of extraordinary talent and experience to an opportunity to create a more resilient supply chain and contribute to our environment's sustainability. We are diverse in our backgrounds and perspectives yet united by a bold vision and shared commitment to our values. Our culture emphasizes the importance of collaboration, respect and agility.

We at Gatik strive to create a diverse and inclusive environment where everyone feels they have opportunities to succeed and grow because we know that together we can do great things. We are committed to an inclusive and diverse team. We do not discriminate based on race, color, ethnicity, ancestry, national origin, religion, sex, gender, gender identity, gender expression, sexual orientation, age, disability, veteran status, genetic information, marital status or any legally protected status.
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior AI Infrastructure Engineer in Santa Clara, CA vacancy
  • $200k - $322k

     ...deep learning ignited modern AI — the next era of computing....  ...to join us today.Design-for-X Engineering at NVIDIA works on groundbreaking...  ....What you'll be doing:As a senior member in our team, you will...  ...cycles as part of the AI Infrastructure requirements at an org-wide level... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

    Joining NVIDIA's DGX Cloud AI Efficiency Team means contributing to the infrastructure that powers our innovative AI research...  ...an AI infrastructure software engineer to join our team. You'll be instrumental...  ...availability of AI systems.As a senior DGX Cloud AI Infrastructure... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    7 hours ago
  • $184k - $287.5k

     ...tapping into the unlimited potential of AI to define the next era of computing....  ...Co-Design Group (SCG) is seeking Senior AI Platform Engineers. They will set the technical direction...  ...platforms at the intersection of ML infrastructure and large-scale systems, this is your... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    7 hours ago
  • $151.8k - $265.35k

     ...organizations to create exceptional content effortlessly. The AI for Engineering team builds a scalable, production-grade AI platform that...  ...engines, tools, and data streams into adaptive AI systems.Mentor senior engineers in modern AI system design, LLM orchestration... 
    Senior
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Jose, CA
    7 hours ago
  • $151.8k - $332.2k

    What you can expectWe're looking for an experienced AI Engineer passionate about building the foundational AI Services Platform that...  ...and platform engineering, designing and implementing core AI infrastructure including context layer intelligence, AI agents, RAG (... 
    Senior
    Full time
    Work at office
    Remote work

    Zoom

    San Jose, CA
    2 days ago
  • $178k - $321k

     ...Audit function has an early but working AI-native capability: a multi-agent platform...  ...setting, and the governed data and AI infrastructure everything else depends on. We hire on demonstrated...  ...just implement it. This is a two-person engineering team: you deploy, debug, and hotfix your... 
    Senior

    OKX

    San Jose, CA
    2 days ago
  • $174.72k - $295.68k

     ...forefront of innovation, integrating advanced AI and autonomous driving technologies into...  ...connectivity.As a core member of our AI Infrastructure team, you will be responsible for...  ...or higher in Computer Science, Software Engineering, Artificial Intelligence, or related fields... 
    Senior
    Full time
    Overseas

    XPENG Motors

    Santa Clara, CA
    7 hours ago
  • $229.9k - $262.4k

     ...Senior Lead AI Engineer (Gen AI Platform Services) Overview: At Capital One, we are creating responsible and reliable AI systems, changing...  ...customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in... 
    Senior
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    1 day ago
  • $229.9k - $262.4k

    Senior Lead AI Engineer (Gen AI Platform Services, Agentic AI) Overview: At Capital One, we are creating responsible and reliable AI systems...  ...customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in... 
    Senior
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    4 days ago
  • $229.9k - $262.4k

    Senior Lead AI Engineer (GenAI Platform, Agentic Infrastructure) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time,... 
    Senior
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    2 days ago
  • Lendistry, LLC. is seeking a Senior AI Engineer to lead the delivery of AI solutions, including document intelligence and risk assessment tools. In this role, you will be responsible for mentoring junior engineers and shaping AI-driven workflows, improving the borrower... 
    Senior

    Lendistry, LLC.

    Santa Clara, CA
    3 days ago
  •  ...GRAIL is seeking a visionary Director of Software Engineering to lead our Data Platform engineering organization in Sunnyvale, CA. You will drive technical strategy, architecture, and delivery of a cloud-native data platform that handles petabyte-scale datasets. You... 
    Senior

    Jobleads-US

    Sunnyvale, CA
    3 days ago
  • $190k - $260k

     ...has developed an artificial intelligence (AI) powered technology stack purpose-built...  ...large-scale world models - depends on infrastructure that turns thousands of hours of multimodal...  ...training throughput. We are looking for engineers who make model training fast: streaming... 
    Senior
    Temporary work
    Work at office
    Visa sponsorship

    Kodiak Robotics

    Mountain View, CA
    7 hours ago
  • $191k - $315k

     ...Overview: The Network Growth and Relationship AI team is at the forefront of creating...  ...in close collaboration with the product, engineering and data science team and has a very exciting...  ...conferences. Responsibilities: As a senior technical leader in the Network Growth AI... 
    Senior
    For contractors
    Work at office
    Flexible hours

    Linkedin

    Mountain View, CA
    3 days ago
  •  ...next-generation computing experiences—from AI and data centers, to PCs, gaming and...  ...THE ROLEWe are hiring AI / ML Platform Engineers to build the platform layer that makes AI...  ...reproducible. This role focuses on the infrastructure and platform systems that support large-... 
    Senior

    AMD

    Santa Clara, CA
    1 day ago
  •  ...Senior Lead Software Engineer Be an integral part of an agile team that's constantly pushing the envelope...  ...Chase within the Corporate Sector, Infrastructure Platforms team, you are an integral...  ...infrastructure platforms optimized for AI and machine learning workloads.... 
    Senior
    For contractors

    Hackajob

    Palo Alto, CA
    1 day ago
  •  ...Nubank is seeking a Staff Software Engineer to lead the technical architecture for a private AI banking platform. You’ll own from Flutter mobile apps to distributed backends, enabling real-time AI-driven personalization at scale across web and mobile. You will collaborate... 
    Senior

    Jobleads-US

    Palo Alto, CA
    1 day ago
  • Voltai is a Bay Area-based AI hardware company seeking backend engineers to deliver our platform to leading semiconductor and electronics customers. You will own deployments across the SF Bay Area, Cupertino, and Santa Clara, and port features from our cloud product to... 
    Senior

    Voltai Inc.

    Palo Alto, CA
    1 day ago
  • $229.9k - $262.4k

    Sr. Lead AI Engineer (GenAI Platform) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking...  ...customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in machine... 
    Senior
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    3 days ago
  • Athelas is seeking a Senior Software Engineer for its R&D team in Mountain View, CA. In this hybrid role, you will build infrastructure for sub-agents, deploy advanced AI tools, and collaborate closely with product teams. The ideal candidate has strong fullstack skills,... 
    Senior

    Athelas

    Mountain View, CA
    5 days ago
  • $200k - $400k

     ...intelligent machines at scale. At Scout AI, we’re developing Fury, the first...  .... The Role We're looking for a Senior or Staff AI Engineer to join the Fury Orchestration Team with...  ...PyTorch, and modern machine learning infrastructure ~ Experience training, finetuning,... 
    Senior
    Full time
    Relocation package

    Scout Ai

    Sunnyvale, CA
    1 day ago
  • $170.5k - $315.49k

    Job Details:Job Description: We are looking for a performance-obsessed AI Infrastructure Engineer to push LLM inference to its absolute limits on Intel's next-generation GPU architectures.In this role, you will dive deep into the inference stack and redefine peak performance... 
    Full time
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    4 days ago
  • $124.36k - $146.3k

     ...all from Day One. Job Description Job Summary The Senior Engineer (Generative AI) is responsible for designing, developing, and deploying...  ...design for scalable AI workloads Utilize modern infrastructure practices: Containerization (Docker) Orchestration... 
    Senior
    Full time
    Temporary work
    Work experience placement
    Local area
    3 days per week

    U.S. Bank

    Cupertino, CA
    1 day ago
  • $152k - $241.5k

     ...about redefining how software is built in the age of Generative AI? Join NVIDIA’s TensorRT team to help lead a first-of-its-kind, AI...  ...at an unprecedented scale.If you are a systems-thinking C++ engineer who wants to help scale out an agentic development framework, stay... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

    We are now looking for a Senior Agentic AI Software Engineer! Today, NVIDIA is tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPUs act as the brains of computers, robots, and self-driving cars that can understand the... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

    We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join our AI Efficiency efforts. As an Engineer, you will have...  ...closely with customers to pinpoint and address infrastructure and application deficiencies, facilitating groundbreaking... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    7 hours ago
  • $184k - $287.5k

    We're looking for outstanding AI systems engineers to develop groundbreaking technologies in the inference systems software stack! We build innovative AI systems software to accelerate for AI inference. As a member of the team, you'll develop libraries, code generators... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

     ...computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain...  ...known as “the AI computing company”.NVIDIA is hiring a Senior AI Compiler Engineer. GPUs are driving rapid progress in deep learning—from LLMs... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...Our technology powers everything from generative AI to autonomous systems, and we continue to shape...  ...Managed AI Superclusters (MARS) builds and scales the infrastructure, platforms, and tools that enable researchers and engineers to develop the next generation of AI/ML systems.... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    7 hours ago
  • $152k - $241.5k

    We are now looking for a Senior AI Frameworks Engineer (C++/Python)! NVIDIA's high-performance computing platforms are powering the AI revolution...  ...and deep learning frameworks.Develop robust compilation infrastructure—including AST transformations and JIT-friendly execution... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    7 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior AI Infrastructure Engineer. Be the first to apply!