Senior AI Infrastructure Engineer
$180k - $240kGatik AI
Senior AI Infrastructure EngineerSanta Clara, CAAbout the RoleGatik, the leader in autonomous middle-mile logistics, is revolutionizing the B2B supply chain with its autonomous transportation-as-a-service (ATaaS) solution and prioritizing safe, consistent deliveries while streamlining freight movement by reducing congestion. The company focuses on short-haul, B2B logistics for Fortune 500 retailers and in 2021 launched the world's first fully driverless commercial transportation service with Walmart. Gatik's Class 3-7 autonomous trucks are commercially deployed across major markets, including Texas, Arkansas, and Ontario, Canada, driving innovation in freight transportation.The company's proprietary Level 4 autonomous technology, Gatik Carrier™, is custom-built to transport freight safely and efficiently between pick-up and drop-off locations on the middle mile. With robust capabilities in both highway and urban environments, Gatik Carrier™ serves as an all-encompassing solution that integrates advanced software and hardware powering the fleet, facilitating effortless integration into customers' logistics operations.We are seeking a Senior AI Infrastructure Engineer to design, build, and scale the high-performance AI platform powering our autonomous driving models. While researchers focus on developing perception, planning, and world models, you will be responsible for the underlying infrastructure that enables distributed training, experiment tracking, and seamless model deployment. You will bridge the gap between research and production, ensuring our AI stack is scalable, resilient, and highly efficient.This role is onsite 5 days a week at our Santa Clara, CA office!What You'll DoDistributed Training & ML Systems SupportScale Research Workloads: Enable researchers to scale complex models (VLA, World Models) across multi-node setups using PyTorch Distributed, and Ray Train.Performance Optimization: Architect and optimize multi-GPU setups, ensuring efficient model parallelism and data parallelism techniques across H100/A100 clusters.Networking & Hardware Tuning: Optimize low-level communication (e.g., NCCL tuning, InfiniBand, or RoCE v2) to minimize latency for 3D Gaussian Splatting (3DGS) and large-scale training.Intelligent Resource Scheduling: Optimize hardware utilization and cost-efficiency through Kubernetes-native GPU scheduling (NVIDIA GPU Operator, KubeFlow).Inference Performance Engineering: Deploy and scale optimized model artifacts using TensorRT, ONNX Runtime, and Triton Inference Server, fine-tuning pipelines for both real-time and batch processingAgentic Infrastructure & AutomationSelf-Healing AI Infrastructure: Architect and deploy Autonomous AI Agents (LangGraph, CrewAI, or AutoGen) to monitor GPU cluster health, enabling automated real-time triage of hardware failures and NCCL timeouts.Agentic DevOps & CI/CD: Develop agent-driven automation, such as Agentic PR Reviewers for infrastructure code and AI agents that proactively suggest model-specific Kubernetes resource optimizations.Agentic Data Curation: Support researchers in building "Data Machines" where AI agents autonomously curate, label, and verify high-priority edge cases from raw data.Model Management & Lifecycle (MLOps)Automated Lifecycle Management: Design and maintain ML infrastructure leveraging MLFlow, Argo Workflows, and Kubernetes to automate the end-to-end model lifecycle.Experiment & Model Tracking: Integrate feature stores and experiment tracking systems to provide a robust system of record for every model iteration.Deployment Strategies: Implement robust serving mechanisms, including A/B testing, shadow deployments, and rollback mechanismsCloud-Native Foundations & Data IntegrationInfrastructure as Code: Drive the "Everything as Code" philosophy using Terraform and Helm.Data Pipelines: Collaborate with data teams to scale ETL pipelines using Apache Airflow, Kafka, and Spark for large-scale dataset management.Integrated Data Factories: Collaborate with data engineering teams to scale high-bandwidth ETL pipelines using Apache Airflow, Kafka, and Spark, ensuring seamless data flow from raw sensor logs to optimized storage in S3, GCS, or Delta LakeMonitoring & ObservabilitySystem Metrics: Define and track key ML system metrics, including training convergence, latency, throughput, and drift detection.Infrastructure Health: Maintain deep visibility into platform health using Prometheus, Grafana, OpenTelemetry, and ELK Stack.Deep Stack Observability: Develop comprehensive monitoring using Prometheus, Grafana, and OpenTelemetry to track low-level infrastructure health alongside high-level ML metrics like training convergence and throughput.AI-Specific Metrics & Drift: Define and monitor critical ML system KPIs, including model latency, inference throughput, and feature drift detectionWhat We're Looking ForExperience: 5+ years in ML infrastructure, MLOps, or DevOps supporting high-scale compute environments.ML Expertise: Deep understanding of multi-GPU training strategies (FSDP, DeepSpeed, Ray Train) and high-performance networking (NCCL, InfiniBand).Infrastructure Automation: Mastery of Kubernetes, Terraform, and Helm, with a focus on GPU-native orchestration.AI Agent Frameworks: Proven experience building or supporting Agentic Workflows for infrastructure or data automation (e.g., using LLMs to drive DevOps tasks).Platform Mastery: Expertise in MLFlow, Argo Workflows, and Kubernetes.Containerization: Strong experience with Docker, Kubernetes, and Helm.Data & CI/CD: Proficiency in Apache Airflow, Kafka, Spark, and GitOps automation.Core Skills: Proficiency in Python and Bash; experience with Go or Rust is a plusBonus QualificationsAdvanced AI Protocols: Familiarity with the Model Context Protocol (MCP) to standardize how AI agents interact with internal databases and orchestration APIs.Hybrid & Physical AI: Experience in hybrid cloud and on-prem GPU cluster management for Physical AI workloads (e.g., 3DGS, World Models).Agentic Observability: Experience utilizing LLMs for semantic monitoring and log analysis to detect complex distributed system failures that traditional threshold-based alerts miss.Salary Ranges - $180,000- $240,000More About GatikFounded in 2017 by experts in autonomous vehicle technology, Gatik has rapidly expanded its presence to Mountain View, Dallas-Fort Worth, Arkansas, and Toronto. As the first and only company to achieve fully driverless middle-mile commercial deliveries, Gatik holds a unique and defensible position in the AV industry, with a clear trajectory toward sustainable growth and profitability.We have delivered complete, proprietary AV technology - an integration of software and hardware - to enable earlier successes for our clients in constrained Level 4 autonomy. By choosing the middle mile – with defined point-to-point delivery, we have simplified some of the more complex AV challenges, enabling us to achieve full autonomy ahead of competitors. Given extensive knowledge of Gatik's well-defined, fixed route ODDs and hybrid architecture, we are able to hyper-optimize our models with exponentially less data, establish gate-keeping mechanisms to maintain explainability, and ensure continued safety of the system for unmanned operations.Visit us at Gatik for more company information and Careers at Gatik for more open roles.
$286.2k - $326.7k
...Senior Director, AI Engineering -Agentic AI PlatformAt Capital One, we are creating responsible and reliable AI systems, changing banking for... ...personalized customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience...SeniorFull timePart timeRemote work$229.9k - $262.4k
Senior Lead AI Engineer (GenAI Platform Services) Overview At Capital One, we are creating responsible and reliable AI systems, changing banking... ...customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience...SeniorFull timePart timeLocal area- ...Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing... ...visit Position Overview We are seeking a Senior AI Storage Infrastructure Engineer to build the critical data-delivery fabric of our...SeniorRemote jobFull timeLocal area
$229.9k - $262.4k
...Sr. Lead AI Engineer (Gen AI Platform Services) Overview At Capital One, we are creating responsible and reliable AI systems,... ...personalized customer experiences. Our investments in technology infrastructure and world‑class talent — along with our deep experience in...SeniorLocal area- ...Nubank is seeking a Staff Software Engineer to lead the technical architecture for a private AI banking platform. You’ll own from Flutter mobile apps to distributed backends, enabling real-time AI-driven personalization at scale across web and mobile. You will collaborate...Senior
$184k - $287.5k
Joining NVIDIA's DGX Cloud AI Efficiency Team means contributing to the infrastructure that powers our innovative AI research... ...an AI infrastructure software engineer to join our team. You'll be instrumental... ...availability of AI systems.As a senior DGX Cloud AI Infrastructure...SeniorFull timeRemote work- ...JPMorganChase is seeking a Senior Director of Software Engineering to lead Data and AI platforms within the Commercial and Investment Bank Operations group. You will oversee multiple departments, set technical strategy, and drive enterprise-wide AI-enabled engineering...Senior
$184k - $287.5k
...tapping into the unlimited potential of AI to define the next era of computing.... ...Co-Design Group (SCG) is seeking Senior AI Platform Engineers. They will set the technical direction... ...platforms at the intersection of ML infrastructure and large-scale systems, this is your...SeniorFull time- ...Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing comprehensive... ...We are seeking a highly skilled and motivated Cloud Senior DevOps Engineer to join our AI Cloud team. In this high-impact role,...SeniorRemote jobFull timeLocal area
$178k - $321k
...Audit function has an early but working AI-native capability: a multi-agent platform... ...setting, and the governed data and AI infrastructure everything else depends on. We hire on demonstrated... ...just implement it. This is a two-person engineering team: you deploy, debug, and hotfix your...Senior$200k - $400k
...intelligent machines at scale. At Scout AI, we’re developing Fury, the first... .... The Role We're looking for a Senior or Staff AI Engineer to join the Fury Orchestration Team with... ...PyTorch, and modern machine learning infrastructure ~ Experience training, finetuning,...SeniorFull timeRelocation package- ...patients worldwide.We’re a team of engineers, clinicians, and innovators... ...is building a new AI-powered platform to support how... ...automate. We're looking for a senior full-stack engineer to help build... ...the First 90 DaysCore platform infrastructure — authentication, the...SeniorLocal areaWorldwideFlexible hours
$145.6k - $246.4k
...forefront of innovation, integrating advanced AI and autonomous driving technologies into... ...connectivity.As a core member of our AI Infrastructure team, you will be responsible for... ...or higher in Computer Science, Software Engineering, Artificial Intelligence, or related fields...SeniorFull timeOverseas$250.8k - $286.2k
Senior Lead AI Engineer (MLX Emerging AI Patterns) Overview: At Capital One, we are creating responsible and reliable AI systems, changing... ...customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in...SeniorFull timePart timeLocal area$192.1k - $249.6k
...Senior AI Inference Infrastructure Software EngineerNIO is a pioneer and a leading company in the premium smart electric vehicle market. Founded in... ...looking for a senior AI Inference Infrastructure Software Engineer with strong hands-on experience building, optimizing,...Full timeTemporary workImmediate startFlexible hours$170.5k - $315.49k
...AI Infrastructure EngineerWe are looking for a performance-obsessed AI Infrastructure Engineer to push LLM inference to its absolute limits on Intel's next-generation GPU architectures. In this role, you will dive deep into the inference stack and redefine peak performance...Local areaImmediate start- ...next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded... ...ROLE: We are seeking a DevOps / Platform Engineer to join our team building and operating large-scale GPU compute infrastructure that powers AI and ML workloads. THE...
- ...AI Infrastructure EngineerWe're looking for an AI Infrastructure Engineer to build and operate the serving infrastructure. You will work hands-on with vLLM and SGLang on Kubernetes. This is a foundational infrastructure role on a small, high-autonomy team.Essential Duties...Work at officeImmediate startRelocation package
$187k - $215k
...the intelligence layer for global trade and logistics. We turn fragmented, messy data into instant, reliable intelligence, enabling AI to reason and act in high-stakes, real-world environments. We are a small, sharp team operating at the frontier, and we are looking for...SeniorWork at officeFlexible hours$123.24k - $200k
...Overview Of Role As a Sr./Principal AI Engineer within TSMC's Artificial Intelligence for... ...Architect and implement the underlying MLOps infrastructure, including model serving, automated... ...concepts to both engineering teams and senior management audiences. Education: B.S....SeniorWork at office- ...Overview We are seeking expertise in Generative AI and Python to work directly with... ...’ll serve as a trusted advisor, hands-on engineer, and delivery lead — driving real-world AI... ...in a startup or high-growth environment. Seniority level Mid-Senior level Employment type Full...SeniorFull time
$98k - $182k
...Want To Make An Impact On The World Of Technology We are looking for a talented Software Engineer with experience in Machine Learning. You will work at the intersection of AI, high-performance software engineering, and electronic design automation (EDA) to build...Senior$255.65k - $299k
...Senior AI Software EngineerImmigration sponsorship is not available for this position.What you can expect:As a Senior AI Software Engineer, you will collaborate to design, implement, and optimize AI algorithms and software applications. You will ensure AI training, inference...SeniorWork at officeRemote work- ...Job Title: Senior AI/ML Engineer Work Location with ZIP: Sunnyvale, CA 94085 (Hybrid) Technical Hiring Criteria (Must Haves) Top 3 Required skills: Machine Learning, AI, Python Years of experience in each of the must-have skills: 7 Years Job description...Senior
$147k - $237.5k
Palo Alto Networks, Inc. in Santa Clara, California, is looking for a skilled technical professional to develop and support cloud-based services. The role includes responsibilities such as requirements analysis, technical design, and collaboration with QA teams. Ideal ...Senior- ...Senior AI Data Science Engineer Advancing embodied AI in robotic surgery requires high-quality data across the robotic platform, the surgical... ...data matters, drive the quality bar, and build the infrastructure that delivers research-ready datasets. This role calls...Senior
- ...Palo Alto Networks, Inc. is seeking a Senior Principal Software Engineer to lead development of tools, platforms, and infrastructure enabling engineering excellence at scale. This role combines deep software engineering expertise with strategic thinking and cross-team...Senior
- ...Senior Infrastructure EngineerWe are seeking a Senior Infrastructure Engineer with a strong focus on development and automation to build scalable, reliable, and efficient infrastructure solutions. The ideal candidate will have deep expertise in Java and Python, infrastructure...Senior
$170.6k - $261.3k
...Model Numerics Engineer The Compression and Parity team in GM's Autonomous Vehicle organization makes aggressive model optimization safe enough to ship repeatedly. We compress and quantize models headed for the car, and we own the analytical machinery that proves the...SeniorFlexible hours- ...Palo Alto Networks, Inc. is seeking a Senior Principal Backend Engineer for the Cortex group. You will lead backend development for Cortex XSOAR, XDR, and XSIAM, shaping data models, APIs, and scalable services in collaboration with cross-functional teams. The role...Senior
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior AI Infrastructure Engineer. Be the first to apply!
- senior ai engineer Santa Clara, CA
- ai engineer remote Santa Clara, CA
- ai developer Santa Clara, CA
- ai prompt engineer Santa Clara, CA
- ai engineer Santa Clara, CA
- infrastructure developer Santa Clara, CA
- principal infrastructure engineer Santa Clara, CA
- remote infrastructure engineer Santa Clara, CA
- infrastructure engineering manager Santa Clara, CA
- infrastructure engineer Santa Clara, CA


