Senior Cloud Infrastructure Engineer
$180k - $240kGatik AI
Senior Cloud Infrastructure EngineerSanta Clara, CAAbout the RoleWe are seeking a Senior Cloud Infrastructure Engineer to architect and manage the large-scale compute and data infrastructure powering our autonomous driving stack. While researchers develop perception, planning, and world models, your mission is to build the high-performance systems and pipelines that make their work possible. You will be the backbone of our AI platform, ensuring that multi-GPU clusters, distributed training frameworks, and automated workflows are scalable, resilient, and cost-effective. This role is onsite 5 days a week at our Santa Clara, CA office!What You'll DoCloud-Native Orchestration & Kubernetes Advanced K8s Management: Architect and maintain mission-critical Kubernetes clusters optimized for heavy GPU/TPU workloads.GPU Scheduling: Implement and optimize Kubernetes-native GPU scheduling (NVIDIA GPU Operator) to ensure maximum hardware utilization.Infrastructure as Code: Drive the "Everything as Code" philosophy using Terraform, Helm, and cloud-native tools.Self-Healing Infrastructure: Deploy Autonomous AI Agents (LangGraph, CrewAI) to monitor cluster health and enable automated triage of hardware failures and NCCL timeouts.Data Engineering & CI/CD Pipelines Autonomy Data Pipelines: Build large-scale pipelines using Apache Airflow, Kafka, and Spark to process raw sensor data into training-ready formats.GitOps: Implement robust GitOps workflows using ArgoCD, Gitlab CI/CD to automate the deployment of both infrastructure and model artifacts.Observability: Maintain deep visibility into infrastructure health and model serving performance using Prometheus, Grafana, and OpenTelemetry.Agentic DevOps & CI/CD: Develop agent-driven workflows to optimize the developer experience, such as automated PR reviewers for Terraform and AI agents that proactively suggest Kubernetes resource-limit adjustments based on model training telemetry.Model Management & Lifecycle (MLOps) Experiment & Model Tracking: Design and maintain MLFlow and feature store integrations to provide a robust system of record for every model iteration.Workflow Automation: Build complex, automated model lifecycles using Airflow and Kubernetes to streamline the transition from training to simulation.High-Performance Serving: Support the deployment of models into simulation and production environments using Triton Inference Server, Ray Serve, and ONNX Runtime.Distributed Training & ML Systems Support Training Systems Support: Enable researchers to scale models (VLA, World Models) across multi-node setups using PyTorch Distributed (TorchElastic), Ray Train, and Horovod.Networking Optimization: Optimize low-level communication (e.g., NCCL tuning, InfiniBand, or RoCE v2) to minimize latency for 3D Gaussian Splatting (3DGS) and large-scale training.Hardware-Aware Orchestration: Partner with researchers to fine-tune performance across multi-node GPU clusters for FSDP and DeepSpeed workloads.What We're Looking ForExperience: 5+ years in Cloud Infrastructure, DevOps, or MLOps supporting high-scale compute environments.Kubernetes Mastery: Deep expertise in K8s, Helm, and container orchestration.Orchestration & Tooling: Strong background in Apache Airflow, Argo Workflows, MLFlow, and Terraform.Distributed Systems: Practical experience supporting frameworks like Ray and PyTorch Distributed.Core Skills: Proficiency in Python, Bash scripting, and a solid understanding of IAM/RBAC.Bonus QualificationsDistributed Training Expertise: Deep understanding of FSDP, and DeepSpeed.AI Agent Orchestration: Experience building Agentic Workflows (LangGraph, AutoGen) for infrastructure automation or data curation.Advanced Protocols: Familiarity with Model Context Protocol (MCP) to connect AI agents with infrastructure tools.Salary Range - $180,000- $240,000More About GatikFounded in 2017 by experts in autonomous vehicle technology, Gatik has rapidly expanded its presence to Mountain View, Dallas-Fort Worth, Arkansas, and Toronto. As the first and only company to achieve fully driverless middle-mile commercial deliveries, Gatik holds a unique and defensible position in the AV industry, with a clear trajectory toward sustainable growth and profitability.We have delivered complete, proprietary AV technology - an integration of software and hardware - to enable earlier successes for our clients in constrained Level 4 autonomy. By choosing the middle mile – with defined point-to-point delivery, we have simplified some of the more complex AV challenges, enabling us to achieve full autonomy ahead of competitors. Given extensive knowledge of Gatik's well-defined, fixed route ODDs and hybrid architecture, we are able to hyper-optimize our models with exponentially less data, establish gate-keeping mechanisms to maintain explainability, and ensure continued safety of the system for unmanned operations.Visit us at Gatik for more company information and Careers at Gatik for more open roles.
- ...A leading cybersecurity firm based in Santa Clara is seeking a Sr Site Reliability Engineer. The candidate will be responsible for maintaining highly reliable cloud infrastructure and will lead cross-functional initiatives. Ideal candidates should have at least 5 years...Senior
$152k - $241.5k
...Vehicles Platform team is seeking a Senior System Software Engineer to help bring NVIDIA's autonomous vehicle... ...of hardware-in-the-loop (HIL) infrastructure to support the robust deployment and... ...g., Jenkins, Docker, Kubernetes) in cloud-native or hybrid environments for mission...SeniorFull time$174k - $253k
...accessible technologies. Google's software engineers develop the next-generation... ...enhance software solutions.The AI and Infrastructure team is redefining what’s possible. We... ...Our customers include Googlers, Google Cloud customers, and billions of Google users...SeniorWorldwide$176k - $276k
Cloud Foundations Reliability (CFR) is part of NVIDIA’s Global Network Infrastructure (GNI) organization. We deploy, integrate, and operate the Kubernetes-based platform... ...environments.We are looking for a hands-on senior engineer to own the lifecycle and automation of the...SeniorFull timeRemote workWeekend work- Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers... ...s designated work from home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our...SeniorWork at officeLocal areaWork from homeFlexible hours
$184k - $287.5k
Joining NVIDIA's DGX Cloud AI Efficiency Team means contributing to the infrastructure that powers our innovative AI research. This... ...an AI infrastructure software engineer to join our team. You'll be... ...availability of AI systems.As a senior DGX Cloud AI Infrastructure software...SeniorFull timeRemote work$280k - $380k
...runs an active-active, multi-cloud platform on AWS and GCP to... ...reliability and automation, engineering systems that perform under stress... ...can, and turning complex infrastructure into reliable, well‑... ...Site Reliability Engineering) Senior Software Engineer to join our...SeniorWork at officeLocal areaRemote workMonday to ThursdayFlexible hours$262k - $364k
...Lead and coach a distributed engineering team, fostering innovation... ...Systems or Machine Learning Infrastructure.Google's software engineers... ...workloads for Google and Google Cloud .TheHPN team is at the core... ...Storage (GDS).As a Senior Staff Software Engineer, you...SeniorRemote workWorldwide$159k - $231k
...qualifications:Bachelor’s degree in Electrical Engineering, Computer Engineering, Physics, a... ...Integrity Engineer within Platforms Infrastructure Engineering, you will play a pivotal role... ...Our customers include Googlers, Google Cloud customers, and billions of Google users...SeniorWorldwide$178k - $321k
...runtime, and the harness: a resilient cloud platform, the agentic runtime and evaluation... ...setting, and the governed data and AI infrastructure everything else depends on. We hire on... ...implement it. This is a two-person engineering team: you deploy, debug, and hotfix your...Senior$174k - $253k
...development code. Review code developed by other engineers and provide feedback to ensure best... ...enhance software solutions.The AI and Infrastructure team is redefining what’s possible. We... ...Our customers include Googlers, Google Cloud customers, and billions of Google users...SeniorWorldwide- ...A leading technology company is seeking a Software Engineer to develop next-generation technologies that change how billions of users... ...degree in Computer Science or related fields. Join a dynamic team at the forefront of AI and infrastructure innovation.#J-18808-Ljbffr...Senior
- ...Career At Palo Alto Networks, Secure Cloud and AI infrastructure is the foundation of our mission to protect... ...’re seeking a world‑class Principal Engineer (Sr Manager‑equivalent) to lead the... ...augmented cloud platforms, mentoring senior engineers and infusing industry‑...SeniorFull timeWork at office3 days per week
$159k - $230k
...reproduction.Design network, power and cooling infrastructure for end-to-end system testing.Drive... ...of hardware/software designers, test engineers, on project planning within hardware... ...manage the needs of multiple lab users.As a Senior Hardware Performance Test Engineer, you...Senior- ...Applied Intuition Inc. is seeking a software engineer to design, develop, and maintain backend platform infrastructure that supports deployment, authentication, and APIs. You will build distributed systems to accelerate internal app development and implement extensible...Senior
$150k - $218k
...programs; design the test fixture for better engineering efficiency and data accuracy.Set up and... ....Google designs powerful computing infrastructures with custom-built machines. The Test... ...Our customers include Googlers, Google Cloud customers, and billions of Google users...SeniorWorldwide$159k - $230k
...design issues.Support design engineers with debug, component validation... ...using test equipment. As a Senior Hardware Test Engineer, you... ...sustaining efforts.The AI and Infrastructure team is redefining what’s possible... ...include Googlers, Google Cloud customers, and billions of...SeniorWorldwide$159k - $231k
...align on Robotics initiatives and roadmap, engineering execution with business goals.Build and... ...Hardware Engineer to join our Physical infrastructure Robotics team. This team is dedicated... ...Our customers include Googlers, Google Cloud customers, and billions of Google users...SeniorContract workWorldwide$236k - $329k
...align on Robotics initiatives and roadmap, engineering execution with business goals.Identify,... ...specialized liquid-cooled AI hardware infrastructure.Oversee the end-to-end deployment of... ...Our customers include Googlers, Google Cloud customers, and billions of Google users...SeniorContract workRemote workWorldwideFlexible hours$166k - $244k
...Senior Software Engineer, Infrastructure, Platforms Infrastructure Engineering Mid Experience driving progress, solving problems, and mentoring... ...and velocity. Our customers include Googlers, Googler Cloud customers, and billions of Google users worldwide. We’re...SeniorFull timeWorldwide$184k - $287.5k
We are seeking a Senior DevOps / Cloud Simulation Infrastructure Engineer to own the complete end-to-end cloud execution pipeline for SimReady assets! This role is critical to our product strategy, enabling us to transition from local, workstation-driven validation to...SeniorFull timeLocal area$159k - $230k
...QUALIFICATIONS: Bachelor’s degree in Electrical Engineering, Computer Engineering, Physics, a... ..., efficiency, and integration. As a Senior Hardware Engineer in the Board and System... ...in the data center.Our Platforms Infrastructure Engineering team designs and builds the...SeniorWorldwide$174k - $252k
...connectivity solution for hybrid/multi cloud, involving data plane and control plane... ...years of experience developing large-scale infrastructure, distributed systems or networking, or... ...Balancing, etc.).Google's software engineers develop the next-generation technologies...Senior$163k - $237k
...qualifications:Bachelor's degree in Electrical Engineering, Computer Engineering, Computer Science... ...for all usage scenarios.The AI and Infrastructure team is redefining what’s possible. We... ...Our customers include Googlers, Google Cloud customers, and billions of Google users...SeniorWorldwide$184k - $287.5k
...advanced large language model workloads. We are looking for a Senior Software Engineer to lead the bring-up, triage, benchmarking, analysis, and... ..., validation, and debugging of large-scale AI clusters, infrastructure, and end-to-end workloads, setting the standard for how...SeniorFull timeRemote work$174k - $252k
...years of experience developing large-scale infrastructure, distributed systems or networks, or... ...technologies. Google's software engineers develop the next-generation technologies... ...and fastest experience possible.Google Cloud accelerates every organization’s ability...Senior$203.3k - $305.6k
Overview Summary The Apple Services Engineering team is one of the most exciting examples of Apple’s long-held passion... ...to a broad range of opportunities. The Apple Cloud Services Infrastructure team is seeking a senior engineering program manager to drive the...SeniorRelocation packageFlexible hours$193.93k - $291.15k
...machine-learning-first approach to autonomous driving, and the ML Infrastructure team builds and operates the infrastructure that makes that... ....About YouBS, MS, or PhD in Computer Science, Electrical Engineering, or a closely related field, plus 3+ years of relevant work...SeniorWork experience placementImmediate startFlexible hours$174k - $252k
...experience building and developing large-scale infrastructure or distributed systems.1 year of... ...(e.g., Python).Google's software engineers develop the next-generation technologies... ...to push technology forward.The Google Cloud AI Research team addresses AI challenges...Senior$160k - $260k
...) encompasses our mission and ethos for why we build.About The RoleWe're looking for a talented, driven Senior Software Engineer to build the critical infrastructure that powers the Aptos blockchain ecosystem.This role will have significant impact and visibility - the systems...SeniorFull timeWork experience placementWork at officeLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Cloud Infrastructure Engineer. Be the first to apply!
- senior cloud security engineer Santa Clara, CA
- senior cloud solutions architect Santa Clara, CA
- senior cloud data engineer Santa Clara, CA
- cloud operations engineer Santa Clara, CA
- cloud engineering manager Santa Clara, CA
- informatica cloud developer Santa Clara, CA
- senior cloud engineer Santa Clara, CA
- cloud architect Santa Clara, CA
- cloud solutions architect Santa Clara, CA
- aws cloud architect Santa Clara, CA


