Senior AI Infrastructure Engineer - Model Training
$190k - $260kKodiak Robotics
Kodiak Robotics, Inc. was founded in 2018 and has become a leader in autonomous ground transportation committed to a safer and more efficient future for all. The company has developed an artificial intelligence (AI) powered technology stack purpose-built for commercial trucking and the public sector. The company delivers freight daily for its customers across the southern United States using its autonomous technology. In 2024, Kodiak became the first known company to publicly announce delivering a driverless semi-truck to a customer. Kodiak is also leveraging its commercial self-driving software to develop, test and deploy autonomous capabilities for the U.S. Department of Defense.Kodiak's AI is only as good as the speed at which we can train it. Every improvement to our models – from GigaFusionNet to large-scale world models – depends on infrastructure that turns thousands of hours of multimodal driving data into training throughput. We are looking for engineers who make model training fast: streaming massive camera, LiDAR, and radar datasets without stalling a single GPU, sharding data and models efficiently across nodes, and extracting every FLOP from the latest hardware. If you measure your impact in tokens per second and GPU utilization, this role is for you.In this role, you will:Design high-throughput data loading and streaming systems for multimodal sensor data (camera, LiDAR, radar), including dataset formats, sharding strategies, and prefetching pipelines that keep GPUs saturatedBuild and optimize distributed training infrastructure across multi-node GPU clusters, applying data, tensor, pipeline, and fully sharded (FSDP/ZeRO) parallelism to models that don't fit on a single deviceMaximize utilization of modern accelerators such as NVIDIA B200s through mixed-precision training (BF16/FP8), fused kernels, memory optimization, and communication/computation overlapProfile end-to-end training pipelines to find and eliminate bottlenecks across storage, network, CPU preprocessing, and GPU computeDevelop scalable dataset construction pipelines that convert petabytes of raw driving logs into training-ready, streamable formatsPartner with ML teams to scale new architectures from prototype to full-cluster training runs efficiently and reliablyWhat you’ll bring:BS, MS, or PhD in Computer Science or a related field, and at least 2-3 years of industry experience in ML systems or infrastructureHands-on experience with distributed training frameworks and techniques (PyTorch DDP/FSDP, DeepSpeed, Megatron, NCCL) and a strong grasp of parallelism trade-offsExperience building high-performance data pipelines for large-scale training, including streaming dataset formats (WebDataset, MosaicML Streaming/MDS, or similar), sharding, and storage/network-aware loadingDeep understanding of GPU performance: mixed precision, memory hierarchy, kernel fusion, profiling tools (Nsight, PyTorch Profiler), and interconnects (NVLink, InfiniBand)Strong Python skills and proficiency in PyTorch internals; systems-level experience (C++/CUDA/Triton) a plusPassion for building the infrastructure that lets AI for the physical world train faster, scale further, and improve continuouslyWhat we offer:Competitive compensation package including equity and annual bonusesExcellent Medical, Dental, and Vision plans through Kaiser Permanente, Cigna, and MetLife (including a medical plan with infertility benefits)MetLife Legal Services, Identity & Fraud Protection, Hospital Indemnity Insurance, Accident Insurance, & Critical Illness InsuranceFlexible PTO, 10 paid holidays, and generous parental leave policiesOur office is centrally located in Mountain View, CAOffice perks: dog-friendly, free catered lunch, a fully stocked kitchen, and free EV chargingLong Term Disability, Short Term Disability, Life InsuranceWellbeing Benefits - Headspace through Cigna, Calm through Kaiser, One Medical, Gympass, Spring Health through Cigna, Rula (mental health navigation) Fidelity 401(k)Commuter, FSA, Dependent Care FSA, HSAVarious incentive programs (referral bonuses, patent bonuses, etc.)The pay range listed below reflects the base salary in our SF/Silicon Valley location, across several internal levels. Actual starting pay will be based on job-related factors including: work location, experience, relevant training, education, skill level and performance during interview. Total compensation at Kodiak includes base pay, equity, bonus and a competitive benefits packageCalifornia Pay Range$190,000—$260,000 USDAt Kodiak, we strive to build a diverse community working towards our common company goals in a safe and collaborative environment where harassment of any kind is strictly prohibited. Kodiak is committed to equal opportunity employment regardless of race, ethnicity, religion, gender identity, sexual orientation, age, disability, or veteran status, or any other basis protected by applicable law.In alignment with its business operations, Kodiak adheres to all relevant statutes, regulations, and administrative prerequisites. Accordingly, roles that carry more sensitive requirements may be limited to candidates that can satisfy additional scrutiny and eligibility for such positions may hinge on verification of a candidate’s residence, U.S. person status, and/or citizenship status. Should the position require, and Kodiak determines that a candidate’s residence, U.S. person status, and/or citizenship status necessitate an export license, bar the candidate from the position, or otherwise fall under national security-related restrictions, Kodiak will consider the candidate for alternative positions unaffected by such restrictions, under terms and conditions set forth at Kodiak’s sole discretion, or, as an alternative, opt not to proceed with the candidate’s application. If applicable, Kodiak may provide visa sponsorship for eligible candidates.We use a third-party AI tool (Endorsed) to assist in the initial screening of applications. As part of the evaluation process, we provide Endorsed with job requirements and candidate-submitted applications. Final hiring decisions are made by our human recruitment team, and no automated system makes the ultimate decision regarding hiring. Certain features of the platform may qualify it as an Automated Employment Decision Tool (AEDT) under applicable regulations. We began using Endorsed on January 1, 2026. You can review the independent bias audit report covering our use of Endorsed [here](). By submitting your application, you acknowledge that your application may be processed by AI systems as part of the screening and selection process. If you have any questions or would like to request a separate review of your application, please contact View email address on click.appcast.io with "Separate Review Request" in the email subject line.
$174k - $252k
...developing large-scale infrastructure or distributed systems... ...Java.Experience in ML model coding languages (e.g.... ...).Google's software engineers develop the next-generation... ....The Google Cloud AI Research team addresses... ...relevant education or training. US: $174000 - $252000...SeniorTraining- ...patients worldwide.We’re a team of engineers, clinicians, and innovators... ...robotic platforms. As a Senior AI/ML Research Engineer, you will... ...and fine-tune the foundation models—VFMs, VLMs, and VLA models—that... ...stable.Build and maintain training and data pipelines that combine...SeniorTrainingLocal areaWorldwideFlexible hours
$160k - $225k
Databricks is hiring a Senior Software Engineer for AI Runtime in Mountain View, California. In this role... ...and evolution of AIR's managed GPU training platform, ensuring scalable and... ...knowledge of GPU performance and training infrastructure. This position comes with a...SeniorTraining$151.3k - $283.8k
...advancements such as cloud, AI, and network security.... ...context of Large Language Model (LLM) inference and training.2.Operator & Performance... ...technologies within cloud infrastructure.Who We Look For1.... ...Ph.D. degree in Computer Engineering, Electronic Engineering,...SeniorTrainingFull timeRelocation package$184k - $287.5k
...that powers innovative AI research and... ...developing scalable AI infrastructure services globally. We... ...infrastructure software engineer to join our team. You'... ...enable large-scale AI training, inferencing, fine-tuning... ...AI in production.As a senior DGX Cloud AI Infrastructure...SeniorTrainingFull timeRemote work$184k - $287.5k
Joining NVIDIA's DGX Cloud AI Efficiency Team means contributing to the infrastructure that powers our... ...of AI workloads - pre-training, post-training, inference... ...infrastructure software engineer to join our team. You'll... ...availability of AI systems.As a senior DGX Cloud AI...SeniorTrainingFull timeRemote work$180k - $240k
...About the role We are seeking a Senior AI Infrastructure Engineer to design, build, and scale the high... ...powering our autonomous driving models. While researchers focus on developing... ...infrastructure that enables distributed training, experiment tracking, and seamless...SeniorTrainingOdd jobWork at office- NVIDIA is seeking senior engineers to advance its AI platform, focusing on performance optimizations in deep learning frameworks using JAX. The role... ...modular, fast, and coordinated tooling to handle data, training and analysis for diverse DL solutions. Strong programming...SeniorTraining
$174.72k - $295.68k
...innovation, integrating advanced AI and autonomous driving... ....As a core member of our AI Infrastructure team, you will be responsible... ...preprocessing dataset production model training / simulation input. In... ...Computer Science, Software Engineering, Artificial Intelligence, or...SeniorTrainingFull timeOverseas$209k - $238.5k
Senior Lead AI Engineer, Gen AI Platform Overview: At Capital One, we are... ...investments in technology infrastructure and world-class talent —... ...millions of customers. Our AI models and platforms empower... ...including foundation model training, large language model inference...SeniorTrainingFull timePart timeLocal area$224k - $356.5k
NVIDIA is searching for a senior or principal engineer who specializes in building cutting-edge infrastructure for large-scale foundation model training in the Generalist Embodied Agent Research (GEAR... ...-scale robot learning, embodied AI, and physics simulation. Our past projects...SeniorTrainingFull time- ...Systems builds the world's largest AI chip, 56 times larger than... ...to deliver industry-leading training and inference speeds; over 10... ...Cerebras works with the leading model labs, global enterprises, and... ...a loop." You'll sit between engineering, product, and customer-facing...TrainingFull time
$229.9k - $262.4k
...Overview Senior Lead AI Engineer (GenAI Platform Services) Overview:... ...investments in technology infrastructure and world-class talent — along... ...of customers. Our AI models and platforms empower teams... ...including foundation model training, large language model inference...SeniorTrainingFull timePart timeLocal area$204k - $259k
...and deploying advanced ML models that interpret traffic lights... ..., you will report to the Senior Engineering Manager of Semantics. You... ...infra. Collaborate with ML infrastructure teams and the modeling team... ...location, experience, relevant training and education, and skill...SeniorTrainingFull timeWork at officeRemote work$165k - $185k
...part of the global research, our AI research in Silicon Valley focuses on Foundation Models, Big Data Visual Analytics,... ...Robotics, Data Science, AI System Engineering, Time-series Analysis. We... ...on foundation models, including training, fine-tuning, and promptingIn-depth...SeniorTrainingWork experience placementWorldwide$200k - $400k
...at scale. At Scout AI, we’re developing Fury... ...robotic foundation model for defense, to give... ...We're looking for a Senior or Staff AI Engineer to join the Fury Orchestration... ...the stack: model training and evaluation,... ...machine learning infrastructure ~ Experience training...SeniorTrainingFull timeRelocation package$117.7k - $221.4k
...machine learning, data infrastructure, and developer... ...important scenarios, prepare training-ready data, and support... ...for embodied AI systems. We believe the... ...not only on stronger models, but also on better infrastructure... ...reflects how Cola engineers think: build durable intermediate...TrainingFull timeLocal areaRemote workWork from homeRelocation packageFlexible hours$230.95k
Crusoe is looking for a Senior Software Engineer to join their AI Model Lifecycle team in Sunnyvale, California. This role involves building a platform for application development, focusing on Machine Learning models, including Large Language Models (LLMs). The ideal candidate...Senior- A tech company specializing in AI infrastructure is seeking a Software Engineer to build a scalable compute platform for its generative video models. The ideal candidate will have over 5 years of experience in MLOps or AI infrastructure management, along with strong Python...Senior
$144k - $236k
...needs of the team.Responsibilities: AI is at the core of how LinkedIn connects... ..., content, and trust platforms. As a Senior AI Software Engineer you will own end-to-end machine... ...production at LinkedIn scale. You won't just train models, you will own the recommender and...SeniorTrainingFor contractorsWork at officeImmediate startFlexible hours- As a Senior AI Engineer at Intapp, you’ll work on an AI-powered, industry-specific cloud platform... ...and orchestration systems, not training models, working on LLM research, or... ...of how to deploy and work in a cloud infrastructure Bachelor's or advanced degree in engineering...SeniorTrainingFull timeWork at officeLocal areaFlexible hours
- ...computing experiences—from AI and data centers, to... ...AI / ML Platform Engineers to build the platform... ...This role focuses on the infrastructure and platform systems that... ..., distributed training and inference, experiment... ...kernel benchmarking, model serving, or distributed...SeniorTraining
- ...Systems builds the world's largest AI chip, 56 times larger than... ...to deliver industry-leading training and inference speeds; over 10... ...Cerebras works with the leading model labs, global enterprises, and... ...a highly skilled WAN Network Engineer to design, implement, manage,...SeniorTraining
$227.5k - $300k
...driving the transformation to AI-enabled software-defined... ...Summary: We are seeking a Senior Staff AI Engineer with a combination of... ...prompt registry allowing for model-agnostic routing and A/B testing... ...development, including modeling, training, tuning, validating,...SeniorTrainingFull timeWork at officeWorldwideFlexible hoursShift work$120k - $160k
...information powered by advanced AI, recommendation systems,... ...mission: building the infrastructure layer for content... ...RoleWe're looking for a Senior Technical Recruiter, AI & Engineering to help build the teams... ...and relevant education or training. At NewsBreak, we design...SeniorTrainingFull timeWork at officeLocal areaWork from homeMonday to Friday$227k - $300k
...driving the transformation to AI-enabled software-defined... ...We are looking for a great Senior Staff AI Engineer to join our seasoned AI team... ...you will build and deploy AI models that analyze continuous data... ...from data ingestion and model training to deployment on resource-constrained...SeniorTrainingWork at officeWorldwideFlexible hoursShift work3 days per week- ...the best job for you. Role: Senior Agentic AI Engineer (Palo Alto Networks Ecosystem)... ...Unlike traditional AI roles focused on model training, this position is centered on... ...agent-to-tool communication. Infrastructure: Hands-on experience with Kubernetes...SeniorTrainingPermanent employmentContract workRemote work
- Google DeepMind in Mountain View seeks a Senior Product Manager to be embedded in the research and model training process for Gemini models. You will participate in evaluations, read outputs, and make judgment calls on quality alongside researchers. You will translate...SeniorTraining
$174k - $252k
...with developing large-scale infrastructure, distributed systems or networks... ....Google's software engineers develop the next-generation technologies... ...to projects enabling AI networking and high-performance... ...experience, and relevant education or training. US: $174000 - $252000 (USD)...SeniorTraining- A leading tech company is seeking a Senior Software Engineer for AI and Infrastructure. The ideal candidate will possess strong programming expertise in C++, Java, or Python, with a focus on software design and architecture. Responsibilities include writing and testing...SeniorFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior AI Infrastructure Engineer - Model Training. Be the first to apply!
- machine learning ai engineer Mountain View, CA
- ai developer Mountain View, CA
- senior ai engineer Mountain View, CA
- ai engineer Mountain View, CA
- ai ml engineer Mountain View, CA
- ai prompt engineer Mountain View, CA
- ai engineer remote Mountain View, CA
- data infrastructure engineer Mountain View, CA
- infrastructure engineering manager Mountain View, CA
- senior infrastructure engineer Mountain View, CA



