AI Infrastructure Intern: Distributed Training & ML Systems
Bytedance
ByteDance Seed's Infrastructure team in San Jose is seeking an Infrastructure Intern within the Technology group. You will contribute to distributed training, RL pipelines, and high-performance AI infrastructure, gaining practical, on-the-job experience in a fast-paced environment. You will collaborate with researchers and engineers to translate model requirements into scalable system solutions while developing core skills in ML systems during this short-term internship. #J-18808-Ljbffr Bytedance
- ByteDance Seed Infrastructure Internship in San Jose invites PhD candidates to contribute to distributed training, reinforcement learning frameworks, and... ...performance inference for AI foundation models. You’ll... .... You will work on system-level performance analysis...InternshipTraining
- ...Jose is seeking an intern to contribute to the Seed Infrastructures team. You will... ...infrastructure for large-scale AI foundation models,... ...engineers across training platforms, inference systems, compilers, and distributed components. The... ...and cutting-edge ML workloads, with...InternshipTraining
- ...Employment Type: Intern Job Code: A157139... ...About the teamThe Seed Infrastructures team oversees the distributed training, reinforcement... ...compilation technologies for AI foundation models.... ...training systems (e.g., data/model parallelism... ...long-term work in ML systems or AI...InternshipTrainingTemporary work
$40 - $85 per hour
...and languages. The AI Platform team builds the infrastructure that Netflix's ML and AI systems run on, from large-scale training platforms and post-training... ...in Computer Science, Distributed Systems, Systems,... ...personalized experience for interns, and our aim is to offer...InternshipTrainingHourly payFull timeImmediate startRemote workFlexible hours$193.3k - $261.5k
...Neuron includes an ML compiler, runtime,... ...so customers can train frontier-scale models... ...their stack.The Distributed Training team is... ...engineers build the infrastructure that large-scale... ..., and distributed systems, where you will help... ...the direction of AI acceleration technology...InternshipTrainingLocal areaFlexible hours- ByteDance\'s Seed Infrastructures team is offering a campus intern role in San Jose. You will contribute to large-scale model infrastructure, training platforms, and distributed systems while working with researchers and engineers. The role requires coursework in CS/EE,...InternshipTraining
- ByteDance invites an intern to join the Seed Infrastructures team in San Jose to contribute to AI compiler optimizations for training and inference workloads. You will help... ...hands‑on experience with distributed training, large-scale ML systems, and #J-18808-Ljbffr ByteDanceInternshipTraining
$184k - $287.5k
...powers innovative AI research and... ...building the AI/ML platform for... ...developing scalable AI infrastructure services... ...large-scale AI training, inferencing, fine... ...and improve system and service reliability... ...large-scale distributed systems.... ...DL frameworks internal PyTorch, TensorFlow...TrainingFull timeRemote work$160k - $198k
...You’ll DoAs a Senior AI Systems Engineer, you will architect... ...manage the critical infrastructure services required for large-scale AI model training and inference. You... ...services tailored for distributed AI model training and... ...dedicated focus on AI/ML systems, high-performance...TrainingLocal area$190k - $237k
...What You’ll DoAs a Staff AI Systems Engineer, you will... ...and manage the critical infrastructure services required for large-scale AI model training and inference. You... ...services tailored for distributed AI model training and... ...dedicated focus on AI/ML systems, high-performance...TrainingLocal area- ...computing experiences—from AI and data centers, to... ...gaming and embedded systems. Grounded in a... ...Scientist - Infrastructure Engineer, Reinforcement... ...at scale—including distributed policy and value training, rollout generation,... ...in machine learning (ML) platforms with deep...Training
$180k - $240k
...Senior AI Infrastructure EngineerSanta Clara, CAAbout the roleWe are... ...infrastructure that enables distributed training, experiment tracking, and... ...doDistributed Training & ML Systems Support Scale Research Workloads... ...AI agents interact with internal databases and...TrainingWork at office- ...potential of generative AI to power the... ...The role: Principal System Software Engineer,... ...the deployment infrastructure, working closely with... ...other software (ML and compilers) and... ...toolsExperience with distributed, high-performance... ..., education, and training. We also offer...Training3 days per week
- ByteDance Seed Infrastructure Intern role focuses on distributed training, RL framework work, high-performance inference, and compiler/runtime optimizations for AI foundation models. You will contribute to large-scale systems and gain hands-on experience in a fast-paced...InternshipTraining
$168k - $258.75k
...capacity operations for large-scale AI infrastructure programs and support tooling and... ...closely related experience in AI/ML platforms, distributed systems, cloud infrastructure, compute capacity... ...bring-up, or large-scale AI training and inference environments.Experience...TrainingFull time$185k - $215k
...you will operate as a Staff AI Systems Engineer, setting... ...Hands-on expertise across the ML lifecycle—including dataset curation, model training, fine-tuning, optimization... ...Experience designing and scaling distributed systems, cloud infrastructure, or orchestration tools (e...InternshipTrainingOdd jobWork experience placementLocal areaWorldwide- ...discovery to powering AI and the... ...and server-class systems. In this... ...evolve the test infrastructure that gates every... ...infrastructure — distributed test runners, GitHub... .../ Jenkins / internal CI fleets, hardware... ...— LLM training and inference (... ...oneAPI, SYCL)AI/ML frameworks (PyTorch...TrainingContract workShift work
$144k - $180k
...You’ll Do As a Senior AI Systems Engineer, you will architect... ...manage the critical infrastructure services required for large-scale AI model training and inference. You... ...services tailored for distributed AI model training and... ...dedicated focus on AI/ML systems, high-performance...TrainingLocal area$224k - $356.5k
...NVIDIA’s Networking Systems & Software Architecture... ...is solving some of AI’s hardest infrastructure problems. The team builds... ...libraries for distributed AIDriving hardware-software... ....Understanding of ML systems concepts—transformer... ..., or distributed training and inference...TrainingFull time$40 - $85 per hour
...creative tooling, system optimization,... ...learning infrastructure, another strong... ...learning Reliable ML: Robustness,... ...modeling Agentic AI: LLM agents,... ...& Efficiency: Training/inference efficiency... ...with distributed training/inference... ...experience for interns, and our aim is...InternshipTrainingHourly payFull timeImmediate startRelocation packageFlexible hours$198k - $326k
...scaling LinkedIn's AI model training, feature... ...queries.Model Training Infrastructure: As an engineer on... ...PyTorch, enable distributed training over 100... ...advanced support for internal AI teams in areas... ...hundreds of new ML models per... ...understandability of various systems with a focus on...TrainingFor contractorsWork at officeFlexible hours- We are looking to hire a Senior Infrastructure, SRE & AI Platforms Manager to help set the long... ...technical domain expertise—spanning distributed systems, Kubernetes, and AI workload... ...running large-scale systems for AI/ML training and inference workloads, including...Training
$153.2k - $234.1k
...hardware and battery systems to intuitive... ...the Embodied AI team at General... .... As a Senior ML Infra Engineer,... ...dataset generation, training, evaluation and... ...engineers and interns.What You’ll... ...on large-scale distributed systems, applications, or ML infrastructure. Experience designing...TrainingFull timeLocal areaRemote workWork from homeRelocation packageFlexible hours$127.1k - $185k
...the network stack for EC2 distributed AI/ML systems. You'll work on software that... ...'s largest AI models to train across massive GPU... ...networking, and machine learning infrastructure - building the systems that... ...Operating Systems (Linux internals, kernel concepts, memory management...InternshipLocal areaFlexible hours$165.2k - $223.6k
...Inferentia and Trainium ML accelerators.... ...inference and training performance.The... ...systematic infrastructure, innovate new... ...'s possible in AI acceleration.As... ...computing, and distributed architectures,... ...the stack from system level optimizations... ...with internal and external stakeholders...InternshipTrainingWork experience placementLocal areaFlexible hours- ...Technology Employment Type: Intern Job Code: A104504A Share... ...About the teamThe Seed Infrastructures team oversees the distributed training, reinforcement learning... ...technologies for AI foundation models.Responsibilities... ...to infrastructure and systems for large-scale models.-...InternshipTraining
$198k - $326k
...the team. LinkedIn’s AI Infrastructure organization is responsible... ...layer between model training and production serving... ...the intersection of systems, machine learning, GPU... ...improvementsPartner closely with ML, infrastructure, and... ...software engineering, distributed systems,...TrainingFor contractorsWork at officeFlexible hours$136.3k - $231.7k
...us. KLA invents systems and solutions... ...developing scalable distributed systems or... ...Intelligence (AI) and Machine Learning (ML) technologies, frameworks, or infrastructure.Experience leveraging... ...bonding leave.Interns are eligible... ...education level or training. We are...TrainingMinimum wageFull timeFlexible hours$250k - $350k
...era of pervasive AI has arrived. In this... ...the team The System Architecture team... ...engineers, SambaFlow, ML performance, and datacenter... ...RDUs into larger training and inference... ...density/48V-class distribution) and the cooling... ...hyperscale, HPC, or AI/ML infrastructure Deep, hands-on...TrainingFull timeTemporary workLocal areaFlexible hours$132.4k - $217.6k
...are embedding AI-driven intelligence... ...-grade systems used in regulated... ...Enterprise Data, and Infrastructure teams to align... ...performance of distributed services... ...considerations for internal and cross-functional... ...Science, AI/ML, or related... ...engineering, model training/tuning,...InternshipTrainingFull timeTemporary workWorldwideFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Infrastructure Intern: Distributed Training & ML Systems. Be the first to apply!
- internship machine learning San Jose, CA
- machine learning scientist San Jose, CA
- machine learning research scientist San Jose, CA
- data engineer machine learning San Jose, CA
- machine learning part time San Jose, CA
- artificial intelligence - machine learning intern San Jose, CA
- machine learning San Jose, CA
- machine learning remote San Jose, CA
- machine learning intern San Jose, CA
- machine learning researcher San Jose, CA


