Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Infrastructure Intern: Distributed Training & ML Systems

Bytedance

ByteDance Seed's Infrastructure team in San Jose is seeking an Infrastructure Intern within the Technology group. You will contribute to distributed training, RL pipelines, and high-performance AI infrastructure, gaining practical, on-the-job experience in a fast-paced environment. You will collaborate with researchers and engineers to translate model requirements into scalable system solutions while developing core skills in ML systems during this short-term internship. #J-18808-Ljbffr Bytedance

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the AI Infrastructure Intern: Distributed Training & ML Systems in San Jose, CA vacancy
  • ByteDance Seed Infrastructure Internship in San Jose invites PhD candidates to contribute to distributed training, reinforcement learning frameworks, and...  ...performance inference for AI foundation models. You’ll...  .... You will work on system-level performance analysis... 
    Internship
    Training

    ByteDance

    San Jose, CA
    1 day ago
  •  ...Jose is seeking an intern to contribute to the Seed Infrastructures team. You will...  ...infrastructure for large-scale AI foundation models,...  ...engineers across training platforms, inference systems, compilers, and distributed components. The...  ...and cutting-edge ML workloads, with... 
    Internship
    Training

    ByteDance

    San Jose, CA
    2 days ago
  •  ...Employment Type: Intern Job Code: A157139...  ...About the teamThe Seed Infrastructures team oversees the distributed training, reinforcement...  ...compilation technologies for AI foundation models....  ...training systems (e.g., data/model parallelism...  ...long-term work in ML systems or AI... 
    Internship
    Training
    Temporary work

    Bytedance

    San Jose, CA
    2 days ago
  • $40 - $85 per hour

     ...and languages. The AI Platform team builds the infrastructure that Netflix's ML and AI systems run on, from large-scale training platforms and post-training...  ...in Computer Science, Distributed Systems, Systems,...  ...personalized experience for interns, and our aim is to offer... 
    Internship
    Training
    Hourly pay
    Full time
    Immediate start
    Remote work
    Flexible hours

    Netflix

    Los Gatos, CA
    3 days ago
  • $193.3k - $261.5k

     ...Neuron includes an ML compiler, runtime,...  ...so customers can train frontier-scale models...  ...their stack.The Distributed Training team is...  ...engineers build the infrastructure that large-scale...  ..., and distributed systems, where you will help...  ...the direction of AI acceleration technology... 
    Internship
    Training
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    4 days ago
  • ByteDance\'s Seed Infrastructures team is offering a campus intern role in San Jose. You will contribute to large-scale model infrastructure, training platforms, and distributed systems while working with researchers and engineers. The role requires coursework in CS/EE,... 
    Internship
    Training

    ByteDance

    San Jose, CA
    4 days ago
  • ByteDance invites an intern to join the Seed Infrastructures team in San Jose to contribute to AI compiler optimizations for training and inference workloads. You will help...  ...hands‑on experience with distributed training, large-scale ML systems, and #J-18808-Ljbffr ByteDance
    Internship
    Training

    ByteDance

    San Jose, CA
    1 day ago
  • $184k - $287.5k

     ...powers innovative AI research and...  ...building the AI/ML platform for...  ...developing scalable AI infrastructure services...  ...large-scale AI training, inferencing, fine...  ...and improve system and service reliability...  ...large-scale distributed systems....  ...DL frameworks internal PyTorch, TensorFlow... 
    Training
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $160k - $198k

     ...You’ll DoAs a Senior AI Systems Engineer, you will architect...  ...manage the critical infrastructure services required for large-scale AI model training and inference. You...  ...services tailored for distributed AI model training and...  ...dedicated focus on AI/ML systems, high-performance... 
    Training
    Local area

    Archer Aviation

    San Jose, CA
    3 days ago
  • $190k - $237k

     ...What You’ll DoAs a Staff AI Systems Engineer, you will...  ...and manage the critical infrastructure services required for large-scale AI model training and inference. You...  ...services tailored for distributed AI model training and...  ...dedicated focus on AI/ML systems, high-performance... 
    Training
    Local area

    Archer Aviation

    San Jose, CA
    1 day ago
  •  ...computing experiences—from AI and data centers, to...  ...gaming and embedded systems. Grounded in a...  ...Scientist - Infrastructure Engineer, Reinforcement...  ...at scale—including distributed policy and value training, rollout generation,...  ...in machine learning (ML) platforms with deep... 
    Training

    AMD

    Santa Clara, CA
    4 days ago
  • $180k - $240k

     ...Senior AI Infrastructure EngineerSanta Clara, CAAbout the roleWe are...  ...infrastructure that enables distributed training, experiment tracking, and...  ...doDistributed Training & ML Systems Support Scale Research Workloads...  ...AI agents interact with internal databases and... 
    Training
    Work at office

    Gatik AI

    Santa Clara, CA
    2 days ago
  •  ...potential of generative AI to power the...  ...The role: Principal System Software Engineer,...  ...the deployment infrastructure, working closely with...  ...other software (ML and compilers) and...  ...toolsExperience with distributed, high-performance...  ..., education, and training. We also offer... 
    Training
    3 days per week

    d-Matrix

    Santa Clara, CA
    1 day ago
  • ByteDance Seed Infrastructure Intern role focuses on distributed training, RL framework work, high-performance inference, and compiler/runtime optimizations for AI foundation models. You will contribute to large-scale systems and gain hands-on experience in a fast-paced... 
    Internship
    Training

    ByteDance

    San Jose, CA
    2 days ago
  • $168k - $258.75k

     ...capacity operations for large-scale AI infrastructure programs and support tooling and...  ...closely related experience in AI/ML platforms, distributed systems, cloud infrastructure, compute capacity...  ...bring-up, or large-scale AI training and inference environments.Experience... 
    Training
    Full time

    Nvidia

    Santa Clara, CA
    19 hours ago
  • $185k - $215k

     ...you will operate as a Staff AI Systems Engineer, setting...  ...Hands-on expertise across the ML lifecycle—including dataset curation, model training, fine-tuning, optimization...  ...Experience designing and scaling distributed systems, cloud infrastructure, or orchestration tools (e... 
    Internship
    Training
    Odd job
    Work experience placement
    Local area
    Worldwide

    Bosch USA

    Sunnyvale, CA
    1 day ago
  •  ...discovery to powering AI and the...  ...and server-class systems. In this...  ...evolve the test infrastructure that gates every...  ...infrastructure — distributed test runners, GitHub...  .../ Jenkins / internal CI fleets, hardware...  ...— LLM training and inference (...  ...oneAPI, SYCL)AI/ML frameworks (PyTorch... 
    Training
    Contract work
    Shift work

    AMD

    San Jose, CA
    19 hours ago
  • $144k - $180k

     ...You’ll Do As a Senior AI Systems Engineer, you will architect...  ...manage the critical infrastructure services required for large-scale AI model training and inference. You...  ...services tailored for distributed AI model training and...  ...dedicated focus on AI/ML systems, high-performance... 
    Training
    Local area

    AlleyCorp

    San Jose, CA
    2 days ago
  • $224k - $356.5k

     ...NVIDIA’s Networking Systems & Software Architecture...  ...is solving some of AI’s hardest infrastructure problems. The team builds...  ...libraries for distributed AIDriving hardware-software...  ....Understanding of ML systems concepts—transformer...  ..., or distributed training and inference... 
    Training
    Full time

    Nvidia

    Santa Clara, CA
    19 hours ago
  • $40 - $85 per hour

     ...creative tooling, system optimization,...  ...learning infrastructure, another strong...  ...learning Reliable ML: Robustness,...  ...modeling Agentic AI: LLM agents,...  ...& Efficiency: Training/inference efficiency...  ...with distributed training/inference...  ...experience for interns, and our aim is... 
    Internship
    Training
    Hourly pay
    Full time
    Immediate start
    Relocation package
    Flexible hours

    Netflix

    Los Gatos, CA
    3 days ago
  • $198k - $326k

     ...scaling LinkedIn's AI model training, feature...  ...queries.Model Training Infrastructure: As an engineer on...  ...PyTorch, enable distributed training over 100...  ...advanced support for internal AI teams in areas...  ...hundreds of new ML models per...  ...understandability of various systems with a focus on... 
    Training
    For contractors
    Work at office
    Flexible hours

    Linkedin

    Sunnyvale, CA
    19 hours ago
  • We are looking to hire a Senior Infrastructure, SRE & AI Platforms Manager to help set the long...  ...technical domain expertise—spanning distributed systems, Kubernetes, and AI workload...  ...running large-scale systems for AI/ML training and inference workloads, including... 
    Training

    Socket.dev

    Cupertino, CA
    2 days ago
  • $153.2k - $234.1k

     ...hardware and battery systems to intuitive...  ...the Embodied AI team at General...  .... As a Senior ML Infra Engineer,...  ...dataset generation, training, evaluation and...  ...engineers and interns.What You’ll...  ...on large-scale distributed systems, applications, or ML infrastructure. Experience designing... 
    Training
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    4 days ago
  • $127.1k - $185k

     ...the network stack for EC2 distributed AI/ML systems. You'll work on software that...  ...'s largest AI models to train across massive GPU...  ...networking, and machine learning infrastructure - building the systems that...  ...Operating Systems (Linux internals, kernel concepts, memory management... 
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    15 hours ago
  • $165.2k - $223.6k

     ...Inferentia and Trainium ML accelerators....  ...inference and training performance.The...  ...systematic infrastructure, innovate new...  ...'s possible in AI acceleration.As...  ...computing, and distributed architectures,...  ...the stack from system level optimizations...  ...with internal and external stakeholders... 
    Internship
    Training
    Work experience placement
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  •  ...Technology Employment Type: Intern Job Code: A104504A Share...  ...About the teamThe Seed Infrastructures team oversees the distributed training, reinforcement learning...  ...technologies for AI foundation models.Responsibilities...  ...to infrastructure and systems for large-scale models.-... 
    Internship
    Training

    Bytedance

    San Jose, CA
    2 days ago
  • $198k - $326k

     ...the team. LinkedIn’s AI Infrastructure organization is responsible...  ...layer between model training and production serving...  ...the intersection of systems, machine learning, GPU...  ...improvementsPartner closely with ML, infrastructure, and...  ...software engineering, distributed systems,... 
    Training
    For contractors
    Work at office
    Flexible hours

    Linkedin

    Mountain View, CA
    19 hours ago
  • $136.3k - $231.7k

     ...us. KLA invents systems and solutions...  ...developing scalable distributed systems or...  ...Intelligence (AI) and Machine Learning (ML) technologies, frameworks, or infrastructure.Experience leveraging...  ...bonding leave.Interns are eligible...  ...education level or training. We are... 
    Training
    Minimum wage
    Full time
    Flexible hours

    KLA-Tencor

    Milpitas, CA
    3 days ago
  • $250k - $350k

     ...era of pervasive AI has arrived. In this...  ...the team The System Architecture team...  ...engineers, SambaFlow, ML performance, and datacenter...  ...RDUs into larger training and inference...  ...density/48V-class distribution) and the cooling...  ...hyperscale, HPC, or AI/ML infrastructure Deep, hands-on... 
    Training
    Full time
    Temporary work
    Local area
    Flexible hours

    SambaNova Systems

    San Jose, CA
    3 days ago
  • $132.4k - $217.6k

     ...are embedding AI-driven intelligence...  ...-grade systems used in regulated...  ...Enterprise Data, and Infrastructure teams to align...  ...performance of distributed services...  ...considerations for internal and cross-functional...  ...Science, AI/ML, or related...  ...engineering, model training/tuning,... 
    Internship
    Training
    Full time
    Temporary work
    Worldwide
    Flexible hours

    Bio-Techne

    San Jose, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Infrastructure Intern: Distributed Training & ML Systems. Be the first to apply!