Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

LLM Training & Inference Scientist (GPU-Optimized)

ByteDance

ByteDance is seeking talented PhD candidates to develop and optimize machine learning training frameworks in San Jose, California. As part of AML-MLsys, you'll work on cutting-edge AI projects involving large language models (LLMs). The role requires proficiency in algorithms, data structures, and deep learning principles, alongside knowledge of Python and CUDA. Employees will enjoy comprehensive benefits including medical, dental, and vision insurance from day one. #J-18808-Ljbffr ByteDance

Vacancy posted 23 hours ago
Similar jobs that could be interesting for youBased on the LLM Training & Inference Scientist (GPU-Optimized) in San Jose, CA vacancy
  •  ...seeking a Senior Research Scientist to lead machine learning system...  .... You will help build and optimize large-scale distributed ML training and inference infrastructure, integrating GPU/NPU/RDMA and storage to run...  ...contribute to cutting-edge LLM/AI system work while... 
    Training

    ByteDance

    San Jose, CA
    4 days ago
  • $212.8k

     ...maintain massively distributed ML training and inference systems around the world,...  ..., and scalable systems for LLM/AIGC/AGI. Topic Content...  ...Responsibilities Develop and optimize LLM training & inference & reinforcement...  ...the next level. Optimize GPU and CUDA performance to... 
    Training
    Temporary work
    Local area
    Shift work

    ByteDance

    San Jose, CA
    2 days ago
  • d-Matrix in Santa Clara, CA is seeking a Sr. Staff ML Researcher to advance LLM algorithmic optimization on our DNN accelerators. You will design and implement efficient inference algorithms, collaborating with mathematicians, ML researchers and engineers on high-impact... 
    Suggested
    3 days per week

    Entrada Ventures

    Santa Clara, CA
    2 days ago
  •  ...The role: Sr. Staff, ML Researcher - LLM Algorithmic OptimizationWhat You...  ...efficient algorithms that will be used to optimize large language model inference on DNN accelerators we develop. You...  ...skills, experience, education, and training. We also offer incentive... 
    Training
    3 days per week

    d-Matrix

    Santa Clara, CA
    3 days ago
  • ByteDance’s Seed LLM team is advancing the next generation of large language models, tackling pre-training, post-training, inference, memory, learning, and interpretability. The role focuses...  ..., instruction tuning, and optimization to improve reasoning and related abilities... 
    Training

    ByteDance

    San Jose, CA
    4 days ago
  • $244.8k

    The Seed LLM team is dedicated to advancing the next generation of large language...  ...development. It focuses on model pre‑training, post‑training, inference, memory capabilities, learning,...  ...Responsibilities Explore large‑scale models and optimize associated systems. Perform data... 
    Training
    Temporary work
    Local area

    ByteDance

    San Jose, CA
    2 days ago
  • $212.8k - $387.6k

    Responsibilities The Seed LLM team is dedicated to advancing the...  ...pretraining, posttraining, inference, memory capabilities, learning...  ...large-scale models and optimize systems. Data construction, instruction...  ...in LLM-related areas such as training, evaluation, or applications.... 
    Training
    Temporary work
    Internship

    ByteDance

    San Jose, CA
    3 days ago
  • $187.04k - $359.72k

    Senior Research Scientist, Foundation Model (LLM/ VLLM), TikTok - Trust and Safety Senior...  ...II Do 1. Lead the design, training, and deployment of...  ...data, and platform teams to optimize large-scale training pipelines...  ...and leverage centralized GPU resources effectively. 5.... 
    Training
    Full time
    Temporary work
    Local area

    TikTok

    San Jose, CA
    2 days ago
  • $244.8k

    Research Engineer - LLM/VLM Inference Optimization (Seed Infra) Location San Jose Team Technology Employment...  ...team oversees the distributed training, reinforcement learning framework, high...  ..., or serving cost. Familiarity with GPU architecture and experience optimizing... 
    Training
    Temporary work
    Local area

    Bytedance

    San Jose, CA
    5 days ago
  • $192k - $304.75k

     ...Deep Learning Research Scientist, Efficiency!Join our...  ...and algorithms to optimize neural networks for training and deployment....  ...effect on neural network inference and training...  ...architectures, optimizers and LLM training.Experience...  ...with GPU computing, kernels,... 
    Training
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $224k - $356.5k

     ...era in which our GPU acts as the brains...  ...Applied Research Scientist with experience researching...  ...foundation-model training, including...  ...training of Nemotron LLM models which lead...  ...the expansion and optimization of curation methodologies...  ...of NVIDIA Inference Microservices (NIMs... 
    Training
    Full time
    Work at office
    Remote work
    Flexible hours

    Nvidia

    Santa Clara, CA
    2 days ago
  • $171.6k - $222.2k

     ...assembling an elite team of world-class scientists and engineers to pioneer the next...  ...development tools. Join the Amazon Kiro LLM-Training team and help create groundbreaking generative...  ...data structures, parsing, numerical optimization, data mining, parallel and distributed... 
    Training
    Work at office
    Local area
    Worldwide
    Flexible hours

    Amazon

    Santa Clara, CA
    9 hours ago
  • ByteDance is seeking a Senior Research Scientist for PICO in San Jose to advance multimodal large language...  ...) for mixed reality. You will lead architecture optimization, data construction, and end-to-end training/inference acceleration tailored to MR scenarios. The role... 
    Training

    ByteDance

    San Jose, CA
    4 days ago
  • ByteDance is seeking talented graduate students for a role focusing on LLM training and inference. Successful candidates will develop high-performance, scalable ML systems while collaborating closely with model researchers. Applicants need to be pursuing a Ph.D. in a related... 
    Training

    ByteDance

    San Jose, CA
    2 days ago
  • $212.8k - $387.6k

    Senior Research Scientist - Machine Learning System...  ...massively distributed ML training and Inference system/services...  ...scalable systems for LLM/AIGC/AGI. In our team...  ...system integrating with GPU/NPU/RDMA/Storage and...  ...Responsible for developing and optimizing LLM inference... 
    Training
    Temporary work
    Local area

    ByteDance

    San Jose, CA
    4 days ago
  • $212.8k - $387.6k

     ...control plane for large‑scale LLM inference. We are building the next‑...  ...networking, and storage—covering GPU virtualization, RDMA/DPDK‑...  ...infrastructure that can self‑optimize via AI (AI for systems) Driving...  ...such as large‑scale model training and inference, heterogeneous... 
    Training
    Temporary work
    Local area

    ByteDance

    San Jose, CA
    5 days ago
  •  ...large-scale MoE architecture training and routing optimization, cross-modal alignment and...  ...of cutting-edge LLM research (e.g., long context...  ...verification for training/finetuning/inference; Being familiar with PEFT,...  ...Have a deep understanding of GPU and/or other AI... 
    Training
    Flexible hours
    Shift work

    SpeedyApply LLC

    San Jose, CA
    23 hours ago
  • $60 per hour

     ...and retrieval to ranking and user experience optimization. Project Overview, Challenges & Value With the rapid evolution of pre-trained large models, TikTok Search is advancing...  ...search models (video, text, image). Explore LLM-based approaches for complex and multi-turn... 
    Training
    Hourly pay
    Summer work
    Internship
    Local area

    TikTok

    San Jose, CA
    2 days ago
  • $145k - $250k

    Research Scientist Graduate- CV/NLP/Multimodal LLM,(Trust and Safety) - 2026 Start(PhD) Join to apply for the Research Scientist Graduate...  ...in TikTok, and will also be responsible for optimizing our distributed model training framework continuously. Responsibilities... 
    Training
    Full time
    Temporary work
    Fixed term contract
    Summer work
    Internship
    Local area
    Flexible hours

    TikTok

    San Jose, CA
    2 days ago
  • $60 per hour

     ...Encode hundreds of millions of items and massive user behavior data for large‑scale training, optimizing via advanced architectures, post‑training (e.g., RLVR), and efficient inference. Multimodal Large Models for E‑commerce: Build multilingual, multimodal models achieving... 
    Training
    Hourly pay
    Internship
    Local area

    TikTok

    San Jose, CA
    2 days ago
  •  ...Clara, CA. This hybrid position involves working on-site 3 days a week. The successful candidate will develop algorithms for optimizing LLM inference on our DNN accelerators. Ideal applicants should have a MSc or PhD in a relevant field and 5+ years of experience in... 
    3 days per week

    d-Matrix inc.

    Santa Clara, CA
    1 day ago
  • $156k - $316.8k

    Research Scientist — Privacy-Preserving Large-Scale Model Training & Architecture Optimization Location: San Jose Employment Type: Regular Job Code: DW1L Responsibilities Design...  ...Flow, hybrid AR + diffusion systems). Lead GPU-centric performance optimization, including... 
    Training
    Temporary work
    Local area

    Ellis Technologies, Inc.

    San Jose, CA
    2 days ago
  •  ...3 days per week. The role Senior Staff ML Researcher - LLM Algorithmic Optimization What You Will Do d-Matrix is seeking machine learning researchers...  ...that will be used to optimize large language model inference on DNN accelerators we develop. You would be part of a... 
    3 days per week

    d-Matrix inc.

    Santa Clara, CA
    1 day ago
  •  ...headquarters 3 days per week. The role: Sr. Staff, ML Researcher - LLM Algorithmic Optimization What You Will Do d-Matrix is seeking machine learning...  ...that will be used to optimize large language model inference on DNN accelerators we develop. You would be part of a close... 
    3 days per week

    Entrada Ventures

    Santa Clara, CA
    2 days ago
  • $168k - $264.5k

     ...looking for a Senior Scientist to join our team and help...  ...data generation for training frontier models. You will...  ...pipelines using LLM-based methods and automated...  ...post-training, and inference frameworks such as vLLM...  ...or audio).Building and optimizing scalable data... 
    Training
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $192k - $304.75k

     ...looking for a passionate scientist at the intersection of...  ...will span synthetic training data generation, surrogate modeling, and co-optimized calibration-decoding...  ...prediction and parameter inference without full...  ...and modalities.Develop GPU-accelerated implementations... 
    Training
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...Our core invention, the GPU, serves as the visual...  ...hiring Senior Deep Learning Scientists to advance our efforts...  ...come join our Nemotron LLM team. For more details...  ...research to develop, train, fine-tune, and deploy...  ...instruction tuning, preference optimization, and RLHF/RLVR/MOPD to... 
    Training
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    2 days ago
  • $171.6k - $222.2k

     ...development with systems and hardware optimization.QualificationsStrong...  ...similar frameworks.Experience training, fine-tuning, or serving large-scale models.Knowledge of GPU architecture, CUDA, Triton,...  ...execution.Optimize training and inference using CUDA, Triton, custom... 
    Training
    Local area
    Flexible hours

    AmazonWebServices

    Santa Clara, CA
    2 days ago
  •  ...edge of what’s possible with LLM inference on heterogeneous hardware....  ...deployment patterns to deep optimization of inference kernels, to building...  ...level.• Familiarity with GPU kernel programming (CUDA/...  ...understanding of modern LLM training and inference.Why D-Matrix Frontier... 
    Training

    d-Matrix

    Santa Clara, CA
    9 hours ago
  • $195.2k - $361.2k

     ...hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for...  ...and edge environments — GPU/iGPUs, Vulkan backends — not...  ...quality impact with the Post-Training teamCut CPU overhead and improve...  ...-level codeExperience with LLM inference. (attention, KV cache... 
    Training
    Full time
    Internship
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to LLM Training & Inference Scientist (GPU-Optimized). Be the first to apply!