LLM Training & Inference Scientist (GPU-Optimized)
ByteDance
ByteDance is seeking talented PhD candidates to develop and optimize machine learning training frameworks in San Jose, California. As part of AML-MLsys, you'll work on cutting-edge AI projects involving large language models (LLMs). The role requires proficiency in algorithms, data structures, and deep learning principles, alongside knowledge of Python and CUDA. Employees will enjoy comprehensive benefits including medical, dental, and vision insurance from day one. #J-18808-Ljbffr ByteDance
- ...seeking a Senior Research Scientist to lead machine learning system... .... You will help build and optimize large-scale distributed ML training and inference infrastructure, integrating GPU/NPU/RDMA and storage to run... ...contribute to cutting-edge LLM/AI system work while...Training
$212.8k
...maintain massively distributed ML training and inference systems around the world,... ..., and scalable systems for LLM/AIGC/AGI. Topic Content... ...Responsibilities Develop and optimize LLM training & inference & reinforcement... ...the next level. Optimize GPU and CUDA performance to...TrainingTemporary workLocal areaShift work- d-Matrix in Santa Clara, CA is seeking a Sr. Staff ML Researcher to advance LLM algorithmic optimization on our DNN accelerators. You will design and implement efficient inference algorithms, collaborating with mathematicians, ML researchers and engineers on high-impact...Suggested3 days per week
- ...The role: Sr. Staff, ML Researcher - LLM Algorithmic OptimizationWhat You... ...efficient algorithms that will be used to optimize large language model inference on DNN accelerators we develop. You... ...skills, experience, education, and training. We also offer incentive...Training3 days per week
- ByteDance’s Seed LLM team is advancing the next generation of large language models, tackling pre-training, post-training, inference, memory, learning, and interpretability. The role focuses... ..., instruction tuning, and optimization to improve reasoning and related abilities...Training
$244.8k
The Seed LLM team is dedicated to advancing the next generation of large language... ...development. It focuses on model pre‑training, post‑training, inference, memory capabilities, learning,... ...Responsibilities Explore large‑scale models and optimize associated systems. Perform data...TrainingTemporary workLocal area$212.8k - $387.6k
Responsibilities The Seed LLM team is dedicated to advancing the... ...pretraining, posttraining, inference, memory capabilities, learning... ...large-scale models and optimize systems. Data construction, instruction... ...in LLM-related areas such as training, evaluation, or applications....TrainingTemporary workInternship$187.04k - $359.72k
Senior Research Scientist, Foundation Model (LLM/ VLLM), TikTok - Trust and Safety Senior... ...II Do 1. Lead the design, training, and deployment of... ...data, and platform teams to optimize large-scale training pipelines... ...and leverage centralized GPU resources effectively. 5....TrainingFull timeTemporary workLocal area$244.8k
Research Engineer - LLM/VLM Inference Optimization (Seed Infra) Location San Jose Team Technology Employment... ...team oversees the distributed training, reinforcement learning framework, high... ..., or serving cost. Familiarity with GPU architecture and experience optimizing...TrainingTemporary workLocal area$192k - $304.75k
...Deep Learning Research Scientist, Efficiency!Join our... ...and algorithms to optimize neural networks for training and deployment.... ...effect on neural network inference and training... ...architectures, optimizers and LLM training.Experience... ...with GPU computing, kernels,...TrainingFull time$224k - $356.5k
...era in which our GPU acts as the brains... ...Applied Research Scientist with experience researching... ...foundation-model training, including... ...training of Nemotron LLM models which lead... ...the expansion and optimization of curation methodologies... ...of NVIDIA Inference Microservices (NIMs...TrainingFull timeWork at officeRemote workFlexible hours$171.6k - $222.2k
...assembling an elite team of world-class scientists and engineers to pioneer the next... ...development tools. Join the Amazon Kiro LLM-Training team and help create groundbreaking generative... ...data structures, parsing, numerical optimization, data mining, parallel and distributed...TrainingWork at officeLocal areaWorldwideFlexible hours- ByteDance is seeking a Senior Research Scientist for PICO in San Jose to advance multimodal large language... ...) for mixed reality. You will lead architecture optimization, data construction, and end-to-end training/inference acceleration tailored to MR scenarios. The role...Training
- ByteDance is seeking talented graduate students for a role focusing on LLM training and inference. Successful candidates will develop high-performance, scalable ML systems while collaborating closely with model researchers. Applicants need to be pursuing a Ph.D. in a related...Training
$212.8k - $387.6k
Senior Research Scientist - Machine Learning System... ...massively distributed ML training and Inference system/services... ...scalable systems for LLM/AIGC/AGI. In our team... ...system integrating with GPU/NPU/RDMA/Storage and... ...Responsible for developing and optimizing LLM inference...TrainingTemporary workLocal area$212.8k - $387.6k
...control plane for large‑scale LLM inference. We are building the next‑... ...networking, and storage—covering GPU virtualization, RDMA/DPDK‑... ...infrastructure that can self‑optimize via AI (AI for systems) Driving... ...such as large‑scale model training and inference, heterogeneous...TrainingTemporary workLocal area- ...large-scale MoE architecture training and routing optimization, cross-modal alignment and... ...of cutting-edge LLM research (e.g., long context... ...verification for training/finetuning/inference; Being familiar with PEFT,... ...Have a deep understanding of GPU and/or other AI...TrainingFlexible hoursShift work
$60 per hour
...and retrieval to ranking and user experience optimization. Project Overview, Challenges & Value With the rapid evolution of pre-trained large models, TikTok Search is advancing... ...search models (video, text, image). Explore LLM-based approaches for complex and multi-turn...TrainingHourly paySummer workInternshipLocal area$145k - $250k
Research Scientist Graduate- CV/NLP/Multimodal LLM,(Trust and Safety) - 2026 Start(PhD) Join to apply for the Research Scientist Graduate... ...in TikTok, and will also be responsible for optimizing our distributed model training framework continuously. Responsibilities...TrainingFull timeTemporary workFixed term contractSummer workInternshipLocal areaFlexible hours$60 per hour
...Encode hundreds of millions of items and massive user behavior data for large‑scale training, optimizing via advanced architectures, post‑training (e.g., RLVR), and efficient inference. Multimodal Large Models for E‑commerce: Build multilingual, multimodal models achieving...TrainingHourly payInternshipLocal area- ...Clara, CA. This hybrid position involves working on-site 3 days a week. The successful candidate will develop algorithms for optimizing LLM inference on our DNN accelerators. Ideal applicants should have a MSc or PhD in a relevant field and 5+ years of experience in...3 days per week
$156k - $316.8k
Research Scientist — Privacy-Preserving Large-Scale Model Training & Architecture Optimization Location: San Jose Employment Type: Regular Job Code: DW1L Responsibilities Design... ...Flow, hybrid AR + diffusion systems). Lead GPU-centric performance optimization, including...TrainingTemporary workLocal area- ...3 days per week. The role Senior Staff ML Researcher - LLM Algorithmic Optimization What You Will Do d-Matrix is seeking machine learning researchers... ...that will be used to optimize large language model inference on DNN accelerators we develop. You would be part of a...3 days per week
- ...headquarters 3 days per week. The role: Sr. Staff, ML Researcher - LLM Algorithmic Optimization What You Will Do d-Matrix is seeking machine learning... ...that will be used to optimize large language model inference on DNN accelerators we develop. You would be part of a close...3 days per week
$168k - $264.5k
...looking for a Senior Scientist to join our team and help... ...data generation for training frontier models. You will... ...pipelines using LLM-based methods and automated... ...post-training, and inference frameworks such as vLLM... ...or audio).Building and optimizing scalable data...TrainingFull timeRemote work$192k - $304.75k
...looking for a passionate scientist at the intersection of... ...will span synthetic training data generation, surrogate modeling, and co-optimized calibration-decoding... ...prediction and parameter inference without full... ...and modalities.Develop GPU-accelerated implementations...TrainingFull timeRemote work$184k - $287.5k
...Our core invention, the GPU, serves as the visual... ...hiring Senior Deep Learning Scientists to advance our efforts... ...come join our Nemotron LLM team. For more details... ...research to develop, train, fine-tune, and deploy... ...instruction tuning, preference optimization, and RLHF/RLVR/MOPD to...TrainingFull timeWork experience placement$171.6k - $222.2k
...development with systems and hardware optimization.QualificationsStrong... ...similar frameworks.Experience training, fine-tuning, or serving large-scale models.Knowledge of GPU architecture, CUDA, Triton,... ...execution.Optimize training and inference using CUDA, Triton, custom...TrainingLocal areaFlexible hours- ...edge of what’s possible with LLM inference on heterogeneous hardware.... ...deployment patterns to deep optimization of inference kernels, to building... ...level.• Familiarity with GPU kernel programming (CUDA/... ...understanding of modern LLM training and inference.Why D-Matrix Frontier...Training
$195.2k - $361.2k
...hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for... ...and edge environments — GPU/iGPUs, Vulkan backends — not... ...quality impact with the Post-Training teamCut CPU overhead and improve... ...-level codeExperience with LLM inference. (attention, KV cache...TrainingFull timeInternshipLocal areaImmediate startShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to LLM Training & Inference Scientist (GPU-Optimized). Be the first to apply!
- molecular biology scientist San Jose, CA
- water quality scientist San Jose, CA
- machine learning scientist San Jose, CA
- image scientist San Jose, CA
- machine learning research scientist San Jose, CA
- materials scientist San Jose, CA
- health scientist San Jose, CA
- scientist San Jose, CA
- graduate scientist San Jose, CA
- quality control scientist San Jose, CA
