Senior Research Scientist (Multimodal Large Language Model) - PICO
$212.8kByteDance
Senior Research Scientist (Multimodal Large Language Model) - PICO Location: San Jose Team: Technology Employment Type: Regular Job Code: A57637 About the Team PICO-MR team is dedicated to pioneering core technologies for intelligent human-computer interaction in MR environments, with a focus on integrating multimodal large language models (MLLM) and tool-use capabilities to redefine user experiences. Our R&D directions cover cutting-edge fields including multimodal scene understanding, MLLM-based agent systems, tool-augmented MR interaction, 3D environment perception, and AIGC-driven content generation. Within MR scenarios, our work spans MLLM optimization and adaptation for MR, intelligent task execution with tool use, multimodal scene understanding (vision, point clouds, text), AIGC-based scene generation, depth estimation (Mono/Stereo/MVS), 3D environment perception, large-scale 3D scene reconstruction (3DGS, NeRF, etc.), visual localization, and lighting estimation - encompassing both fundamental research breakthroughs and industrial-grade solution deployment. Responsibilities Lead the R&D of multimodal large language models (MLLM) tailored for MR scenarios, integrating vision, point clouds, text, and other multimodal information—including model architecture optimization, cross-modal alignment, data construction, evaluation system enhancement, and end-to-end training/inference acceleration. Drive the research and implementation of MLLM tool-use capabilities in MR environments, enabling models to proficiently utilize spatial interaction and spatial computing-related professional tools, support tool calls for both single-turn and multi-turn conversations, and solve complex user tasks through interaction. Address key challenges in long-horizon, multi-turn tool-augmented tasks in MR, such as context memory management, tool selection strategy, and error correction mechanisms. Keep abreast of cutting-edge technologies in MLLM, multimodal intelligence, and tool-use research, and lead the application and deployment of innovative technologies in PICO's MR products. Collaborate with cross-functional teams (including software engineering, product design, and hardware development) to translate research outcomes into practical features that enhance user experience. Qualifications Minimum Qualifications Master's or Ph.D. degree in Computer Science, Electrical Engineering, Machine Learning, Artificial Intelligence, or a related quantitative field. Expertise in multimodal large model pre-training, post-training, fine-tuning, or cross-modal fusion technologies, with hands-on experience in model optimization, training workflow design, and performance tuning. Proven research experience in LLM tool use, reinforcement learning, LLM agents, or interactive learning, with a deep understanding of single-turn and multi-turn interaction mechanisms. Proficiency in core 2D/3D computer vision tasks, including detection, segmentation, depth estimation, image matching, and 3D scene perception. Skilled in Python and C++, with solid programming capabilities and experience in developing large-scale models using mainstream deep learning frameworks (PyTorch/TensorFlow). Excellent problem-solving and independent research abilities, capable of addressing complex technical challenges in the integration of MR and MLLM tool use. Preferred Qualifications Publications in AI/ML/CV conferences (e.g., NeurIPS, ICML, ICLR, CVPR, ICCV, ECCV, ACL, EMNLP) focusing on multimodal large models, LLM tool use, or agent systems. Hands-on experience in building large-scale MLLM training pipelines, tool-use evaluation systems, or multimodal agent platforms. Familiarity with MR/AR/VR technologies, spatial computing, or 3D scene reconstruction (3DGS, NeRF, etc.) is a strong plus. Experience in addressing long-horizon reasoning or asynchronous agent behavior challenges is highly valued. Award winners of competitions such as ACM-ICPC, NOI/IOI, TopCoder, or AI/ML contests (e.g., Kaggle) are preferred. Strong collaboration and communication skills, able to lead research initiatives and drive cross-team technical alignment. Job Information The base salary range for this position in the selected city is $212,800 - $450,000 annually. Compensation may vary outside of this range depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units. Benefits Employees have day-one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short-term and long-term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure). The Company reserves the right to modify or change these benefits programs at any time, with or without notice. For Los Angeles County (unincorporated) Candidates Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state, and local laws including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Our company believes that criminal history may have a direct, adverse and negative relationship on the following job duties, potentially resulting in the withdrawal of the conditional offer of employment: Interacting and occasionally having unsupervised contact with internal/external clients and/or colleagues. Appropriately handling and managing confidential information including proprietary and trade secret information and access to information technology systems. Exercising sound judgment. Reasonable Accommodation ByteDance is committed to providing reasonable accommodations in our recruitment processes for candidates with disabilities, pregnancy, sincerely held religious beliefs or other reasons protected by applicable laws. If you need assistance or a reasonable accommodation, please reach out to us at #J-18808-Ljbffr ByteDance
$187.04k
Research Scientist - Multimodal Interaction and World Model - Pre-Training Location: San Jose Responsibilities About Seed... ...intelligence. Our research spans large language models, GenMedia, AI for... ...including TikTok, Lemon8, CapCut and Pico as well as platforms specific...LanguageTemporary workLocal area$244.8k
...About the team The Seed Multimodal Interaction and World Model team is dedicated to developing... ...integrating vision, language, audio, and environment... ...systems. Familiarity with large-scale model training or... ...Preferred Qualifications Strong research track record in relevant...LanguageTemporary workLocal area$192k - $304.75k
We are now looking for a Senior Research Scientist focused on Multimodal Foundation Models and Robotics! NVIDIA is searching for... ...multimodal foundation models, large-scale robot learning, game AI,... ...following topics: LLMs; Large vision-language models; Video generative models...SeniorLanguageFull time- ByteDance is seeking a Senior Research Scientist for PICO in San Jose to advance multimodal large language models (MLLM) for mixed reality. You will lead architecture optimization, data construction, and end-to-end training/inference acceleration tailored to MR scenarios...SeniorLanguage
$192.2k - $260k
...Delivery Foundation Model team, where you... ...world-class scientists and engineers... ...an exceptional Senior Applied Scientist... ...for specific research initiatives, ensuring... ..., from multimodal training using... ..., C++ or other languages- Strong publication... ...with large-scale distributed...SeniorLanguageLocal areaWorldwideFlexible hours$212.8k
...Architecture team of the PICO-Interactive Perception... ...deployment of AI models, strive for excellence... ...seamless integration with large models. Topic Value: Breakthroughs in this research direction will enable intelligent... ...least one programming language (Python is preferred)....LanguageTemporary workInternshipLocal area$244.8k
...artificial general intelligence, with research spanning MLLM, GenMedia, AI for Science... ...industry-leading general foundation models and multimodal capabilities, powering over 50 application... ..., video understanding, or vision-language modeling Preferred Qualifications Expertise...Language$212.8k - $387.6k
...development. Our areas of focus include model pretraining, posttraining, inference, memory... ...technological innovation. Explore large-scale models and optimize systems. Data construction... ...such as reasoning, code, math. In-depth research and exploration of future use cases....LanguageTemporary workInternship$60 per hour
...building machine learning models and systems to protect... ...in multilingual and multimodal content, and upgraded... ...model: We study large-scale MoE architecture... ...generalization across 200+ languages/strategies, adversarial... ...direction and produce research outcomes with...LanguageHourly paySummer workInternshipLocal areaFlexible hoursShift work$184k - $287.5k
...advancement.NVIDIA is hiring Senior Deep Learning Scientists to advance our... ...and agentic multimodal AI. You will demonstrate... ...to help develop models capable of reasoning... ..., high-visibility large language models and multimodal... ...and applied research to develop, train,...SeniorLanguageFull timeWork experience placement- ...Beam team in Mountain View is seeking a Senior Research Scientist to advance foundation models and multimodal AI research. You will design large-scale experiments, prototype new... ...Strong background in deep learning, vision-language models, and responsible AI principles is...SeniorLanguage
$254.4k
...AI sector, the Seed team's research spans MLLM, GenMedia, AI... ...leading general foundation models and cutting-edge multimodal capabilities. Our technology... ...and audio, music, natural language understanding, and... ...synthesis, audio generation, large language model, computer vision...LanguageTemporary workLocal areaWorldwide$187.04k - $359.72k
Senior Research Scientist, Foundation Model (LLM/ VLLM), TikTok - Trust and Safety Senior Research Scientist, Foundation... ...challenges in multilingual, multimodal, and low-resource scenarios by... ...data, and platform teams to optimize large-scale training pipelines, improve...SeniorFull timeTemporary workLocal area$174.72k - $295.68k
...strong expertise in generative modeling and large-scale deep learning systems... .... In this role, you will research, implement, and evaluate... ...physical world from large-scale multimodal data — predicting how a... ...training to improve Vision-Language-Action (VLA) driving performance...SeniorLanguageFull time$184k - $287.5k
...is built. We are seeking a senior vision language model engineer to design and... ...be doing:Partner with our researchers to develop and evaluate prototypes... ...and maintain high‑quality multimodal datasets (e.g., video,... ...through contributions to large internal or open-source projects...SeniorLanguageFull time$224k - $356.5k
...is searching for a senior or principal engineer... ...infrastructure for large-scale foundation model training in the Generalist... ...Embodied Agent Research (GEAR) group. Our team... ...works on multimodal foundation models, large... ...a high-performance language such as C++ for efficient...SeniorLanguageFull time$156k - $316.8k
Research Scientist — Privacy-Preserving Large-Scale Model Training & Architecture Optimization Location: San Jose Employment Type: Regular Job Code: DW1L Responsibilities... ...models, enabling shared context, long-sequence multimodal reasoning, and scalable training without...Temporary workLocal area- ...that enhance content discovery. The role requires a strong foundation in AI/ML, with experience in training and deploying vision-language models. You will work closely with cross-functional teams and need to demonstrate excellent communication skills. The compensation...SeniorLanguageFlexible hours
$184k - $287.5k
...accelerated computing and AI. We are seeking passionate researchers to advance longitudinal multimodal foundation models for healthcare. Recent progress in medical AI has... ..., and multimodal generative modeling.Build large-scale datasets, benchmarks, and open-source...SeniorFull time$165k - $185k
Company DescriptionThe Bosch Research and Technology Center North America with offices in Sunnyvale,... ...research in Silicon Valley focuses on Foundation Models, Big Data Visual Analytics, Explainable AI (XAI), Natural Language Processing, Computer Vision & Mixed Reality,...SeniorLanguageWork experience placementWorldwide$212.8k - $387.6k
...for a talented candidate to develop and scale vision foundation models, focusing on image and video modalities. Applicants should be... ...'s degree in a relevant field, have excellent coding skills in languages like C/C++ or Python, and possess strong problem-solving abilities...Language- ...Number: P25F11 Honda Research Institute USA (HRI-... ...seeking a Research Scientist to push the frontiers... ...will develop multimodal and foundation-model-based approaches that... ...reason over vision, language, action, and other embodied... ...evaluate models on large-scale, diverse...LanguageWork experience placementShift work
- ...Institute of Foundation Models We are a dedicated research lab for building,... ...researchers, data scientists, and engineers,... ...Scientist in the Vision Language Model (VLM) team,... ...state-of-the-art multimodal foundation models that... ...and development of large-scale VLM systems,...Language
- Advanced Micro Devices (AMD) is seeking an Applied Research Scientist to advance large language and multimodal models, including image/video generation. You will train, fine‑tune, and align LLMs, LMMs, and diffusion models, while pushing the state of the art and shaping...Language
- ...We are looking for a Research Scientist to join the Multi-Embodiment... ...building foundation models for general-purpose... ...embodied intelligence and large-scale machine learning... ...work may span vision-language-action models, world and action models, multimodal and omni models, video...LanguageFull timeWork from home
$272k - $431.25k
...generating it! Our world model team is pushing the boundaries of multimodal AI, robotics, and... ...We are looking for a Senior Research Manager to lead world-... ...Lead a team of Research Scientists focused on world-model... ...video models, vision-language-action models, diffusion...SeniorLanguageFull time$174.72k - $295.68k
...full-time Machine Learning Engineer / Research Scientist to drive the modeling and algorithmic development of XPENG’s next-generation Vision-Language-Action (VLA) Foundation Model — the core... ...experts to design, train, and deploy large-scale multi-modal models that unify...SeniorLanguageFull time- ...collaboration. We are looking for a Senior ML Research Engineer, Embodied... ...Design and deploy vision-language(-action) models (VLM/VLA) for contextual... ...PyTorch, JAX). Expertise in large foundation models (VLM,... ...Decent understanding of multimodal models, modern ML architectures...SeniorLanguageWork at officeVisa sponsorship
$192k - $304.75k
NVIDIA is searching for an outstanding Senior Researcher working on efficient deep learning to... ...excited about methods for post-training model optimization (pruning, quantization,... ...backbones is required.Experience with large language models and large vision-language models...SeniorLanguageFull time$192k - $304.75k
We are now looking for a Senior Research Scientist for Generative AI!NVIDIA is searching... ...impacts with generative AI models. You will be building... ...and scaling them with large datasets and compute. After... ..., computer vision, natural language processing, or computer graphicsTrack...SeniorLanguageFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Research Scientist (Multimodal Large Language Model) - PICO. Be the first to apply!
- molecular biology scientist San Jose, CA
- water quality scientist San Jose, CA
- machine learning scientist San Jose, CA
- image scientist San Jose, CA
- machine learning research scientist San Jose, CA
- materials scientist San Jose, CA
- health scientist San Jose, CA
- scientist San Jose, CA
- graduate scientist San Jose, CA
- quality control scientist San Jose, CA

