Multimodal Vision Foundation Models Research Scientist
ByteDance
ByteDance in San Jose seeks a PhD candidate for the Seed Vision team focused on visual generation models. Responsibilities include developing and scaling foundation models, optimizing architectures, and exploring real-world applications. Ideal candidates have excellent coding skills in C/C++ or Python, experience in computer vision or multimodal learning, and a strong understanding of deep learning methodologies. The position offers robust benefits including healthcare, a 401(k) plan, and generous paid time off. #J-18808-Ljbffr ByteDance
$192k - $304.75k
We are now looking for a Senior Research Scientist focused on Multimodal Foundation Models and Robotics! NVIDIA is searching for an outstanding research scientist... ...at least one of the following topics: LLMs; Large vision-language models; Video generative models and diffusion...FoundationFull time$212.8k - $387.6k
ByteDance is looking for a talented candidate to develop and scale vision foundation models, focusing on image and video modalities. Applicants should be pursuing a Bachelor's or Master's degree in a relevant field, have excellent coding skills in languages like C/C++...Foundation$244.8k
...artificial general intelligence, with research spanning MLLM, GenMedia, AI for... ...has launched industry-leading general foundation models and multimodal capabilities, powering over 50 application... ...learning, video understanding, or vision-language modeling Preferred...Foundation$212.8k - $387.6k
Team Overview The Seed Multimodal Interaction and World Model team is dedicated to developing models that have human-level multimodal understanding... ...products. Responsibilities Develop multimodal foundation models integrating vision, language, audio, and environment signals....FoundationTemporary workInternshipLocal area$212.8k - $387.6k
Responsibilities Develop and scale vision foundation models across image and video modalities. Design data pipelines, pre‑training strategies... ...core capabilities such as perception, reasoning, and multimodal understanding. Optimize model architectures, training efficiency...FoundationTemporary workInternshipLocal area$184k - $287.5k
NVIDIA is redefining healthcare through accelerated computing and AI. We are seeking passionate researchers to advance longitudinal multimodal foundation models for healthcare. Recent progress in medical AI has improved the interpretation of individual images and clinical...FoundationFull time$165k - $185k
...DescriptionThe Bosch Research and Technology Center... ...Silicon Valley focuses on Foundation Models, Big Data Visual... ...Processing, Computer Vision & Mixed Reality, Cloud... ...a Research Scientist- Vision-Language-Action... ...building and applying multimodal transformer-based sequence...FoundationWork experience placementLocal areaWorldwide- Google Beam team in Mountain View is seeking a Senior Research Scientist to advance foundation models and multimodal AI research. You will design large-scale... ...technologies at Google. Strong background in deep learning, vision-language models, and responsible AI principles is...Foundation
- ...Labor for dull, dirty, and dangerous work. The team develops vision-language models and world models to enable safe, real-world robot... ...modal perception, planning, and action policies, advancing foundation models, and delivering production-grade solutions on RoboForce...FoundationWork at office
- ...deployment of innovative AI solutions that enhance content discovery. The role requires a strong foundation in AI/ML, with experience in training and deploying vision-language models. You will work closely with cross-functional teams and need to demonstrate excellent...FoundationFlexible hours
- ...About the Institute of Foundation Models We are a dedicated research lab for building, understanding... ...class researchers, data scientists, and engineers, tackling... ...Scientist in the Vision Language Model (VLM) team... ...advancing state-of-the-art multimodal foundation models that...Foundation
- ...TeamThe Content Representation Models team creates a single,... ...library by developing foundation models that understand everything... ...RoleWe are looking for a Research Scientist specializing in embeddings... ....Experience in computer vision or multimodal AIIndustry experience in...FoundationHourly payFull timeImmediate startFlexible hours
- Apple in Sunnyvale is seeking a Multimodal AI Researcher to push the boundaries of foundation models for real-time multimodal data, including video, audio, and text. You will work on interactive models, audio-to-audio modeling, and streaming multimodal systems, driving...Foundation
$165k - $195k
Company DescriptionThe Bosch Research and Technology Center North America with offices in Sunnyvale, California... ..., our AI research in Silicon Valley focuses on Foundation Models, Natural Language Processing, Computer Vision & Mixed Reality, Cloud Robotics, Big Data Visual...FoundationFull timeWork experience placementLocal areaWorldwide- Institute of Foundation Models, operating the AllWorld Team at MBZUAI, seeks researchers to develop the PAN world models that simulate physical environments. You will... ...engineering and research to push state-of-the-art multimodal AI. Candidate requirements include MSc/PhD in...Foundation
- Responsibilities Conduct research on multimodal foundation models and related systems. Explore methods to improve model capabilities across modalities, including areas such as vision-language modeling, world modeling, and representation learning. Design and prototype algorithms...Foundation
- Research Scientist: Robotics Foundation Models - Honda Research Institute USA Honda Research Institute USA (HRI‑US)... ...developing robotic AI systems that combine multimodal foundation models (e.g., VLMs,... .... Train, fine‑tune, and design vision‑language and multimodal foundation...FoundationWork experience placement
- Job Number: P25F11 Honda Research Institute USA (HRI-US) is seeking a Research Scientist to push the frontiers of machine... ...candidate will develop multimodal and foundation-model-based approaches that unify... ...multimodal systems that reason over vision, language, action, and...FoundationWork experience placementShift work
$60 per hour
Responsibilities About the team: The Seed Vision team focuses on foundational models for visual generation, developing multimodal generative models, and carrying out leading research and application development to solve fundamental computer vision challenges in GenAI....FoundationHourly payInternshipLocal area- ...in Milpitas, CA is seeking researchers to advance AI-powered robot control through vision-language models and world-models. You will design... ...planning, and integrate multimodal data for natural human-robot... ...and a strong track record in foundation models, PyTorch/JAX, and...FoundationWork at office
- ...Overview The Content Representation Models team develops unified foundation models that understand video,... ...the Role We are seeking a Research Scientist specializing in embeddings and... ...in LLMs. Experience in computer vision or multimodal AI. Industry experience in recommendation...FoundationFull timeFlexible hours
$244.8k
...post‑training (SFT/RL) of Seed multimodal models and LLM. Participate in unified... ...horizon task capabilities of agentic foundation models, and conduct in‑depth research on agentic RL. Develop large‑... ...access to medical, dental, and vision insurance; a 401(k) savings plan...FoundationTemporary workLocal area$244.8k
...Global GenAI team focuses on applied research in Generative AI, delivering intelligent... ...content with ease. The team works on multimodal foundation models, image and video generation and... ...Benefits include medical, dental, and vision insurance; 401(k) with company match;...FoundationTemporary workLocal area$156k - $316.8k
Research Scientist — Privacy-Preserving Large-Scale Model Training & Architecture Optimization Location... ...-running, multi-stage foundation model training. Build... ..., long-sequence multimodal reasoning, and scalable... ...medical, dental, and vision insurance, a 401(k) savings...FoundationTemporary workLocal area$162k - $316.8k
World Model Research Scientist (Intelligent Creation) - Global Frontier Tech Recruitment... ...in generative AI and multimodal models (e.g., image, video). Develop large-scale foundation models (LLMs/VLMs), through... ...to medical, dental, and vision insurance, a 401(k) savings...FoundationTemporary workLocal area$254.4k
...both technology and society. With a long-term vision for the AI sector, the Seed team's research spans MLLM, GenMedia, AI for Science, and... ...we have launched industry-leading general foundation models and cutting-edge multimodal capabilities. Our technology powers over 50...FoundationTemporary workLocal areaWorldwide$187.04k - $359.72k
Senior Research Scientist, Foundation Model (LLM/ VLLM), TikTok - Trust and Safety Senior Research Scientist... ...Tackle challenges in multilingual, multimodal, and low-resource scenarios by... ...one access to medical, dental, and vision insurance, a 401(k) savings plan with...FoundationFull timeTemporary workLocal area$60 per hour
...for building machine learning models and systems to protect our... ...complexity in multilingual and multimodal content, and upgraded generative... ...: (1) Multimodal moderation foundation model: We study large-scale... ...for this direction and produce research outcomes with significant...FoundationHourly paySummer workInternshipLocal areaFlexible hoursShift work$187.04k
A leading tech company is seeking a Research Scientist for Multimodal Interaction and World Model. This role focuses on enhancing models for multimodal data understanding... ..., plus benefits that include medical, dental, and vision insurance, a 401(k) plan, and paid personal time....- ByteDance in San Jose seeks a PhD candidate to join the Seed Multimodal Interaction team, focused on developing innovative multimodal models that integrate vision, language, and more. The successful candidate will work on improving decision-making and reasoning capabilities...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Multimodal Vision Foundation Models Research Scientist. Be the first to apply!
- molecular biology scientist San Jose, CA
- water quality scientist San Jose, CA
- machine learning scientist San Jose, CA
- image scientist San Jose, CA
- machine learning research scientist San Jose, CA
- materials scientist San Jose, CA
- health scientist San Jose, CA
- scientist San Jose, CA
- graduate scientist San Jose, CA
- quality control scientist San Jose, CA

