Research Scientist - Vision Language Model
Institute of Foundation Models
Job Description
Job Description
About the Institute of Foundation Models
We are a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.
As part of our team, you’ll have the opportunity to work on the core of cutting-edge foundation model training, alongside world-class researchers, data scientists, and engineers, tackling the most fundamental and impactful challenges in AI development. You will participate in the development of groundbreaking AI solutions that have the potential to reshape entire industries. Strategic and innovative problem-solving skills will be instrumental in establishing MBZUAI as a global hub for high-performance computing in deep learning, driving impactful discoveries that inspire the next generation of AI pioneers.
Position SummaryAs a Research Scientist in the Vision Language Model (VLM) team, your role will be central to advancing state-of-the-art multimodal foundation models that integrate visual understanding, reasoning, and agentic capabilities. You will work on the research and development of large-scale VLM systems, spanning model architectures, data recipes for pre-training and post-training, and evaluation benchmarks. The role combines cutting-edge research with practical engineering, emphasizing large-scale data processing, filtering, and weighting pipelines, distributed training systems, and reinforcement learning algorithms and systems for multimodal reasoning and agent development.
Key ResponsibilitiesResearch and development of next-generation Vision Language Models across pre-training, instruction tuning, reasoning, and agents.
Develop novel architectures and training methodologies for integrating visual understanding, language reasoning, and tool-use capabilities.
Research efficient multimodal learning techniques, including data-efficient training, long-context modeling, model modularity, and inference optimization.
Build and improve large-scale multimodal datasets, synthetic data generation pipelines, and evaluation benchmarks for VLM capabilities.
Investigate multimodal reasoning, agentic behavior, OCR, grounding, document understanding, chart understanding, and visual question answering capabilities.
Contribute to technical reports, research publications, and open-source software.
Represent MBZUAI at research conferences and industry events, showcasing advancements in multimodal foundation models and large-scale AI systems.
Mentor junior researchers and collaborate across teams to drive impactful research initiatives.
PhD or equivalent research experience in Machine Learning, Computer Vision, Natural Language Processing, or Multimodal AI.
Salary Range
The posted salary range represents the company’s good faith estimate of the compensation for this position upon hire. The actual compensation offered may vary within this range depending on individual qualifications, including but not limited to relevant skills, experience, education, certifications, geographic location, and specific business needs.
Professional Experience Minimum
Experience working with large language models and/or vision-language models, including pre-training, fine-tuning, evaluation, or inference.
Strong Python and PyTorch development skills for large-scale machine learning research.
Experience with distributed training systems and large-scale model optimization.
Familiarity with multimodal datasets and data processing pipelines involving images, text, and video.
Understanding of modern deep learning architectures, including Transformers, attention mechanisms, and multimodal fusion techniques.
Experience with ML infrastructure, including model evaluation, debugging, optimization, and large-scale experimentation.
Problem-solving and research skills with the ability to independently drive research/engineering projects.
Effective communication and collaboration skills for working across research and engineering teams.
Hands-on experience training or fine-tuning large Vision Language Models or multimodal foundation models at scale.
Experience with distributed learning frameworks and infrastructure such as PyTorch Distributed, Megatron, Triton, or CUDA.
Research experience in multimodal reasoning, agentic systems, tool use, OCR, grounding, document understanding, or multimodal retrieval.
Experience with synthetic data generation, multimodal data curation, or automated evaluation frameworks for VLMs.
Familiarity with efficient training and inference techniques such as FlashAttention, quantization, tensor parallelism, pipeline parallelism, or memory optimization.
Experience contributing to open-source ML software and large-scale research codebases.
Strong publication record in leading AI conferences such as NeurIPS, ICML, ICLR, CVPR, ICCV, ECCV, ACL, EMNLP, or related venues.
Experience collaborating across research, infrastructure, and product-oriented teams to deliver state-of-the-art multimodal systems.
$165k - $185k
...Company Description The Bosch Research and Technology Center North America with offices... ...Silicon Valley focuses on Foundation Models, Big Data Visual Analytics, Explainable AI (XAI), Natural Language Processing, Computer Vision & Mixed Reality, Cloud Robotics, Data...LanguageWork experience placementWorldwide$185k - $215k
Company Description The Bosch Research and Technology Center... ...focuses on Foundation Models, Big Data Visual Analytics... ...Explainable AI (XAI), Natural Language Processing, Computer Vision & Mixed Reality, Cloud... ...As a Senior Research Scientist - Vision‑Language‑Action...LanguageWork experience placementLocal areaWorldwide$35 per hour
...services in translation, localization, and adaptation for over 250 languages with a growing network of over 400,000 in-country linguistic... ...: ▪️ Medical Insurance ▪️ Dental Insurance ▪️ Vision Insurance ▪️ FSA and HSA ▪️ Voluntary Life Insurance ▪️...LanguageRemote jobHourly payFull time- ...as our ability to measure it. At Sanas, model quality spans dimensions that automated... ...-world disfluency. We’re looking for a Research Scientist who can define what "better" actually... ...Noise Cancellation, Speech Enhancement, Language Translation, and more — ensuring each captures...Language
$184k - $287.5k
...new AI-powered application is built. We are seeking a senior vision language model engineer to design and build agentic data and training... ...companies in Physical AI.What you'll be doing:Partner with our researchers to develop and evaluate prototypes of our latest models,...LanguageFull time$192.2k - $260k
...s Delivery Foundation Model team, where you'll work... ...alongside world-class scientists and engineers to... ...direction for specific research initiatives, ensuring... ...combines ambitious research vision with real-world impact... ...Python, C++ or other languages- Strong publication record...LanguageLocal areaWorldwideFlexible hours$39 - $66 per hour
...Company Description The Bosch Research and Technology Center North America with offices... ...Silicon Valley focuses on Foundation Models, Big Data Visual Analytics, Explainable AI (XAI), Natural Language Processing, Computer Vision & Mixed Reality, Cloud Robotics, Data...LanguageWork experience placementInternshipLocal areaWorldwide$100k - $120k
...Job Summary The Research Scientist II is responsible for driving laboratory operations and collaborating... .... Knowledge of one or more scripting languages (e.g. Python). Fluorescence imaging and... ...paid time off, medical/dental/vision insurance and 401(k) to eligible employees...LanguageHourly pay- ...What You’ll Do Lead hands‑on research at the intersection of classical... ...image processing, computer vision, graphics, and content... ...signal processing, spectral/3D modeling, geometry, and calibration. Deep... ...segmentation, synthesis, captioning, language models). Experience with GPU...LanguageLocal areaWorldwideFlexible hours
$150k - $300k
...at scale. Silicon Valley Research Lab The Silicon Valley Research... ...and evaluate algorithms, models and prototypes of AI... ...learning, natural language processing, computer vision, reinforcement learning,... ...this full-time Research Scientist position is between $150,...LanguageFull timeH1bWork at office3 days per week- ...for multiple passionate Research Interns to join the... ...world-action foundation model with various world modalities including vision and physics associated... ...human data incorporation, language modality, and spatial reasoning... ...with the Research Scientists and Engineers on high...LanguageFor contractorsFor subcontractorCasual workInternshipWork at officeImmediate startRemote workDay shift
$301.75k - $355k
...The Senior Director for the Model LifeCycle team will undertake... ...Learning models, including Large Language Models (LLMs). What You’ll... ...field strongly preferred. Research publications at NeurIPS, ICML... ...Comprehensive health, dental & vision insurance Employer...LanguageTemporary work- ...software and foundation models enable vehicles to... ...driving systems. Our vision is to create autonomy... ...re looking for Applied Scientists to join Wayve Labs and... ...are a high‑conviction research team with the strategic... ...inputs, using vision, language, and active sensors. Key...LanguageFull timeWork at officeWork from homeVisa sponsorshipRelocation packageFlexible hours
$100k - $120k
...Technologies (IDT), a Danaher company, the Research Scientist II drives laboratory operations and... ...instruments. Knowledge of one or more scripting languages (e.g., Python). Fluorescence imaging... ...include paid time off, medical/dental/vision insurance, and a 401(k) plan. Danaher...Language- ...world-class scientific research, exploration, and... ...community of researchers, scientists, engineers, and innovators... ...time series. Develop Models & Simulations:Run... ...equivalent programming language. What Will Make You Stand... ...medical, dental, and vision coverage, retirement...LanguageLocal area
$190k - $250k
...developing large-scale generative world models that learn to predict realistic,... ...autonomous trucks. We are looking for a research scientist to lead the design and development of world... ...bonusesExcellent Medical, Dental, and Vision plans through Kaiser Permanente, Cigna,...Temporary workWork at officeVisa sponsorship$158k - $185k
...experience • Bilingual in one of the following languages preferred: Mandarin, Cantonese, Spanish,... ...full covered! • Medical, dental, and vision insurance • Paid time off and holidays •... ...Opportunity to work within an integrated care model • Meaningful work serving diverse...LanguageFull timePart time3 days per week- Language Specialist (Customer Support (Email)) Contract InterSources Inc was founded in 2007 providing intelligent data solutions to clients across industries and geographies. Over the years, we have built products on Business Intelligence & Big Data platform simplifying...LanguageContract workWork experience placement
$192.2k - $260k
...LLC Overview Are you a passionate scientist in the computer vision area who aspires to apply your skills... ...Multi‑modal LLMs and/or Vision Language Models and collaborating with different Amazon... ...for computer vision applications. Research and implement state‑of‑the‑art computer...LanguageFlexible hours$192.2k - $260k
...and experienced Applied Scientist to support adoption... ...services and tools for model customization, including... ...across large language models. As an Applied... ...and 6+ years of applied research experience Experience... ...insurance (medical, dental, vision, prescription, Basic Life...LanguageLocal areaFlexible hours$192.2k - $260k
...seeking a Sr. Applied Scientist to focus on Robot Navigation... ...In this role, you'll research and develop advanced... ...and foundation models—to build robust solutions... ...Python, C++,or related language - Experience with sim... ...(medical, dental, vision, prescription, Basic Life...LanguageLocal areaFlexible hours$171.6k - $222.2k
...team is looking for an Applied Scientist to work on the intersection of... ...to optimize application models across diverse domains, including Large Language and Vision, originating from leading frameworks... ...-critical tooling, publish research, and mentor a brilliant team of...LanguageLocal areaFlexible hours$171.6k - $222.2k
...Description Applied Scientists in AWS Automated Reasoning are dedicated to making AWS the... ...proving, symbolic simulation, programming language type systems, program analysis. Preferred... ...health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance...LanguageLocal areaFlexible hours$192.2k - $260k
...an experienced Applied Scientist who will join a team... ...you can pursue applied research, with many peta-bytes... ...and evaluation of AI models for predictive learning... ...C++, Python or related language - Experience with neural... ...(medical, dental, vision, prescription, Basic Life...LanguageLocal areaFlexible hours$171.6k - $222.2k
...We are seeking a Senior Applied Scientist to join our team in developing pioneering AI research, Generative AI, Agentic AI, Large Language Models (LLMs), Diffusion and Flow Models, and other... ...Modeling, Multi-modality Computer Vision, Diffusion Models, Reinforcement Learning...LanguageWorldwideFlexible hours$33 - $39 per hour
...occasional client site visits required. Requirements ~ Language Proficiency: Fluency in both English and Spanish is required... ...Benefits Choice of select medical plans Dental Plan Vision Plan Paid time off Bereavement Leave IRA Life Insurance...LanguageFull timeWork at office$192.2k - $260k
...and resourceful Applied Scientist to bring diverse... ...learning practitioner and a research leader. You will play... ...machine learning models from the ground up. At... ...C++, Python or related language - Experience with neural... ...insurance (medical, dental, vision, prescription, Basic...LanguageLocal areaWorldwideFlexible hours$171.6k - $222.2k
...building the next generation models for intelligent automation. Description... ...that will give Applied Scientists like you endless opportunities to see your research have a positive and immediate... ...multimodal models (especially vision-language models), reinforcement learning...LanguageLocal areaImmediate startFlexible hours$167.1k - $226.1k
...talented, and resourceful Applied Scientist in the field of Large Language Models (LLMs), Artificial Intelligence (AI... ...technical expertise to set the research agenda for how we measure conversational... ...insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance...LanguageLocal areaFlexible hours$126k - $248k
...fine‑tuned embedding models and rerankers to enable... ...by a strong team of AI researchers from Stanford, MIT,... ...seeking a Senior Research Scientist to join our team and... ...learning, and natural language processing. Familiarity... ...to utilize research vision to innovate the entire...LanguageLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Research Scientist - Vision Language Model. Be the first to apply!
- scientist Sunnyvale, CA
- lab scientist Sunnyvale, CA
- research scientist - biology Sunnyvale, CA
- r&d scientist Sunnyvale, CA
- applied scientist Sunnyvale, CA
- molecular biology scientist Sunnyvale, CA
- senior research scientist Sunnyvale, CA
- senior scientist Sunnyvale, CA
- health scientist Sunnyvale, CA
- applied sports scientist Sunnyvale, CA




