Research Scientist- Vision-Language-Action (VLA) Models
$165k - $185kRobert Bosch
Company DescriptionThe Bosch Research and Technology Center North America with offices in Sunnyvale, California, Pittsburgh, Pennsylvania, and Cambridge, Massachusetts is a part of the global Bosch Group ( a company with over 70 billion euro revenue, 400,000 employees worldwide, a very diverse product portfolio, and a history spanning over 125 years. The Research and Technology Center North America (RTC-NA) is dedicated to providing technologies and system solutions for various Bosch business fields, primarily in the field of artificial intelligence, energy technologies, internet technologies, circuit design, semiconductors and wireless, as well as advanced MEMS design.As a part of the global research, our AI research in Silicon Valley focuses on Foundation Models, Big Data Visual Analytics, Explainable AI (XAI), Natural Language Processing, Computer Vision & Mixed Reality, Cloud Robotics, Data Science, AI System Engineering, Time-series Analysis. We develop scalable, intelligent, and trustworthy AIoT solutions for Bosch products and services in application areas such as automated driving, advanced driver assistance systems (ADAS), robotics, smart manufacturing, enterprise AI, health care, smart home and building solutions.Originating from the AI research in Silicon Valley, our Intelligent Autonomous Systems group is responsible for enabling future autonomous Bosch products by pushing the boundaries of automated driving, advanced driver assistance systems (ADAS), robotics and automation through key innovations that encompass system architecture and AI components. These include methods for motion planning, high level task planning and decision making as well as systems for making these technologies work on real products by building frameworks that take advantage of technologies in the field of reliable distributed computing. We work with internal partners of different Bosch business units to transfer our solutions into future products. We also actively collaborate with leading groups in academia and industry to promote research ideas and publish research findings in internationally renowned conferences and journals such as CVPR, ICRA, IROS, RSS, NeurIPS and CoRL.Job DescriptionAs a Research Scientist- Vision-Language-Action (VLA) Models, you contribute to research projects at the forefront of the ADAS/AD industry. Key responsibilities include:Conduct research and engineering in core AI and machine learning fields to enable Embodied AI (including computer vision, autonomous planning, open-world learning, and so on) for related business domains of ADAS/AD, industrial automation, robotics etc.Push the boundaries in (modular) end-to-end perception and planning for ADAS/AD, incorporating advancements in large vision-language-(action) models to aid reasoning capabilities and explainability.Collaborate cross-functionally with global research and engineering teams to ensure seamless technology transfer and system integration.Implement research results to solve real-world challenges, ensuring high-quality system integration within Bosch's existing platforms.Stay at the forefront of innovation by actively engaging with academic and industry communities through conferences, workshops, and technical events.Document and disseminate research findings through high-caliber publications and/or patent submissions.QualificationsBasic QualificationsPh.D. in Computer Science, Robotics or a related discipline or Master's degree with >= 2 years industry experience after graduation.A minimum of 3 years of R&D experience, or an equivalent graduate research background, primarily in AI technologies including Computer Vision and Robotic or Automotive Motion and Behavioral Planning.Proficiency in one or more programming languages commonly used in machine learning (e.g., Python, C++, Rust).Strong interpersonal, communication, and teamwork capabilities.Knowledge of major machine learning frameworks like TensorFlow or PyTorch.Hands-on experience in reinforcement learning for behavior or motion planning or other applicable contexts and familiarity with common RL techniques (e.g. PPO, DQN, DDPG).A strong portfolio of publications in premier machine learning, deep learning, robotics and computer vision journals and conferences.Preferred QualificationsExperience with real-world product development and deployment of autonomous systems.Hands-on experience building and applying multimodal transformer-based sequence-to-sequence models, especially multimodal vision-language-action models.Hands-on experience in computer vision and deep learning, with work in any of the following areas: multimodal transformers, multimodal language models, diffusion models, NeRF, gaussian splatting, object detection / segmentation, 3D scene understanding, sensor calibration, SfM, voxel/BEV grid-based feature representation.Additional InformationWe offer a competitive base salary for this position with a range in US-California of --$165,000 - $185,000 along with an annual corporate bonus, and a long-term incentive bonus designed to reward sustained impact and contribution over time. Within the salary range, the individual pay is determined based on several factors, including, but not limited to, work experience and job knowledge, complexity of the role, job location, etc.Your well-being matters at Bosch! We offer a a benefits package designed to empower you in every area of your life. This includes premium health coverage, a 401(k) with generous matching, resources for financial planning and goal setting, ample paid time off, parental leave, and comprehensive life and disability protection. Your Recruiter can share more details for this position during the interview process.Learn more about our full benefits offerings by visiting: .Equal Opportunity Employer, including disability / veterans.*Bosch adheres to Federal, State, and Local laws regarding drug-testing. Employment is contingent upon the successful completion of a drug screen and background check. Candidates who have been offered the position must pass both screenings before their start date.#LI-JM1SummaryType: Full-timeFunction: Research
- ...CAResearch and Development - Computer Vision and Deep Learning /Full-time /... ...decision-making, building the Vision-Language-Action (VLA) models that form SuperDrive's reasoning layer... ...platform teams to bring models from research to productionRequired qualificationsM...LanguageFull time
- ...About the Institute of Foundation Models We are a dedicated research lab for building, understanding, using... ...world-class researchers, data scientists, and engineers, tackling the most... ...Summary As a Research Scientist in the Vision Language Model (VLM) team, your role will...Language
$193.93k - $352.29k
...flexible, partner-led business model, Nuro is working toward a... ...collaborate closely with researchers and engineers on the Learned... ...models. Leverage large language models and world foundation... ...autonomous driving. Experiences in vision-language-action models, reinforcement...LanguageImmediate startFlexible hours- ...collaboration. We are looking for a Senior / Staff AI Research Scientist, Foundation Models to advance robotic embodied intelligence. In... ...tasks. Responsibilities Design and deploy vision-language(-action) models (VLM/VLA) for contextual understanding and generalized...LanguageWork at officeVisa sponsorship
$192k - $304.75k
We are now looking for a Senior Research Scientist focused on Multimodal Foundation Models and Robotics! NVIDIA is... ...following topics: LLMs; Large vision-language models; Video generative models and diffusion algorithms; or Action-based transformers.Outstanding...LanguageFull time$165k - $195k
...The Bosch Research and Technology Center North America with offices in Sunnyvale, California, Pittsburgh, Pennsylvania... ...AI research in Silicon Valley focuses on Foundation Models, Natural Language Processing, Computer Vision & Mixed Reality, Cloud Robotics, Big Data Visual...LanguageFull timeWork experience placementLocal areaWorldwide$218.8k - $335.3k
...ready to redefine mobility and shape the future of autonomous transportation? As a Staff Research Scientist specializing in Vision-Language Models (VLMs), Vision-Language-Action models (VLAs), and Onboard Foundational Models, you will advance the frontier of artificial...LanguageFull timeLocal areaRemote workWork from homeRelocation packageFlexible hours$230k - $380k
...Role We're looking for Research Scientists to join Wayve Labs and help... ...areas: World & Reward Modeling: Building realistic, diverse... ...consequences and costs of actions. Representation Learning... ...and multimodal inputs, using vision, language, and active sensors. Key...LanguageFull timeWork at officeVisa sponsorshipRelocation packageFlexible hours$200k - $287.5k
Senior Machine Learning Researcher At Toyota Research... ...Policy and Large Behavior Models (LBM). We are looking... ...of-the-art, pixels-to-action, end-to-end system for... ...integrating visual-language-action modalities. Beyond... ...a focus on computer vision as the primary sensing...LanguageLocal areaShift work- ...multiple passionate Research Interns to join the Research... ...on pretraining world-action foundation model with various world modalities including vision and physics... ...data incorporation, language modality, and spatial... ...closely with the Research Scientists and Engineers on high...LanguageFor contractorsFor subcontractorCasual workInternshipWork at officeImmediate startRemote workDay shift
- ...TeamThe Content Representation Models team creates a single, unified "language" for Netflix's entire library by... ...personalizationAbout the RoleWe are looking for a Research Scientist specializing in embeddings and... ...in LLMs.Experience in computer vision or multimodal AIIndustry...LanguageHourly payFull timeImmediate startFlexible hours
$35 per hour
...services in translation, localization, and adaptation for over 250 languages with a growing network of over 400,000 in-country linguistic... ...: ▪️ Medical Insurance ▪️ Dental Insurance ▪️ Vision Insurance ▪️ FSA and HSA ▪️ Voluntary Life Insurance ▪️...LanguageRemote jobHourly payFull time$171.6k - $222.2k
...of deep learning and large language models.We leverage advanced robotics... ...candidate will contribute to research that bridges the gap between... ...collaboration.As an Applied Scientist, you will develop and... ...strength in at least one: computer vision, multimodal models,...LanguageLocal areaWorldwideFlexible hours$230k - $270k
...observation alone, is an open research frontier. Our lab... ...robot foundation models on our GPUs, and... ...research : adapting vision-language models to achieve low-... ...understanding : temporal action detection/segmentation... ...Exposure to robot learning (VLA models, imitation...LanguageContract workTemporary workWorldwideFlexible hours$231.5k - $405.1k
...team Our Core AI Research team develops... ...tools, take reliable action across stateful workflows... ...work across LLM model post-training,... ...a Staff Research Scientist, you will independently... ...more modalities - language, documents, images... ...AI, computer vision, speech/audio, or...LanguageFull timeWork at officeImmediate startRemote workFlexible hoursShift work$165k - $185k
Company DescriptionThe Bosch Research and Technology Center North America with offices... ...in Silicon Valley focuses on Foundation Models, Big Data Visual Analytics, Explainable AI (XAI), Natural Language Processing, Computer Vision & Mixed Reality, Cloud Robotics, Data Science...LanguageWork experience placementWorldwide- ...Job Title: Research Scientist We are seeking a dedicated Research Scientist to join our... ...with deep experience on relevant work (vision-action models, robotics + RL, etc.). Key... ...field Proficiency in programming languages relevant to AI research, particularly...LanguageFull time
$215.28k - $364.32k
...connectivity.We are looking for a full-time Machine Learning Engineer / Research Scientist to drive the modeling and algorithmic development of XPENG’s next-generation Vision-Language-Action (VLA) Foundation Model — the core brain that powers our end-to-end autonomous...LanguageFull time$192.2k - $260k
...exceptional Sr. Applied Scientist to lead the... ...where conventional vision alone falls short.... ...insufficient. You will lead research that translates... ...deep learning models for object... ...rigor, and a bias for action. We believe in learning... ...Python or related language ~ PhD in...LanguageLocal areaFlexible hoursNight shift- ...What You’ll Do Lead hands‑on research at the intersection of classical... ...image processing, computer vision, graphics, and content... ...signal processing, spectral/3D modeling, geometry, and calibration. Deep... ...segmentation, synthesis, captioning, language models). Experience with GPU...LanguageLocal areaWorldwideFlexible hours
$207k - $300k
...data pipelines to detect model misbehavior and misuse end-to-end.Research and develop cross-... ...using model activations, actions, chains-of-thought and final... ...infrastructure teams and data scientists to scale your work and... ...generative AI and Large Language Models (LLM).Preferred...Language$192k - $304.75k
We are now looking for a Senior Research Scientist for Generative AI!NVIDIA is searching for... ...great impacts with generative AI models. You will be building research prototypes... ...practice of deep learning, computer vision, natural language processing, or computer...LanguageFull time$174k - $252k
...understanding of foundation model, language and multi-modal technologies... ...Gemma).Experience conducting research and development, including... ...different data types (e.g., vision, audio, text).Proven expertise... ...of work. As a Research Scientist, you'll setup large-scale tests...Language$204k - $259k
...Research Scientist, RL for Autonomous Planning & World Modeling Waymo is an autonomous driving technology company with the mission to be the world's most trusted... ...for this role include: Health, dental, vision, life, disability insurance Retirement Benefits...Full timeTemporary workRemote work$165k - $185k
...Description The Bosch Research and Technology... ...on Foundation Models, Big Data Visual Analytics... ...AI (XAI), Natural Language Processing, Computer Vision & Mixed Reality, Cloud... ...As a Research Scientist- Robotics AI, you contribute... ...vision-language-(action) models to aid reasoning...LanguageWork experience placementLocal areaWorldwide$192.2k - $260k
...practical experience to join the Modeling and Optimization (MOP)... ...learning, robotics, operations research, statistics, mathematics or equivalent... ...- Knowledge of programming languages such as C/C++, Python, Java... ...insurance (medical, dental, vision, prescription, Basic Life &...LanguageLocal areaFlexible hours$171.6k - $222.2k
...Engineer to build efficient, stable foundation models for long-horizon agentic workloads. The... ...in Java, C++, Python or related language- Experience in any of the following areas... ...including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance...LanguageLocal areaFlexible hours$165k - $185k
...The Bosch Research and Technology Center North America with offices in Sunnyvale,... ...in Silicon Valley focuses on Foundation Models, Big Data Visual Analytics, Explainable AI (XAI), Natural Language Processing, Computer Vision & Mixed Reality, Cloud Robotics, Data Science...LanguageFull timeWork experience placementWorldwide$38 per hour
...outputs and provide structured, actionable feedback to improve accuracy... ...edge cases in both human and model outputsParticipate in... ...We're Looking ForNative-level language proficiency and a university... ...company policyMedical, Dental, and Vision Insurance (eligibility applies...LanguageFull timeContract workRemote workVisa sponsorship- ...more than selecting the newest model. It requires disciplined... ...training when justified, speech and language quality, evaluation... ...product-facing modeling role. Research depth matters, but success is... ...Option Plan. Medical, dental, vision, retirement, leave, and disability...LanguageTemporary workImmediate start
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Research Scientist- Vision-Language-Action (VLA) Models. Be the first to apply!
- applied scientist Sunnyvale, CA
- operations research scientist Sunnyvale, CA
- applied sports scientist Sunnyvale, CA
- health scientist Sunnyvale, CA
- drug safety scientist Sunnyvale, CA
- scientist biology Sunnyvale, CA
- safety scientist Sunnyvale, CA
- machine learning research scientist Sunnyvale, CA
- application scientist Sunnyvale, CA
- materials scientist Sunnyvale, CA




