Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Research Scientist- Vision-Language-Action (VLA) Models

$165k - $185k

Robert Bosch

Company DescriptionThe Bosch Research and Technology Center North America with offices in Sunnyvale, California, Pittsburgh, Pennsylvania, and Cambridge, Massachusetts is a part of the global Bosch Group ( a company with over 70 billion euro revenue, 400,000 employees worldwide, a very diverse product portfolio, and a history spanning over 125 years. The Research and Technology Center North America (RTC-NA) is dedicated to providing technologies and system solutions for various Bosch business fields, primarily in the field of artificial intelligence, energy technologies, internet technologies, circuit design, semiconductors and wireless, as well as advanced MEMS design.As a part of the global research, our AI research in Silicon Valley focuses on Foundation Models, Big Data Visual Analytics, Explainable AI (XAI), Natural Language Processing, Computer Vision & Mixed Reality, Cloud Robotics, Data Science, AI System Engineering, Time-series Analysis. We develop scalable, intelligent, and trustworthy AIoT solutions for Bosch products and services in application areas such as automated driving, advanced driver assistance systems (ADAS), robotics, smart manufacturing, enterprise AI, health care, smart home and building solutions.Originating from the AI research in Silicon Valley, our Intelligent Autonomous Systems group is responsible for enabling future autonomous Bosch products by pushing the boundaries of automated driving, advanced driver assistance systems (ADAS), robotics and automation through key innovations that encompass system architecture and AI components. These include methods for motion planning, high level task planning and decision making as well as systems for making these technologies work on real products by building frameworks that take advantage of technologies in the field of reliable distributed computing. We work with internal partners of different Bosch business units to transfer our solutions into future products. We also actively collaborate with leading groups in academia and industry to promote research ideas and publish research findings in internationally renowned conferences and journals such as CVPR, ICRA, IROS, RSS, NeurIPS and CoRL.Job DescriptionAs a Research Scientist- Vision-Language-Action (VLA) Models, you contribute to research projects at the forefront of the ADAS/AD industry. Key responsibilities include:Conduct research and engineering in core AI and machine learning fields to enable Embodied AI (including computer vision, autonomous planning, open-world learning, and so on) for related business domains of ADAS/AD, industrial automation, robotics etc.Push the boundaries in (modular) end-to-end perception and planning for ADAS/AD, incorporating advancements in large vision-language-(action) models to aid reasoning capabilities and explainability.Collaborate cross-functionally with global research and engineering teams to ensure seamless technology transfer and system integration.Implement research results to solve real-world challenges, ensuring high-quality system integration within Bosch's existing platforms.Stay at the forefront of innovation by actively engaging with academic and industry communities through conferences, workshops, and technical events.Document and disseminate research findings through high-caliber publications and/or patent submissions.QualificationsBasic QualificationsPh.D. in Computer Science, Robotics or a related discipline or Master's degree with >= 2 years industry experience after graduation.A minimum of 3 years of R&D experience, or an equivalent graduate research background, primarily in AI technologies including Computer Vision and Robotic or Automotive Motion and Behavioral Planning.Proficiency in one or more programming languages commonly used in machine learning (e.g., Python, C++, Rust).Strong interpersonal, communication, and teamwork capabilities.Knowledge of major machine learning frameworks like TensorFlow or PyTorch.Hands-on experience in reinforcement learning for behavior or motion planning or other applicable contexts and familiarity with common RL techniques (e.g. PPO, DQN, DDPG).A strong portfolio of publications in premier machine learning, deep learning, robotics and computer vision journals and conferences.Preferred QualificationsExperience with real-world product development and deployment of autonomous systems.Hands-on experience building and applying multimodal transformer-based sequence-to-sequence models, especially multimodal vision-language-action models.Hands-on experience in computer vision and deep learning, with work in any of the following areas: multimodal transformers, multimodal language models, diffusion models, NeRF, gaussian splatting, object detection / segmentation, 3D scene understanding, sensor calibration, SfM, voxel/BEV grid-based feature representation.Additional InformationWe offer a competitive base salary for this position with a range in US-California of --$165,000 - $185,000 along with an annual corporate bonus, and a long-term incentive bonus designed to reward sustained impact and contribution over time. Within the salary range, the individual pay is determined based on several factors, including, but not limited to, work experience and job knowledge, complexity of the role, job location, etc.Your well-being matters at Bosch! We offer a a benefits package designed to empower you in every area of your life. This includes premium health coverage, a 401(k) with generous matching, resources for financial planning and goal setting, ample paid time off, parental leave, and comprehensive life and disability protection. Your Recruiter can share more details for this position during the interview process.Learn more about our full benefits offerings by visiting: .Equal Opportunity Employer, including disability / veterans.*Bosch adheres to Federal, State, and Local laws regarding drug-testing. Employment is contingent upon the successful completion of a drug screen and background check. Candidates who have been offered the position must pass both screenings before their start date.#LI-JM1SummaryType: Full-timeFunction: Research

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Research Scientist- Vision-Language-Action (VLA) Models in Sunnyvale, CA vacancy
  • Bosch Group in Sunnyvale, California, invites a Research Scientist- Vision-Language-Action (VLA) Models to advance Embodied AI for ADAS/AD, robotics, and industrial automation. You will conduct research and engineer end-to-end perception and planning, and collaborate with... 
    Language

    Bosch Group

    Sunnyvale, CA
    2 days ago
  • $126k - $423k

     ...looking for multiple passionate Research Scientists to join the Research Group...  ...on pretraining world-action foundation model with various world modalities including vision and physics associated with...  ..., human data incorporation, language modality, and spatial reasoning... 
    Language
    Full time
    For contractors
    For subcontractor
    Casual work
    Work at office
    Immediate start
    Remote work
    Day shift

    Applied Intuition

    Sunnyvale, CA
    2 days ago
  • Bosch USA is seeking a Research Scientist- Vision-Language-Action (VLA) Models to push frontiers in Embodied AI for ADAS/ADAS and robotics. You will conduct research and engineering in core AI fields, exploring modular end-to-end perception and planning, and collaborating... 
    Language

    Bosch USA

    Sunnyvale, CA
    4 days ago
  •  ...Labor for dull, dirty, and dangerous work. The team develops vision-language models and world models to enable safe, real-world robot deployment...  ...and deploying multi-modal perception, planning, and action policies, advancing foundation models, and delivering production... 
    Language
    Work at office

    Embedding VC

    Milpitas, CA
    15 hours ago
  •  ...About the Institute of Foundation Models We are a dedicated research lab for building, understanding, using...  ...world-class researchers, data scientists, and engineers, tackling the most...  ...Summary As a Research Scientist in the Vision Language Model (VLM) team, your role will... 
    Language

    Institute of Foundation Models

    Sunnyvale, CA
    10 days ago
  • $160.36k - $240.54k

     ...flexible, partner-led business model, Nuro is working toward a...  ...models. Leverage large language models and world foundation...  ...lab, in industry, or both.Research experiences in generative models...  ...driving. Experiences in vision-language-action models, reinforcement learning... 
    Language
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    4 days ago
  • $192k - $304.75k

    We are now looking for a Senior Research Scientist focused on Multimodal Foundation Models and Robotics! NVIDIA is...  ...following topics: LLMs; Large vision-language models; Video generative models and diffusion algorithms; or Action-based transformers.Outstanding... 
    Language
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • RoboForce in Milpitas, CA is seeking researchers to advance AI-powered robot control through vision-language models and world-models. You will design and deploy VLMs/VLAs, train world models for long-horizon planning, and integrate multimodal data for natural human-robot... 
    Language
    Work at office

    Socket.dev

    Milpitas, CA
    15 hours ago
  • $218.8k - $335.3k

     ...ready to redefine mobility and shape the future of autonomous transportation? As a Staff Research Scientist specializing in Vision-Language Models (VLMs), Vision-Language-Action models (VLAs), and Onboard Foundational Models, you will advance the frontier of artificial... 
    Language
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    2 days ago
  • The Role We are looking for a Research Scientist to join the Multi-Embodiment Generalist Agent...  ...member. MEGA is building foundation models for general-purpose robots beyond...  ...capabilities. Your day-to-day work may span vision-language-action models, world and action models,... 
    Language
    Full time
    Work from home

    Wayve

    Sunnyvale, CA
    15 hours ago
  • $192k - $304.75k

    We're now looking for a Senior Research Scientist, Multi-Modal Language Models!NVIDIA is seeking a Senior Research Scientist passionate about multi modal...  ...or related areas.4+ years of experiences in computer vision, especially multi-modal LLMs.Proficiency in Python with... 
    Language
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...‑office collaboration. We are looking for a Senior ML Research Engineer, Embodied Intelligence to advance robotic embodied...  ...with humans. Responsibilities Design and deploy vision-language(-action) models (VLM/VLA) for contextual understanding and generalized robot action... 
    Language
    Work at office
    Visa sponsorship

    RoboForce

    Milpitas, CA
    4 days ago
  • $165k - $195k

    Company DescriptionThe Bosch Research and Technology Center North America with offices in Sunnyvale, California,...  ...AI research in Silicon Valley focuses on Foundation Models, Natural Language Processing, Computer Vision & Mixed Reality, Cloud Robotics, Big Data Visual Analytics... 
    Language
    Full time
    Work experience placement
    Local area
    Worldwide

    Robert Bosch

    Sunnyvale, CA
    4 days ago
  • $165k - $185k

    Company DescriptionThe Bosch Research and Technology Center North America with offices...  ...in Silicon Valley focuses on Foundation Models, Big Data Visual Analytics, Explainable AI (XAI), Natural Language Processing, Computer Vision & Mixed Reality, Cloud Robotics, Data Science... 
    Language
    Work experience placement
    Worldwide

    Robert Bosch

    Sunnyvale, CA
    2 days ago
  • Google Beam team in Mountain View is seeking a Senior Research Scientist to advance foundation models and multimodal AI research. You will design large-...  ...at Google. Strong background in deep learning, vision-language models, and responsible AI principles is expected.... 
    Language

    Google

    Mountain View, CA
    2 days ago
  • $151.8k - $218.21k

     ...capabilities development, as well as research impact via open-source...  ...for a driven research scientist with a strong background...  ...large scale foundation models (VLMs, text-to-video models...  ...art in behavior learning, language, and/or computer vision. Some familiarity with robots... 
    Language

    Toyota Research Institute

    Los Altos, CA
    4 days ago
  • $230k - $270k

     ...observation alone, is an open research frontier. Our lab...  ...robot foundation models on our GPUs, and...  ...research : adapting vision-language models to achieve low-...  ...understanding : temporal action detection/segmentation...  ...Exposure to robot learning (VLA models, imitation... 
    Language
    Contract work
    Temporary work
    Worldwide
    Flexible hours

    Samsung SDS America

    Mountain View, CA
    10 days ago
  • $200k - $287.5k

     ...-time /HybridAt Toyota Research Institute (TRI), we’re...  ...Policy and Large Behavior Models (LBM).The OpportunityWe...  ...-of-the-art, pixels-to-action, end-to-end system for...  ...and integrating visual-language-action modalities....  ...with a focus on computer vision as the primary sensing... 
    Language
    Full time
    Local area
    Shift work

    Toyota Research Institute

    Los Altos, CA
    2 days ago
  • Applied Intuition, Inc. in Sunnyvale, CA, seeks multiple Research Scientists to advance next‑gen physical AI for autonomous driving...  ...You will join a high‑caliber team of experts, drive world action foundation models, and contribute to publications and product deployment.... 

    Applied Intuition Inc.

    Sunnyvale, CA
    15 hours ago
  • $192.2k - $260k

     ...practical experience to join the Modeling and Optimization (MOP)...  ...learning, robotics, operations research, statistics, mathematics or equivalent...  ...- Knowledge of programming languages such as C/C++, Python, Java...  ...insurance (medical, dental, vision, prescription, Basic Life &... 
    Language
    Local area
    Flexible hours

    Amazon

    Santa Clara, CA
    4 days ago
  • $165k - $185k

     ...DescriptionThe Bosch Research and Technology...  ...focuses on Foundation Models, Big Data Visual Analytics...  ...AI (XAI), Natural Language Processing, Computer Vision & Mixed Reality,...  ...DescriptionAs a Research Scientist- Robotics AI, you...  ...large vision-language-(action) models to aid... 
    Language
    Work experience placement
    Local area
    Worldwide

    Robert Bosch

    Sunnyvale, CA
    3 days ago
  •  ...TeamThe Content Representation Models team creates a single, unified "language" for Netflix's entire library by...  ...personalizationAbout the RoleWe are looking for a Research Scientist specializing in embeddings and...  ...in LLMs.Experience in computer vision or multimodal AIIndustry... 
    Language
    Hourly pay
    Full time
    Immediate start
    Flexible hours

    Netflix

    Los Gatos, CA
    2 days ago
  • $171.6k - $222.2k

     ...Engineer to build efficient, stable foundation models for long-horizon agentic workloads. The...  ...in Java, C++, Python or related language- Experience in any of the following areas...  ...including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance... 
    Language
    Local area
    Flexible hours

    AmazonWebServices

    Santa Clara, CA
    15 hours ago
  •  ...build frontier foundation models that power intelligent...  ...tasks to precise, action-taking workflows. If you...  ...problems where the research and the product are inseparable...  ...-class engineers and scientists to tackle some of the...  ...learning, including vision-language modeling, image and... 
    Language

    Apple

    Cupertino, CA
    4 days ago
  • $251k - $310k

     ...the use of large multimodal foundation models (e.g., Gemini) to build a powerful...  ...true scene understanding and driving actions—building offboard models that can comprehend...  ...focused deeply on training Large Language Models (LLMs) or Vision-Language Models (VLMs) . Proven... 
    Language
    Full time
    Remote work

    Neura Market

    Mountain View, CA
    15 hours ago
  • $192.2k - $260k

     ...s Delivery Foundation Model team, where you'll work...  ...alongside world-class scientists and engineers to...  ...direction for specific research initiatives, ensuring...  ...combines ambitious research vision with real-world impact...  ...Python, C++ or other languages- Strong publication record... 
    Language
    Local area
    Worldwide
    Flexible hours

    Amazon

    Santa Clara, CA
    1 day ago
  • $230k - $380k

     ...Role We're looking for Research Scientists to join Wayve Labs and help...  ...areas: World & Reward Modeling: Building realistic, diverse...  ...consequences and costs of actions. Representation Learning...  ...and multimodal inputs, using vision, language, and active sensors. Key... 
    Language
    Full time
    Work at office
    Visa sponsorship
    Relocation package
    Flexible hours

    Wayve

    Sunnyvale, CA
    3 days ago
  • $202.5k - $354.4k

     ...DescriptionOur Core AI Research team develops...  ...tools, take reliable action across stateful...  ...We work across LLM model post-training,...  ...a Staff Research Scientist, you will independently...  ...more modalities - language, documents, images...  ...AI, computer vision, speech/audio, multilingual... 
    Language
    Work at office
    Immediate start
    Remote work
    Flexible hours
    Shift work

    ServiceNow

    Santa Clara, CA
    3 days ago
  • $262k - $364k

     ...execute foundational research in learning,...  ...-to-perception-to-action stack calibration....  ...learning, computer vision, and robotics venues...  ...contributing code, models, and datasets...  ...modern programming languages and working with large...  .... As a Research Scientist, you'll setup large... 
    Language

    Google

    Mountain View, CA
    4 days ago
  •  ...Senior / Staff AI Research Scientist, Foundation Models RoboForce is an AI robotics company developing Physical AI–powered Robo-...  ...tasks. Responsibilities Design and deploy vision-language(-action) models (VLM/VLA) for contextual understanding and generalized... 
    Language
    Work at office
    Visa sponsorship

    Embedding VC

    Milpitas, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Research Scientist- Vision-Language-Action (VLA) Models. Be the first to apply!