Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

GenAI Research Scientist — Multimodal Vision & Audio

$244.8k

ByteDance

ByteDance is looking for talented individuals in San Jose to join the Intelligent Creation - Global GenAI team. As a researcher, you will work on generative AI projects, focusing on applied research and innovative solutions for TikTok. The successful candidate will hold a PhD in a related field and will conduct research in computer vision and machine learning. The position offers a competitive salary range of $244,800 to $588,000 annually and several benefits including medical insurance and paid personal time. #J-18808-Ljbffr ByteDance

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the GenAI Research Scientist — Multimodal Vision & Audio in San Jose, CA vacancy
  • $212.8k - $387.6k

    Team Overview The Seed Multimodal Interaction and World Model team is dedicated to developing models that have human-level multimodal...  ...Develop multimodal foundation models integrating vision, language, audio, and environment signals. Design and optimize world models... 
    Audio
    Temporary work
    Internship
    Local area

    ByteDance

    San Jose, CA
    1 day ago
  • SpeedyApply LLC is seeking graduates and PhD candidates to advance GenAI and content understanding within our TikTok business-focused team. You will contribute to research on multimodal models and autonomous planning, shaping how content and recommendations are generated... 
    Suggested

    SpeedyApply LLC

    San Jose, CA
    10 hours ago
  • $212.8k

    TikTok is looking for talented individuals to join our GenAI team in San Jose, CA. As a PhD graduate, you will work on advanced projects involving computer vision and natural language processing, contributing to our mission of empowering TikTok businesses through innovative... 
    Suggested

    TikTok

    San Jose, CA
    1 day ago
  • ByteDance in San Jose seeks a PhD candidate for the Seed Vision team focused on visual generation models. Responsibilities include developing...  ...skills in C/C++ or Python, experience in computer vision or multimodal learning, and a strong understanding of deep learning... 
    Suggested

    ByteDance

    San Jose, CA
    1 day ago
  • $147k - $211k

    SnapshotWe are seeking strong Research Scientists with expertise in AI...  ...sociotechnical modeling to join a multimodal safety research effort within...  ...in deep learning, computer vision, and generative architectures...  ...data types (e.g., vision, audio, text).The US base salary range... 
    Audio
    Full time

    DeepMind

    Mountain View, CA
    1 day ago
  • $180k

     ...intelligence. One that is proactive, multimodal, and capable of interacting...  ...world through speech, text, vision, and persistent memory. We'...  .... Responsibilities Drive research and development to advance...  ...systems (vision + text, vision + audio) or real‑time AI systems is a... 
    Audio
    Full time

    Hark

    San Jose, CA
    3 days ago
  • $244.8k

     ...The Intelligent Creation - Global GenAI team focuses on applied research in Generative AI, delivering intelligent...  ...with ease. The team works on multimodal foundation models, image and video...  ...Benefits include medical, dental, and vision insurance; 401(k) with company match... 
    Temporary work
    Local area

    TikTok

    San Jose, CA
    1 day ago
  • Ifm Us is seeking a Research Scientist to advance Vision Language Models. This role focuses on research and development in multimodal AI, integrating visual understanding with language reasoning. The successful candidate will contribute to technical reports, mentor junior... 

    Ifm Us

    Sunnyvale, CA
    3 days ago
  • Honda Research Institute USA, Inc. is seeking a Research Scientist in San Jose, CA. This role aims to advance human-centric visual intelligence through AI systems...  ...the opportunity to work at the forefront of multimodal AI research. #J-18808-Ljbffr Honda Research Institute... 

    Honda Research Institute USA, Inc.

    San Jose, CA
    1 day ago
  • Apple in Sunnyvale is seeking a Multimodal AI Researcher to push the boundaries of foundation models for real-time multimodal data, including video, audio, and text. You will work on interactive models, audio-to-audio modeling, and streaming multimodal systems, driving... 
    Audio

    Socket.dev

    Sunnyvale, CA
    1 day ago
  • $224k - $356.5k

     ...never been done before takes vision, innovation, and the world’s...  ...is seeking a Senior Applied Research Scientist with experience researching,...  ...data (documents, image, audio and videos) used in the training...  ...targeting petabyte-scale multimodal data run across hundred-node... 
    Audio
    Full time
    Work at office
    Remote work
    Flexible hours

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...is hiring Senior Deep Learning Scientists to advance our efforts in streaming and agentic multimodal AI. You will demonstrate...  ...Apply fundamental and applied research to develop, train, fine-tune,...  ...agentic systems encompassing audio-visual reasoning, tool usage,... 
    Audio
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...understand everything from video and audio to text and artwork at a semantic level...  ...the RoleWe are looking for a Research Scientist specializing in embeddings and representation...  ...them in LLMs.Experience in computer vision or multimodal AIIndustry experience in... 
    Audio
    Hourly pay
    Full time
    Immediate start
    Flexible hours

    Netflix

    Los Gatos, CA
    4 days ago
  • $165k - $195k

    Company DescriptionThe Bosch Research and Technology Center North America with offices in Sunnyvale...  ..., Natural Language Processing, Computer Vision & Mixed Reality, Cloud Robotics, Big Data...  ...AI & big data analytic solutions (e.g., audio, images, sensor logs) for a range of... 
    Audio
    Full time
    Work experience placement
    Local area
    Worldwide

    Robert Bosch

    Sunnyvale, CA
    1 day ago
  •  ...solutions that enhance content discovery. The role requires a strong foundation in AI/ML, with experience in training and deploying vision-language models. You will work closely with cross-functional teams and need to demonstrate excellent communication skills. The... 
    Flexible hours

    Netflix

    Los Gatos, CA
    1 day ago
  • $244.8k

    About the Team The Seed Vision team focuses on foundational models for visual generation, developing multimodal generative models, and carrying out leading research and application development to solve...  ...computer vision challenges in GenAI. Responsibilities Develop and... 
    Temporary work
    Local area

    ByteDance

    San Jose, CA
    1 day ago
  • Job Number: P25F07 Honda Research Institute USA (HRI-US) is seeking a Research Scientist to develop AI methods for sensing, modeling, and interpreting...  ...noisy, sparse, and imperfect real‑world multimodal data, including vision, audio, language, interaction traces, and... 
    Audio

    Honda Research Institute USA, Inc.

    San Jose, CA
    1 day ago
  • $192k - $304.75k

    We are now looking for a Senior Research Scientist focused on Multimodal Foundation Models and Robotics! NVIDIA is searching for an outstanding research...  ...in at least one of the following topics: LLMs; Large vision-language models; Video generative models and diffusion... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...foundation models that understand video, audio, text, and artwork at a semantic...  ...Netflix. About the Role We are seeking a Research Scientist specializing in embeddings and...  ...grounded in LLMs. Experience in computer vision or multimodal AI. Industry experience in recommendation... 
    Audio
    Full time
    Flexible hours

    Netflix, Inc.

    Los Gatos, CA
    20 hours ago
  •  ...surging complexity in multilingual and multimodal content, and upgraded generative adversarial...  ...for multimodality (text/image/video/audio), Unified Understanding & Generation, and...  ...paradigm for this direction and produce research outcomes with significant industry influence... 
    Audio
    Flexible hours
    Shift work

    SpeedyApply LLC

    San Jose, CA
    10 hours ago
  • $156k - $387.6k

     ...representation, retrieval, and multimodal fusion. Partnering closely...  ...product teams, we turn advanced research into scalable production...  ...multimodal data (text, image, audio, behavioral signals), enhancing...  ...access to medical, dental, and vision insurance, a 401(k) savings... 
    Audio
    Temporary work
    Local area

    ByteDance

    San Jose, CA
    1 day ago
  • $254.4k

     ...technology and society. With a long-term vision for the AI sector, the Seed team's research spans MLLM, GenMedia, AI for...  ...models and cutting-edge multimodal capabilities. Our technology powers...  ...product development in speech and audio, music, natural language understanding... 
    Audio
    Temporary work
    Local area
    Worldwide

    ByteDance

    San Jose, CA
    3 days ago
  • $238.9k - $305.5k

     ...Advanced Technology Group (ATG) is the research division of the company. ATG’s...  ..., digital signal processing, audio engineering, image processing, computer vision, data science & analytics,...  ...Applications in vision, audio, or multimodal domains (e.g., source separation... 
    Audio
    Full time
    Local area
    Worldwide
    Flexible hours

    Dolby

    Sunnyvale, CA
    1 day ago
  •  ...Responsibilities About the team The Seed Vision Team focuses on foundational...  ...visual generation, developing multimodal generative models, and carrying out leading research and application development to...  ...computer vision challenges in GenAI. Researching and developing foundational... 
    Internship

    ByteDance

    San Jose, CA
    3 days ago
  • $192k - $304.75k

    We are now looking for a Senior Research Scientist for Generative AI!NVIDIA is searching for a world...  ..., video generation, 3D generation, and audio generation. You will be working with a...  ...practice of deep learning, computer vision, natural language processing, or computer... 
    Audio
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • Adobe seeks Applied Scientists at varying levels to drive AI innovations from research to product validation. You will design and evaluate multimodal architectures, focusing on image/video editing, while directly contributing to product strategy and technical mentorship... 

    Adobe

    San Jose, CA
    2 days ago
  • $187.04k

    A leading tech company is seeking a Research Scientist for Multimodal Interaction and World Model. This role focuses on enhancing models for multimodal...  ..., plus benefits that include medical, dental, and vision insurance, a 401(k) plan, and paid personal time. #J-1880... 

    ByteDance

    San Jose, CA
    2 days ago
  • $55 per hour

    Research Scientist Intern - Multimodal Sensing & On-Device Perception Location: San Jose Team: Technology Employment Type: Intern Responsibilities...  ...Design and train machine learning models, utilizing computer vision and language modeling to interpret complex spatio‑... 
    Hourly pay
    Internship

    Socket.dev

    San Jose, CA
    3 days ago
  •  ...-visibility projects that strengthen competitiveness in generative AI. Responsibilities include designing training pipelines for multimodal models including images and videos, while collaborating closely with various teams to enhance quality. The ideal candidate holds... 

    Adobe Inc.

    San Jose, CA
    1 day ago
  • $192k - $304.75k

    We're now looking for a Senior Research Scientist, Multi-Modal Language Models!NVIDIA is seeking...  ...together, such as text, image, video, audio, etc …Design solutions that improve pareto...  ....4+ years of experiences in computer vision, especially multi-modal LLMs.Proficiency... 
    Audio
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to GenAI Research Scientist — Multimodal Vision & Audio. Be the first to apply!