Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Research Scientist, Video & Multimodal

$160k - $185k
Full-time

Innodata Inc.

Innodata (Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale. We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.

Scope of the Role:

Video is where multimodal models are weakest and hardest to grade. Temporal reasoning, long-form understanding, grounding events in time, and holding audio, video, and text together do not fall out of image benchmarks — and the evaluations for them are still immature. Closing that gap is gated as much by how we design data and evaluation as by architecture. Innodata builds that data and those evaluations for the customers and frontier labs advancing video and multimodal models, and we are hiring a Research Scientist to own the science behind it.

You will partner directly with the customers and frontier labs building video understanding, video-language, and video-generation models, as interested in the data behind them as in the models themselves. Video spans two model families judged in completely different ways: models that understand video — answering questions, localizing events, grounding language in time — where the question is whether the answer is correct; and models that generate it, where fidelity, temporal coherence, and physical plausibility matter and no automatic metric is settled. You own the evaluation science for both, and knowing when model-based scoring can stand in for a human versus when it can't. Your conclusions shape what our partners measure and collect next.

What You’ll Own:

You will define how Innodata designs, structures, and evaluates video data for video and multimodal models, and you will validate those choices experimentally. Concretely, you will:

  • Translate the requirements of video and multimodal models — video understanding, temporal and event localization, action recognition, long-form video, video-language models, video generation, cross-modal reasoning, and multimodal retrieval and grounding — into concrete data specifications: modalities, annotation schemas, sampling, and evaluation criteria.
  • Build evaluation methodology for video understanding — temporal grounding accuracy, long-context and long-horizon reasoning, and dynamic multi-turn, cross-modal, and retrieval-and-grounding evaluation — clear about when model-based scoring is trustworthy and when a human is needed.
  • Build evaluation methodology for video generation — fidelity, temporal coherence, and physical plausibility, including generative video used as a world model — the regime where automatic metrics are weakest and human judgment matters most.
  • Decide how existing and incoming video should be structured, enriched, and sampled to extract the most model value from it, including from messy, domain-specific footage.
  • Run experiments that prove data decisions matter: fine-tune and evaluate models on Innodata data, with ablations tying specific data choices to measurable improvement.
  • Design adversarial and stumping evaluations that surface where video and multimodal systems fail, and turn those failures into better data.
  • Publish. Turn what you learn into benchmarks, methodology, and papers that advance the field and earn the trust of the customers and frontier labs we partner with.
  • Work with annotation teams, subject-matter experts, and the synthetic-data pipeline to turn specifications into operational collection and labeling plans.

You’ll Thrive in This Role If You Have:

  • Roughly 5+ years of hands-on industry experience in video understanding or multimodal ML. We weight practical experience over formal credentials; a PhD with a compelling, current research agenda can offset the lower end.
  • A Bachelor's degree in computer science, electrical engineering, or a related technical or quantitative field is required; an advanced degree (MS or PhD) in a relevant field is preferred.
  • Trained and evaluated video or multimodal models yourself, with strong PyTorch fundamentals.
  • Fluency in the formats and tooling video work runs on: ffmpeg and decord pipelines, temporal and COCO-style annotation, WebDataset, Parquet and Arrow, and HuggingFace datasets.
  • Experience fine-tuning large video or vision-language models with the modern toolchain (HuggingFace transformers, PEFT, efficient inference), and with long-form video, streaming, temporal segmentation, or synthetic video generation.
  • A way of thinking in datasets and benchmarks: you have built evaluation sets, calibrated difficulty, and argued about what makes video data good for a given objective.
  • A track record the field recognizes: first-author publications or strong open-source contributions at venues such as CVPR, ICCV, ECCV, NeurIPS, or ICLR.
  • The ability to work directly with the research scientists at the customers and frontier labs we partner with, and to explain data and modeling decisions clearly to both expert and non-expert audiences, backed by a rigorous, reproducible approach to experiments and documentation.
  • Bonus: interest or hands-on experience in responsible-AI evaluation and red-teaming — safety and robustness testing for video and multimodal systems.

The expected salary range for this position is $160,000 - $185,000 p/year, based on experience, skills, and qualifications.

Please be aware of recruitment scams involving individuals or organizations falsely claiming to represent employers. Innodata will never ask for payment, banking details, or sensitive personal information during the application process. To learn more on how to recognize job scams, please visit the Federal Trade Commission’s guide at

If you believe you’ve been targeted by a recruitment scam, please report it to Innodata at View email address on aiapply.co and consider reporting it to the FTC at ReportFraud.ftc.gov .

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Research Scientist, Video & Multimodal in Remote vacancy
  • $180k - $300k

     ...more details, check out our recent research on synthetic data scaling ( BeyondWeb...  ...Role We're looking for a Research Scientist to investigate how intervening on...  ...~ Training large vision (including video), language, or multimodal models Efficient ML Enough... 
    Video
    Work at office
    Work from home
    Relocation package

    DatologyAI

    San Mateo, CA
    1 day ago
  • $155k - $269k

     ...scalable, controllable, and efficient simulation. As a Research Scientist in World Models, you will develop algorithms and productionize...  ...models for temporal reasoning and generation, including video models, multimodal generative models, LLM/VLM/VLA models, and predictive... 
    Video
    Full time
    Work at office
    Work from home
    Flexible hours

    Waabi

    San Francisco, CA
    2 days ago
  •  ...hours of continuous, time-synchronized multimodal data. Nothing comparable exists publicly...  ...built. We're looking for an experienced researcher to be a driving force behind this work:...  ...post-training multimodal models (vision, video, audio, language) — SFT, RLHF, DPO or GRPO... 
    Video
    Work visa

    Noösphere

    Seattle, WA
    2 days ago
  •  ...Descript's Research team builds the models behind the product's most distinctive features: Video Regenerate and lipsync, video translation, zero-shot voice and roomtone cloning...  ...within months. This role is focused on multimodal understanding: training models to perceive... 
    Video
    Full time

    Descript

    Remote
    18 days ago
  • $251k - $310k

     ...foster collaborations with other research teams in Alphabet. AI...  ...reports to a Principal Research Scientist. You will : Research...  ...& Develop state-of-the-art Multimodal LLMs and World models to...  ...such as world models, images, videos, 3D, using techniques such as... 
    Video
    Full time
    Temporary work
    Remote work

    Waymo

    Kirkland, WA
    2 days ago
  • This AI Research Scientist will lead the design and build biological foundation models that learn...  ...based approaches for high-dimensional or multimodal data (e.g., time-series, imaging, or...  ...preferred. Experience with image processing, video processing, signal processing, and/or... 
    Video
    Remote work
    3 days per week

    Q-state Biosciences

    Cambridge, MA
    3 days ago
  • $218.8k - $335.3k

     ...a global scale. Role : As a Staff Research Scientist in the AI Research organization, you will...  ...models, VLAs, diffusion, image/video generation, self-supervised learning, IL...  ...architectures (transformers, generative AI, multimodal systems). ~ Strong programming skills... 
    Video
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    3 days ago
  • $159.75k - $255.6k

     ...at a company where you matter. Your Impact We are seeking a skilled and innovative Senior AI Research Scientist to join a new team focusing on agentic video and multimodal reasoning systems. As a research scientist at Axon you will play a crucial role in developing AI... 
    Video
    Work experience placement
    Work at office
    Remote work

    University of Georgia- FACS

    Seattle, WA
    1 day ago
  • $142.8k - $193.2k

     ...Prime Video is a first-stop entertainment destination offering customers a vast collection...  ....Key job responsibilitiesAs an Applied Scientist, you will have access to large datasets...  ...or representation learning (embeddings, multimodal models)- Experience with online... 
    Video
    Flexible hours

    Amazon

    Seattle, WA
    1 day ago
  • $164k - $184k

     ...Insights is the team that turns that data into research, executive narrative, and category...  .... We are hiring a Senior Research Scientist to lead studies from question to published...  ...interview with the Hiring Manager ~ Three video interviews with several team members - 2... 
    Video
    Remote work
    Flexible hours

    Syndio

    Washington DC
    4 days ago
  • $231.5k - $405.1k

     ...About the team  Our Core AI Research team develops novel methods...  ...enterprise agents that reason over multimodal information, use tools, take...  ...role  As a Staff Research Scientist, you will independently lead...  ...language, documents, images/video, and speech/audio - and... 
    Video
    Full time
    Work at office
    Immediate start
    Remote work
    Flexible hours
    Shift work

    ServiceNow

    Santa Clara, CA
    27 days ago
  • $213k - $263k

     ...the Waymo Driver. We focus on building an onboard multi-task, multimodal perception model designed to tackle highly complex and unpredictable...  ...temporal reasoning for sequential decision-making or complex video understanding. First-author publications in premier computer... 
    Video
    Full time
    Remote work

    Waymo

    Remote
    more than 2 months ago
  •  ...experience for our TikTok users.Responsibilities:• Conduct cutting-edge research in machine learning algorithms, such as retrieval and...  ...commodity recommendations, live stream recommendations, short video recommendations etc in TikTok.• Build long and short term user... 
    Video
    Temporary work
    Work experience placement

    Tik Tok

    Seattle, WA
    3 days ago
  • $185k - $400k

     ...infrastructure built around real-time, multimodal generation and intelligent agentic platforms. We are seeking accomplished Research Scientists in Foundation Models with expertise in...  .../mid-training (text, image, audio, and video), and drive innovative approaches for foundational... 
    Video
    Remote work

    Pika

    Palo Alto, CA
    4 days ago
  • $185k - $400k

     ...infrastructure built around real-time, multimodal generation and intelligent agentic platforms...  ...are looking for a staff or lead-level Research Engineer, Data to architect and scale...  ...research workflows for text, image, audio, and video datasetsPartner with research and... 
    Video
    Remote work

    Pika

    Palo Alto, CA
    2 days ago
  • $230k - $400k

     ...is a quickly growing group of committed researchers, engineers, policy experts, and business...  ...other modes of data, including images, video and audio. Such models have potential to...  ...about the risks introduced by powerful multimodal AIs. The Multimodal team builds and studies... 
    Video
    Work experience placement
    Work at office
    Home office
    Visa sponsorship
    Relocation package
    Flexible hours

    TalentPros.AI

    San Francisco, CA
    7 days ago
  • $224k - $356.5k

     ...Curator team is seeking a Senior Applied Research Scientist with experience researching,...  ...modal data (documents, image, audio and videos) used in the training of foundation models...  ...methodologies targeting petabyte-scale multimodal data run across hundred-node GPU clusters... 
    Video
    Full time
    Work at office
    Remote work
    Flexible hours

    Nvidia

    Santa Clara, CA
    1 day ago
  • TELUS Digital is seeking a Multimodal AI Content Specialist to join our Global Community. In this role, you'll evaluate the relationship...  ...culturally resonant. Your responsibilities include auditing image and video datasets, conducting safety and bias detection, and helping... 
    Video
    Remote job

    Jazmin Par

    New York, NY
    5 days ago
  • $230k - $270k

     ...floor, in real time and from observation alone, is an open research frontier. Our lab researches the AI systems that make...  ...robot is doing something wrong — built on multi-camera video understanding and multimodal signals from the operating environment and the robot... 
    Video
    Contract work
    Temporary work
    Worldwide
    Flexible hours

    Samsung SDS America

    Mountain View, CA
    9 days ago
  •  ...Physical AI model is trained on petabytes of video, lidar, radar, and sensor data. Today's...  ...engine, Daft , is purpose-built for multimodal AI: 2 PB/day at Amazon, 60-100 PB at...  ...that empower robotics operators and AI researchers to review, annotate, and analyze complex... 
    Video
    Full time
    Work at office
    Immediate start
    Flexible hours
    Night shift

    Eventual

    Remote
    22 days ago
  • $92k - $111k

     ...Johnson Innovative Medicine is recruiting for a Postdoctoral Scientist - Multimodal Representation Learning for Predictive Biology to join the...  ...-date code repository. Draft manuscripts and disseminate research findings internally and externally (e.g., publishing in peer... 
    Temporary work
    Fixed term contract
    Local area
    Remote work

    Johnson & Johnson Innovative Medicine

    Spring House, PA
    2 days ago
  • $86.15k - $106.65k

     ...Scientist I – Neuromodulation Anatomy: Multimodal Data Integration The Allen Institute accelerates science for a healthier world through large-scale research designed to answer some of the most complex questions in biology. Our multi-disciplinary teams generate foundational... 
    Work at office
    Local area
    Remote work
    Visa sponsorship
    Work visa
    Relocation package
    Flexible hours

    Allen Institute

    Washington DC
    3 days ago
  • $86.5k - $106.5k

     ...Scientist I – ML/AI algorithms for Multimodal Foundational Models for Gene Regulation The Allen Institute accelerates science for a healthier world through large-scale research designed to answer some of the most complex questions in biology. Our multi-disciplinary... 
    Work at office
    Local area
    Remote work
    Visa sponsorship
    Relocation package

    Allen Institute

    Seattle, WA
    more than 2 months ago
  • $124.8k - $171.6k

     ...SUMMARY: Natera is hiring a Machine Learning Scientist to join our AI and computational biology...  ...DNA (cfDNA) modalities. You will build multimodal AI systems that integrate imaging,...  ...) - Proven ability to take ownership of research projects and translate prototypes into robust... 
    Full time
    Work at office
    Remote work

    Natera

    Remote
    18 days ago
  • $123.05k

     ...Job Description We have an opening for a Postdoctoral Researcher to work in the field of computational materials science and actively...  ...personal information LLNL collects and for what purpose. The Employee Privacy Notice can be accessed here. Videos To Watch... 
    Video
    For contractors
    Relocation package
    Flexible hours

    LLNL

    Livermore, CA
    1 day ago
  • $35 per hour

     ...referral partner. We refer candidates to our partner that collaborates with world’s leading AI research labs to build and train cutting-edge AI models. Position: Video Filtering Expert Type: Hourly Contract Compensation: $35/hour Location: Remote Duration... 
    Video
    Hourly pay
    Contract work
    Remote work
    Flexible hours

    Crossing Hurdles

    United States
    4 days ago
  • $35 per hour

     ...top candidates to Mercor that work with the world’s leading AI research labs to help build and train cutting‑edge AI models. Base Pay...  ...$35.00/hr – $35.00/hr Organization: Mercor Position: Video Filtering Expert Referral Partner: Crossing Hurdles Type... 
    Video
    Hourly pay
    Contract work
    Remote work

    Crossing Hurdles

    United States
    4 days ago
  • $110k - $130k

     ...Department Overview The New York Times is hiring a full-stack Software Engineer for the News Product Multimodal team. In this role, you will develop our audio and video features into a truly multimodal experience, shaping how tens of millions of people watch, listen to,... 
    Video
    Full time
    Local area
    Flexible hours

    The New York Times Company

    Remote
    more than 2 months ago
  • $35.6 - $54.93 per hour

     ...Job Description: The Medical Technologist/Medical Laboratory Scientist performs a variety of laboratory tests of varying complexity to...  ...it mean to be a caregiver with Intermountain? Check out this video and learn more and discover the “Power of We.” Job Specifics... 
    Video
    Hourly pay
    Full time
    Temporary work
    Internship
    Remote work
    Shift work
    Night shift

    Intermountain Health

    Grand Junction, CO
    19 hours ago
  •  ...of objectives and modalities (LLM, VLM, video, and action) at massive scale. Working alongside...  ...world-class team, you'll help set the research agenda, own high-impact architecture,...  ...pre-training large models—LLM, VLM, multimodal, or VLA—at significant scale, with concrete... 
    Video
    Work from home
    Flexible hours

    Walden Robotics

    Cambridge, MA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Research Scientist, Video & Multimodal. Be the first to apply!