Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Multimodal Hallucination Model Evaluator [Remote]

AuraOne Human Data

Remote
  • Remote job

Multimodal Hallucination Model Evaluator is a remote evaluation track for reviewing multimodal hallucination model evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain.

Why this role matters

AI data reviewers help turn multimodal hallucination model evaluation outputs into auditable labels, rationales, and regression cases for AuraOne Human Data.

Responsibilities

  • Evaluate multimodal hallucination model evaluation model outputs against a versioned rubric and assign severity tags for Multimodal Hallucination Model Evaluator assignments.
  • Compare paired responses and pick the stronger answer with a written rationale.
  • Label hallucinations, instruction-following failures, and unsafe content with structured tags.
  • Capture ambiguous prompts and route them back to the program team for rubric updates.
  • Maintain reviewer-quality scores by calibrating against gold-standard examples each week.
  • Document recurring failure modes so the modeling team can target them in the next training run.

Qualifications

  • Prior evaluation, annotation, or human-rater experience on multimodal hallucination model evaluation or adjacent content for Multimodal Hallucination Model Evaluator work.
  • Comfort applying multi-page rubrics consistently across long batches.
  • Clear written reasoning that names the issue and the rubric clause being applied.
  • Strong attention to detail and the ability to flag when a prompt itself is the problem.
  • Reliable async availability for at least 10 hours per week.

Example tasks

  • Compare two multimodal hallucination model evaluation model responses to the same prompt and pick the stronger one with rationale.
  • Tag an unsafe response with the correct policy category and severity.
  • Audit a 50-row batch for rubric consistency and report drift to the program lead.
  • Propose a rubric clarification after spotting a recurring failure mode.

Nice to have

  • Background in linguistics, content moderation, or trust & safety review.
  • Experience with inter-rater agreement metrics and calibration cycles.
  • Domain expertise that lets you spot subject-matter errors automated checks miss.

Skills

  • Model output evaluation
  • Rubric-based annotation
  • Severity tagging
  • Inter-rater calibration
  • Multimodal Hallucination Model evaluation
  • Multimodal evaluation
  • Cross-modal reasoning
  • Grounding review
  • Multimodal
  • Hallucination

Work model

Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.

Compensation

Hourly rate confirmed after the interview process.

Application process

Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.

Vacancy posted 15 days ago
Similar jobs that could be interesting for youBased on the Multimodal Hallucination Model Evaluator [Remote] in Remote vacancy
  • Receipt and Invoice Understanding Model Evaluator is a remote evaluation track for reviewing receipt and invoice understanding model evaluation...  ...Receipt and Invoice Understanding Model evaluation Multimodal evaluation Cross-modal reasoning Grounding review Receipt... 
    Suggested
    Hourly pay
    For contractors
    Remote work
    10 hours per week

    AI Trainer Jobs

    New York, NY
    5 days ago
  •  ...Workflow Annotator—Product Management & Marketing to remotely review evaluation prompts and responses against the company's quality rubric....  ..., label edge cases, and provide structured feedback the modeling team can use to retrain. As a contractor, you will assess frontier... 
    Suggested
    Remote job
    For contractors

    AI Trainer Jobs

    New York, NY
    5 days ago
  •  ...Medical Document OCR Model Evaluator is a remote clinical-review track for evaluating AI outputs that touch clinical review. Reviewers...  ...assessment Structured clinical writing Clinical review Multimodal evaluation Cross-modal reasoning Grounding review... 
    Suggested
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    12 days ago
  • AI Trainer Jobs is seeking a remote contractor to evaluate people ops / recruiting prompts and responses against a evolving quality rubric. You will compare model outputs, label edge cases, and provide structured feedback for retraining. Responsibilities include evaluating... 
    Suggested
    Remote job
    For contractors

    AI Trainer Jobs

    New York, NY
    5 days ago
  • AI Trainer Jobs is seeking an Illustration Quality Evaluator for a remote contractor role. Review illustration quality evaluation prompts and responses against a published rubric, compare paired outputs, and provide structured feedback to support retraining efforts. You... 
    Suggested
    Remote job
    Part time
    For contractors

    AI Trainer Jobs

    New York, NY
    5 days ago
  • $50 per hour

    Combinatorics Model Evaluator is a remote review track for evaluating AI outputs across combinatorics model research review reasoning, calculations, and research workflows. Reviewers grade derivations and assumptions, reproduce key results, and document the correct method... 
    Hourly pay
    For contractors
    Remote work

    AI Trainer Jobs

    New York, NY
    5 days ago
  • $15 - $20 per hour

     ...public sources and external tools. Generate high-quality human evaluation data by identifying response strengths, areas for improvement,...  ...quality, clarity, tone, and completeness of responses. Ensure model responses align with expected conversational behavior and system... 
    Contract work
    Summer work
    Remote work

    Remote Jobs

    New York, NY
    5 days ago
  • Surgical Planning Safety Evaluator is a remote evaluation track for reviewing surgical planning safety evaluation prompts and responses...  ...edge cases, and write the kind of structured feedback the modeling team can use to retrain. AI data reviewers help turn surgical... 
    Remote job

    AI Trainer Jobs

    New York, NY
    5 days ago
  •  ...Scientific Figure Understanding Model Evaluator is a remote review track for evaluating AI outputs across scientific figure understanding...  ...Scientific Figure Understanding Model research review Multimodal evaluation Cross-modal reasoning Grounding review Scientific... 
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    15 days ago
  • AuraOne is seeking a Latin Bilingual Expert for a remote evaluation track. Reviewers compare paired outputs to a quality rubric, label edge cases, and generate structured feedback to retrain models. Responsibilities include evaluating outputs against rubrics, tagging issues... 
    For contractors
    Remote work
    10 hours per week

    AI Trainer Jobs

    New York, NY
    5 days ago
  • $80 per hour

     .../hour Location: Remote Role Responsibilities Use frontier AI coding agents to complete and evaluate complex data engineering tasks. Review model-generated implementations involving ETL pipelines , data warehouses , analytics platforms , and distributed... 
    Contract work
    Summer work
    Remote work

    Mercor

    New York, NY
    9 days ago
  • Prolific Academic Ltd is seeking Licensed Pharmacists to assist in AI model training and evaluation from a home office. Successful candidates join as Domain Experts and will be paid to train and evaluate models, with tests to assess suitability. Experts review AI-driven... 
    Remote work
    Home office
    Flexible hours

    Prolific Academic Ltd

    New York, NY
    2 days ago
  • $65 per hour

    Prolific is seeking registered nurses in Las Vegas, NV, to assist in training AI models. You will review AI responses, rate their accuracy and safety, and write feedback to enhance AI learning. Candidates must be verified registered nurses, have recent clinical experience... 
    Hourly pay
    Self employment
    Work from home
    Flexible hours

    Prolific

    Las Vegas, NV
    3 days ago
  • $80 - $150 per hour

    Prolific Academic Ltd is recruiting Medical Doctors to help train and evaluate AI models. You’ll complete a quick test to assess suitability and, if successful, join as a Domain Expert, paid to work on AI tasks. Researchers pay $80-$150 per hour per completed task, with... 
    Hourly pay
    Self employment
    Work from home

    Prolific Academic Ltd

    New York, NY
    2 days ago
  •  ...interaction, enabling the human touch where it has been previously unscalable. We achieve this through pioneering research in multimodal AI for modeling human-to-human communication (language, audio, and video), as well as generating audio-visual avatar behavior. Our models... 
    Full time
    Remote work
    Relocation package
    Flexible hours

    Tavus

    San Francisco, CA
    9 days ago
  •  ...Refusal Preference Reward Model Evaluator is a remote red-team track for stress-testing AI systems against adversarial prompts. Reviewers craft attack scenarios, document the failure mode, and pair each successful jailbreak with the rubric clause it violated so the safety... 
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    21 hours ago
  •  ...Quant Finance Reasoning Model Evaluator is a remote review track for evaluating AI outputs across finance and risk workflows. Reviewers grade calculations, narrative reasoning, and policy adherence; flag compliance and reconciliation issues; and document the correct treatment... 
    Remote job
    Hourly pay
    For contractors
    Work experience placement
    10 hours per week

    AuraOne Human Data

    Remote
    1 day ago
  •  ...Policy Preference Reward Model Evaluator is a remote review track for evaluating AI outputs in policy review workflows. Reviewers grade citation accuracy, statutory reasoning, and policy adherence; flag risk; and document the corrected analysis so the modeling team can... 
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    1 day ago
  • $185k - $325k

     ...aviation. About You:You want to build a model that understands what happens next. Not...  ...machinery that turns those models into evaluated candidate futures.Responsibilities:Design...  ...dynamics, aircraft performance or air traffic.Multimodal fusion across vision, state and... 
    Full time
    Remote work
    Relocation

    Merlin Labs

    Boston, MA
    1 day ago
  • $219k - $351k

     ...becoming a memory-bandwidth business. As models scale past what any single GPU can hold...  ...sliding-window/hybrid layers, diffusion and multimodal transformers — and how each changes the...  ...managers to ensure every candidate is evaluated fairly and holistically.Recruiting... 
    Work at office
    Remote work
    Flexible hours

    Samsung Semiconductor

    San Jose, CA
    4 days ago
  • $85 per hour

     ...Location: Remote Role Responsibilities Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks. Review model-generated implementations involving cloud platforms , Kubernetes , CI/CD systems , observability... 
    Contract work
    Summer work
    Remote work

    Mercor

    New York, NY
    a month ago
  • About the role Conversational AI Evaluator is a remote evaluation track for reviewing conversational...  ...the kind of structured feedback the modeling team can use to retrain. AI data...  ...Conversational AI evaluation Voice, language and multimodal AI evaluation Rubric writing Expert... 
    Hourly pay
    For contractors
    Remote work
    10 hours per week

    AI Trainer Jobs

    New York, NY
    3 days ago
  • $162k - $243k

     ...efficient inference of large-scale foundation models.We are seeking a Staff Engineer - AI...  ...for LLMs, VLMs, diffusion, and multimodal models on Qualcomm inference accelerators...  ...performance/accuracy trade-off analysis, or evaluation frameworks.· PhD in a relevant field.Minimum... 
    Work experience placement
    Work from home

    Jobleads-US

    Austin, TX
    1 day ago
  •  ...reasoning and computational problem solving to improve and evaluate large language models. You will design rigorous math problems, produce clear,...  ...specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents. Help enterprises move AI from proof of... 
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $149 per hour

     ...Vice President - Technical AI Foundation Model EngineerRole SummaryThe VP, Technical AI...  ...and semantic search capabilities.Evaluate, benchmark, and recommend foundation models...  ...ManagementComplianceAuditabilityImplement controls for:Hallucination mitigationPrompt securityModel... 
    Full time
    Work at office
    Local area
    Remote work
    1 day per week

    MUFG

    Jersey City, NJ
    1 day ago
  •  ...Apply advanced chemistry expertise to improve large language models by creating, solving, and clearly explaining complex chemistry...  ...contract role combines analytical work, English comprehension, and multimodal scientific communication using text, images, chemical equations... 
    Contract work
    For contractors
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $125k - $150k

     ...DescriptionEverforth ECS is seeking an AI Model Engineer to work in a hybrid remote/...  ...candidate is innovative with a track record of evaluating, experimenting with, and transitioning...  ..., and classification tasks.Utilize multimodal architectures for robust retrieval and data... 
    Contract work
    Work at office
    Remote work

    ECS Federal

    Fairfax, VA
    3 days ago
  • $50 - $300 per day

    About us We are professional, agile and professional. Our work environment includes: Modern office setting Food provided Growth opportunities Responsibilities Perform in various theatrical productions, including plays, musical, and other live performances Memorize lines...
    Contract work
    Internship
    Work at office
    Remote work
    Flexible hours
    Shift work
    Night shift
    Day shift

    Think Entertainment Jobs

    Detroit, MI
    2 days ago
  • $130.1k - $174.91k

     ...purpose.Job Description:Responsible for analytical, statistical modeling, and forecasting methods.The primary requirement is not...  ...Independently determines and develops approach to solutions.Work is evaluated upon completion for adequacy in satisfying objectives.... 
    Part time
    Local area
    Remote work

    Freddie Mac

    McLean, VA
    1 day ago
  •  ...Physician Services PC · Mental HealthValhalla, NYAllied Health Prof/TechnicalPer DiemAll ShiftsAs neededJob Summary: The Mental Health Evaluator (MHE) is a critical role dedicated to delivering comprehensive, clinically appropriate, culturally competent, and trauma-informed... 

    HealthAlliance of the Hudson Valley

    Valhalla, NY
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Multimodal Hallucination Model Evaluator [Remote]. Be the first to apply!