Multimodal Hallucination Model Evaluator [Remote]
AuraOne Human Data
- Remote job
Multimodal Hallucination Model Evaluator is a remote evaluation track for reviewing multimodal hallucination model evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain.
Why this role matters
AI data reviewers help turn multimodal hallucination model evaluation outputs into auditable labels, rationales, and regression cases for AuraOne Human Data.
Responsibilities
- Evaluate multimodal hallucination model evaluation model outputs against a versioned rubric and assign severity tags for Multimodal Hallucination Model Evaluator assignments.
- Compare paired responses and pick the stronger answer with a written rationale.
- Label hallucinations, instruction-following failures, and unsafe content with structured tags.
- Capture ambiguous prompts and route them back to the program team for rubric updates.
- Maintain reviewer-quality scores by calibrating against gold-standard examples each week.
- Document recurring failure modes so the modeling team can target them in the next training run.
Qualifications
- Prior evaluation, annotation, or human-rater experience on multimodal hallucination model evaluation or adjacent content for Multimodal Hallucination Model Evaluator work.
- Comfort applying multi-page rubrics consistently across long batches.
- Clear written reasoning that names the issue and the rubric clause being applied.
- Strong attention to detail and the ability to flag when a prompt itself is the problem.
- Reliable async availability for at least 10 hours per week.
Example tasks
- Compare two multimodal hallucination model evaluation model responses to the same prompt and pick the stronger one with rationale.
- Tag an unsafe response with the correct policy category and severity.
- Audit a 50-row batch for rubric consistency and report drift to the program lead.
- Propose a rubric clarification after spotting a recurring failure mode.
Nice to have
- Background in linguistics, content moderation, or trust & safety review.
- Experience with inter-rater agreement metrics and calibration cycles.
- Domain expertise that lets you spot subject-matter errors automated checks miss.
Skills
- Model output evaluation
- Rubric-based annotation
- Severity tagging
- Inter-rater calibration
- Multimodal Hallucination Model evaluation
- Multimodal evaluation
- Cross-modal reasoning
- Grounding review
- Multimodal
- Hallucination
Work model
Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.
Compensation
Hourly rate confirmed after the interview process.
Application process
Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.
- Receipt and Invoice Understanding Model Evaluator is a remote evaluation track for reviewing receipt and invoice understanding model evaluation... ...Receipt and Invoice Understanding Model evaluation Multimodal evaluation Cross-modal reasoning Grounding review Receipt...SuggestedHourly payFor contractorsRemote work10 hours per week
- ...Workflow Annotator—Product Management & Marketing to remotely review evaluation prompts and responses against the company's quality rubric.... ..., label edge cases, and provide structured feedback the modeling team can use to retrain. As a contractor, you will assess frontier...SuggestedRemote jobFor contractors
- ...Medical Document OCR Model Evaluator is a remote clinical-review track for evaluating AI outputs that touch clinical review. Reviewers... ...assessment Structured clinical writing Clinical review Multimodal evaluation Cross-modal reasoning Grounding review...SuggestedRemote jobHourly payFor contractors10 hours per week
- AI Trainer Jobs is seeking a remote contractor to evaluate people ops / recruiting prompts and responses against a evolving quality rubric. You will compare model outputs, label edge cases, and provide structured feedback for retraining. Responsibilities include evaluating...SuggestedRemote jobFor contractors
- AI Trainer Jobs is seeking an Illustration Quality Evaluator for a remote contractor role. Review illustration quality evaluation prompts and responses against a published rubric, compare paired outputs, and provide structured feedback to support retraining efforts. You...SuggestedRemote jobPart timeFor contractors
$50 per hour
Combinatorics Model Evaluator is a remote review track for evaluating AI outputs across combinatorics model research review reasoning, calculations, and research workflows. Reviewers grade derivations and assumptions, reproduce key results, and document the correct method...Hourly payFor contractorsRemote work$15 - $20 per hour
...public sources and external tools. Generate high-quality human evaluation data by identifying response strengths, areas for improvement,... ...quality, clarity, tone, and completeness of responses. Ensure model responses align with expected conversational behavior and system...Contract workSummer workRemote work- Surgical Planning Safety Evaluator is a remote evaluation track for reviewing surgical planning safety evaluation prompts and responses... ...edge cases, and write the kind of structured feedback the modeling team can use to retrain. AI data reviewers help turn surgical...Remote job
- ...Scientific Figure Understanding Model Evaluator is a remote review track for evaluating AI outputs across scientific figure understanding... ...Scientific Figure Understanding Model research review Multimodal evaluation Cross-modal reasoning Grounding review Scientific...Remote jobHourly payFor contractors10 hours per week
- AuraOne is seeking a Latin Bilingual Expert for a remote evaluation track. Reviewers compare paired outputs to a quality rubric, label edge cases, and generate structured feedback to retrain models. Responsibilities include evaluating outputs against rubrics, tagging issues...For contractorsRemote work10 hours per week
$80 per hour
.../hour Location: Remote Role Responsibilities Use frontier AI coding agents to complete and evaluate complex data engineering tasks. Review model-generated implementations involving ETL pipelines , data warehouses , analytics platforms , and distributed...Contract workSummer workRemote work- Prolific Academic Ltd is seeking Licensed Pharmacists to assist in AI model training and evaluation from a home office. Successful candidates join as Domain Experts and will be paid to train and evaluate models, with tests to assess suitability. Experts review AI-driven...Remote workHome officeFlexible hours
$65 per hour
Prolific is seeking registered nurses in Las Vegas, NV, to assist in training AI models. You will review AI responses, rate their accuracy and safety, and write feedback to enhance AI learning. Candidates must be verified registered nurses, have recent clinical experience...Hourly paySelf employmentWork from homeFlexible hours$80 - $150 per hour
Prolific Academic Ltd is recruiting Medical Doctors to help train and evaluate AI models. You’ll complete a quick test to assess suitability and, if successful, join as a Domain Expert, paid to work on AI tasks. Researchers pay $80-$150 per hour per completed task, with...Hourly paySelf employmentWork from home- ...interaction, enabling the human touch where it has been previously unscalable. We achieve this through pioneering research in multimodal AI for modeling human-to-human communication (language, audio, and video), as well as generating audio-visual avatar behavior. Our models...Full timeRemote workRelocation packageFlexible hours
- ...Refusal Preference Reward Model Evaluator is a remote red-team track for stress-testing AI systems against adversarial prompts. Reviewers craft attack scenarios, document the failure mode, and pair each successful jailbreak with the rubric clause it violated so the safety...Remote jobHourly payFor contractors10 hours per week
- ...Quant Finance Reasoning Model Evaluator is a remote review track for evaluating AI outputs across finance and risk workflows. Reviewers grade calculations, narrative reasoning, and policy adherence; flag compliance and reconciliation issues; and document the correct treatment...Remote jobHourly payFor contractorsWork experience placement10 hours per week
- ...Policy Preference Reward Model Evaluator is a remote review track for evaluating AI outputs in policy review workflows. Reviewers grade citation accuracy, statutory reasoning, and policy adherence; flag risk; and document the corrected analysis so the modeling team can...Remote jobHourly payFor contractors10 hours per week
$185k - $325k
...aviation. About You:You want to build a model that understands what happens next. Not... ...machinery that turns those models into evaluated candidate futures.Responsibilities:Design... ...dynamics, aircraft performance or air traffic.Multimodal fusion across vision, state and...Full timeRemote workRelocation$219k - $351k
...becoming a memory-bandwidth business. As models scale past what any single GPU can hold... ...sliding-window/hybrid layers, diffusion and multimodal transformers — and how each changes the... ...managers to ensure every candidate is evaluated fairly and holistically.Recruiting...Work at officeRemote workFlexible hours$85 per hour
...Location: Remote Role Responsibilities Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks. Review model-generated implementations involving cloud platforms , Kubernetes , CI/CD systems , observability...Contract workSummer workRemote work- About the role Conversational AI Evaluator is a remote evaluation track for reviewing conversational... ...the kind of structured feedback the modeling team can use to retrain. AI data... ...Conversational AI evaluation Voice, language and multimodal AI evaluation Rubric writing Expert...Hourly payFor contractorsRemote work10 hours per week
$162k - $243k
...efficient inference of large-scale foundation models.We are seeking a Staff Engineer - AI... ...for LLMs, VLMs, diffusion, and multimodal models on Qualcomm inference accelerators... ...performance/accuracy trade-off analysis, or evaluation frameworks.· PhD in a relevant field.Minimum...Work experience placementWork from home- ...reasoning and computational problem solving to improve and evaluate large language models. You will design rigorous math problems, produce clear,... ...specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents. Help enterprises move AI from proof of...Contract workFor contractorsFreelanceRemote work
$149 per hour
...Vice President - Technical AI Foundation Model EngineerRole SummaryThe VP, Technical AI... ...and semantic search capabilities.Evaluate, benchmark, and recommend foundation models... ...ManagementComplianceAuditabilityImplement controls for:Hallucination mitigationPrompt securityModel...Full timeWork at officeLocal areaRemote work1 day per week- ...Apply advanced chemistry expertise to improve large language models by creating, solving, and clearly explaining complex chemistry... ...contract role combines analytical work, English comprehension, and multimodal scientific communication using text, images, chemical equations...Contract workFor contractorsRemote work
$125k - $150k
...DescriptionEverforth ECS is seeking an AI Model Engineer to work in a hybrid remote/... ...candidate is innovative with a track record of evaluating, experimenting with, and transitioning... ..., and classification tasks.Utilize multimodal architectures for robust retrieval and data...Contract workWork at officeRemote work$50 - $300 per day
About us We are professional, agile and professional. Our work environment includes: Modern office setting Food provided Growth opportunities Responsibilities Perform in various theatrical productions, including plays, musical, and other live performances Memorize lines...Contract workInternshipWork at officeRemote workFlexible hoursShift workNight shiftDay shift$130.1k - $174.91k
...purpose.Job Description:Responsible for analytical, statistical modeling, and forecasting methods.The primary requirement is not... ...Independently determines and develops approach to solutions.Work is evaluated upon completion for adequacy in satisfying objectives....Part timeLocal areaRemote work- ...Physician Services PC · Mental HealthValhalla, NYAllied Health Prof/TechnicalPer DiemAll ShiftsAs neededJob Summary: The Mental Health Evaluator (MHE) is a critical role dedicated to delivering comprehensive, clinically appropriate, culturally competent, and trauma-informed...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Multimodal Hallucination Model Evaluator [Remote]. Be the first to apply!


