Misconception Detection AI Evaluator [Remote]
AuraOne Human Data
- Remote job
Misconception Detection AI Evaluator is a remote evaluation track for reviewing misconception detection ai evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain.
Why this role matters
AI data reviewers help turn misconception detection ai evaluation outputs into auditable labels, rationales, and regression cases for AuraOne Human Data.
Responsibilities
- Evaluate misconception detection ai evaluation model outputs against a versioned rubric and assign severity tags for Misconception Detection AI Evaluator assignments.
- Compare paired responses and pick the stronger answer with a written rationale.
- Label hallucinations, instruction-following failures, and unsafe content with structured tags.
- Capture ambiguous prompts and route them back to the program team for rubric updates.
- Maintain reviewer-quality scores by calibrating against gold-standard examples each week.
- Document recurring failure modes so the modeling team can target them in the next training run.
Qualifications
- Prior evaluation, annotation, or human-rater experience on misconception detection ai evaluation or adjacent content for Misconception Detection AI Evaluator work.
- Comfort applying multi-page rubrics consistently across long batches.
- Clear written reasoning that names the issue and the rubric clause being applied.
- Strong attention to detail and the ability to flag when a prompt itself is the problem.
- Reliable async availability for at least 10 hours per week.
Example tasks
- Compare two misconception detection ai evaluation model responses to the same prompt and pick the stronger one with rationale.
- Tag an unsafe response with the correct policy category and severity.
- Audit a 50-row batch for rubric consistency and report drift to the program lead.
- Propose a rubric clarification after spotting a recurring failure mode.
Nice to have
- Background in linguistics, content moderation, or trust & safety review.
- Experience with inter-rater agreement metrics and calibration cycles.
- Domain expertise that lets you spot subject-matter errors automated checks miss.
Skills
- Model output evaluation
- Rubric-based annotation
- Severity tagging
- Inter-rater calibration
- Misconception Detection AI evaluation
- Learning design
- Assessment review
- Pedagogy
- Misconception
- Detection
Work model
Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.
Compensation
Hourly rate confirmed after the interview process.
Application process
Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.
$14.5 per hour
...Join to apply for the AI Web Search Evaluator role at Welo Data Welo Data works with technology companies to provide datasets that are high-quality, ethically sourced, relevant, diverse, and scalable to supercharge their AI models. As a Welocalize brand, Welo Data...SuggestedHourly payPart timeImmediate startRemote workWork from home10 hours per weekFlexible hours$14.5 per hour
A technology company is seeking an AI Web Search Evaluator to enhance the quality of search engine results. This flexible, remote, part-time role focuses on analyzing search performance and providing feedback to improve algorithms. Ideal candidates will have strong analytical...SuggestedHourly payPart timeRemote workFlexible hours$14.5 per hour
...AI Web Search Evaluator Welo Data works with technology companies to provide datasets that are high-quality, ethically sourced, relevant, diverse, and scalable to supercharge their AI models. As a Welocalize brand, Welo Data leverages over 25 years of experience in...SuggestedBi-weekly payHourly payPart timeImmediate startRemote workWork from homeFlexible hours- Obsidian is hiring expert Evaluators in Healthcare operations for a remote, hourly engagement. You will review AI-generated work products for accuracy, rigor, and domain quality, leveraging your extensive subject-matter expertise in the field. The ideal candidate has over...SuggestedRemote jobHourly payWork at office
- Mercor is seeking expert Evaluators in Privacy/regulatory compliance to review AI-generated documents, spreadsheets, and slide decks for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. This is a remote, hourly engagement...SuggestedRemote jobHourly payWork at office
- Mercor is seeking expert Evaluators in Compliance / regulatory response with financial-services AI to review AI-generated outputs (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. This remote, hourly engagement leverages your subject-matter...Remote jobHourly pay
- Mercor is hiring expert Evaluators in General Sales / GTM to review AI-generated outputs for accuracy and quality. This remote hourly role requires applying deep domain knowledge to assess documents, spreadsheets, and slide decks. Ideal candidates have 5+ years in General...Remote jobHourly payWork at office
- Modern MedEd is seeking expert Evaluators in Clinical/biomedical/pharma to review AI-generated work products for accuracy and domain quality. This is a remote, hourly engagement that you can complete on your own schedule. You will apply deep subject-matter expertise to...Remote jobHourly payWork at office
- BAM Ventures is seeking Swedish-speaking remote annotators to evaluate AI-generated content, ensuring that it's coherent and aligns with real-world expectations. Your role will involve reviewing outputs, identifying deviations, and providing structured feedback to enhance...Remote job
- ...Population Health Informaticists to bring their domain expertise to AI development, focusing on health data, analytics, and data... ...contribute to cutting-edge AI projects in population health. You will evaluate AI-generated analyses, identify errors, review data pipelines,...Remote jobHourly payContract work10 hours per weekFlexible hours
- Mercor is seeking expert Evaluators in Clinical / biomedical / pharma to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. This is a remote, hourly engagement. You will apply deep subject-matter...Remote jobHourly pay
$30 per hour
...Location: Remote Commitment: 10-40 hours/week Role Responsibilities Evaluate outputs from large language models and autonomous agent systems... ...stakeholders. Requirements Strong experience in LLM evaluation, AI output analysis, QA/testing, UX research, or similar analytical...Remote jobHourly payContract work$20 per hour
A tech company specializing in AI is hiring a Digital Web Designer. In this remote role, you will evaluate AI-generated designs and provide feedback to enhance the model’s understanding of aesthetics. An ideal candidate will have a strong background in UI/UX design and...Remote jobFlexible hours$50 per hour
Prolific seeks Product Designers and UX Specialists to help train AI models using your expertise. You'll evaluate AI-generated designs and ensure usability and accessibility while working from home. Ideal candidates hold a BS, MS, or PhD in a relevant field and have at...Remote jobWork from homeFlexible hours$400k
Legal Contracts / Diligence / Redlines Evaluator (Train AI Models Part Time!) Remote Up to $400,000/ year Legal Counsel Attorney / Lawyer Legal Paralegal Mercor is hiring expert Evaluators in Legal contracts / diligence / redlines to review and assess AI-generated work...Hourly payPart timeWork at officeRemote work- YO AI Labs is seeking a Healthcare Expert for a remote contract to support AI training and evaluation projects. You will apply clinical knowledge to assess AI outputs, workflows, and real-world use of healthcare systems. Key tasks include evaluating EHR applications, reviewing...Remote jobContract work
- A talent marketplace is seeking soccer experts to evaluate live soccer games. The role involves assessing AI-generated commentary, scoring performance, and providing feedback. Qualified candidates will have deep expertise in soccer, strong analytical and communication...Contract work
$24 per hour
Prolific is seeking an AI Trainer with advanced Tamil fluency to evaluate AI models' understanding of the Tamil language's emotional and cultural nuances. Responsibilities include assessing audio clips, reviewing tones for cultural relevancy, and ensuring quality control...Remote jobFlexible hours- Dorado is seeking expert Evaluators in program management / implementation planning to review AI-generated documents, spreadsheets, and slide decks for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. This is a remote,...Remote jobHourly payWork at office
- MERIT Beauty is seeking Turkish-speaking annotators for a contract role evaluating AI-generated content. You will review content for coherence, consistency, and alignment with real-world expectations in Turkish. The ideal candidates are native or fluent Turkish speakers...Remote jobContract workTemporary workImmediate start
$20 - $30 per hour
A remote-focused technology firm is looking for an individual proficient in evaluating large language model outputs. The role involves assessing AI systems, reviewing workflows, and providing actionable feedback to enhance product quality. Ideal candidates will have strong...Remote jobHourly pay- YO AI Labs is seeking an Investment & Finance Expert to work remotely as a contractor. The role focuses on evaluating AI outputs related to valuation, financial modeling, markets, and investments, and crafting expert prompts plus reference answers to improve model reasoning...Remote jobFor contractors
- Obsidian is seeking expert Evaluators for a remote, hourly role focused on reviewing AI-generated work products for quality and accuracy. The successful candidates will leverage their subject matter expertise to assess various documents, spreadsheets, and slide decks....Remote jobHourly payWork at office
- Feedinkoo is seeking a Graphic Designer to help train and improve AI models focused on UI/UX design, visuals, and user experiences. You will critique AI outputs, provide feedback, and guide model improvements to better reflect aesthetics and usability for designers. The...Remote work
- Dorado is seeking an AI Language Quality Evaluator fluent in Greek and English for an ongoing, task-based project. This remote freelance role involves reviewing translated and AI-flagged content to judge accuracy, classify issues, and suggest corrected translations. You...Remote jobFor contractorsFreelanceFlexible hours
- Obsidian is seeking expert Evaluators in Humanities / arts / culture to review AI-generated work products for accuracy, rigor, and domain quality. You will apply subject-matter expertise to grade outputs and provide structured feedback. This is a remote, hourly engagement...Remote jobHourly payWork at office
- Turing is seeking detail-oriented AI Analysts based in the United States for a Google Wallet evaluation project. This role allows you to engage with advanced AI tools while contributing to the future of AI. You will evaluate model responses, review output quality, and provide...Remote jobFull timeContract work
- About OpenTrain OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and... ...States English-language work 20+ hours per week About AI Response Evaluation AI training is the human side of building artificial...Part timeFor contractorsRemote work
- A leading AI research accelerator is hiring a position focused on contributing to projects that evaluate and enhance AI systems. You will design community service scenarios, write structured explanations, and evaluate AI accuracy. The ideal candidate will have 4+ years...Remote jobFull timeFor contractors
$40 - $100 per hour
About OpenTrain OpenTrain AI is the hiring and contracting organization for this opportunity. OpenTrain is the #1 platform for finding... ...English-language work About AI Training and Scientific Evaluation AI training is the human side of building modern artificial intelligence...Hourly payContract workPart timeFor contractorsRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Misconception Detection AI Evaluator [Remote]. Be the first to apply!

