Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Factuality Evaluation Model Evaluation Specialist [Remote]

AuraOne Human Data

Remote
  • Remote job

Factuality Evaluation Model Evaluation Specialist is a remote evaluation track for reviewing factuality evaluation model evaluation evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain.

Why this role matters

AI data reviewers help turn factuality evaluation model evaluation evaluation outputs into auditable labels, rationales, and regression cases for AuraOne Human Data.

Responsibilities

  • Evaluate factuality evaluation model evaluation evaluation model outputs against a versioned rubric and assign severity tags for Factuality Evaluation Model Evaluation Specialist assignments.
  • Compare paired responses and pick the stronger answer with a written rationale.
  • Label hallucinations, instruction-following failures, and unsafe content with structured tags.
  • Capture ambiguous prompts and route them back to the program team for rubric updates.
  • Maintain reviewer-quality scores by calibrating against gold-standard examples each week.
  • Document recurring failure modes so the modeling team can target them in the next training run.

Qualifications

  • Prior evaluation, annotation, or human-rater experience on factuality evaluation model evaluation evaluation or adjacent content for Factuality Evaluation Model Evaluation Specialist work.
  • Comfort applying multi-page rubrics consistently across long batches.
  • Clear written reasoning that names the issue and the rubric clause being applied.
  • Strong attention to detail and the ability to flag when a prompt itself is the problem.
  • Reliable async availability for at least 10 hours per week.

Example tasks

  • Compare two factuality evaluation model evaluation evaluation model responses to the same prompt and pick the stronger one with rationale.
  • Tag an unsafe response with the correct policy category and severity.
  • Audit a 50-row batch for rubric consistency and report drift to the program lead.
  • Propose a rubric clarification after spotting a recurring failure mode.

Nice to have

  • Background in linguistics, content moderation, or trust & safety review.
  • Experience with inter-rater agreement metrics and calibration cycles.
  • Domain expertise that lets you spot subject-matter errors automated checks miss.

Skills

  • Model output evaluation
  • Rubric-based annotation
  • Severity tagging
  • Inter-rater calibration
  • Factuality Evaluation Model Evaluation evaluation
  • Frontier evaluation
  • Rubric calibration
  • Failure analysis
  • Factuality
  • Evaluation

Work model

Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.

Compensation

Hourly rate confirmed after the interview process.

Application process

Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.

Vacancy posted 15 days ago
Similar jobs that could be interesting for youBased on the Factuality Evaluation Model Evaluation Specialist [Remote] in Remote vacancy
  • Dorado is seeking a Physics Specialist to contribute deep scientific expertise to AI model evaluation. You will craft and assess challenging physics problems, probe model reasoning at the frontier, and help identify where models fail under rigorous scientific scrutiny.... 
    Suggested
    Remote job

    Dorado

    New York, NY
    2 days ago
  •  ...Causal Reasoning Model Evaluation Specialist is a remote review track for evaluating AI outputs across causal reasoning model evaluation research review reasoning, calculations, and research workflows. Reviewers grade derivations and assumptions, reproduce key results,... 
    Suggested
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    24 days ago
  • $40 - $50 per hour

     ...Role Overview Apply your linguistics expertise to evaluate large language model outputs and help train next-generation AI systems. You will analyze model-human conversations for linguistic precision, assess adherence to instructions, and deliver clear, actionable feedback... 
    Suggested
    Remote job
    Hourly pay
    For contractors
    Immediate start

    SaidGig

    Remote
    4 days ago
  • $15 - $20 per hour

     ...external tools. ~Generate high-quality human evaluation data by identifying response strengths, areas for improvement, and factual inaccuracies. ~Assess reasoning quality,...  ..., and completeness of responses. ~Ensure model responses align with expected conversational... 
    Suggested
    Part time
    Summer work

    Mercor

    Remote
    a month ago
  • $15 - $20 per hour

     ...external tools . Generate high-quality human evaluation data by identifying response strengths, areas for improvement, and factual inaccuracies. Assess reasoning quality,...  ..., and completeness of responses. Ensure model responses align with expected conversational... 
    Suggested
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    21 days ago
  •  ...first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world...  ...City, Montreal, Seoul, Germany and Paris. Join us!Why this role?Evaluation is critical to making progress in scaling intelligence. As... 
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    2 days ago
  • $60 - $90 per hour

     ...Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey . Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract Compensation: $60–$90/hour Location: Remote Commitment:... 
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    Remote
    16 hours ago
  •  ...is seeking Biology Experts and Life Science Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities... 
    Hourly pay
    Remote work
    Flexible hours

    Prolific

    Jacksonville, FL
    3 days ago
  • $120 per hour

     ...engagement you will apply that operational experience to evaluate outputs from AI models, assess field-specific content, and provide clear, structured...  ...traffic control scenarios and terminology. Identify factual errors, safety risks, ambiguous language, and operational... 
    Hourly pay
    Temporary work
    Part time
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  •  ...Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant... 
    Remote work
    Flexible hours

    Prolific

    Charlotte, NC
    4 days ago
  • $89 - $95 per hour

    Role Description ORAU is seeking a fully remote Senior Advisor – Payment Model Evaluation to support the Centers for Medicare and Medicaid Services Innovation Center (CMMI) as an ORAU employee. This is a part-time, temporary role expected to last 8 months or longer.... 
    Hourly pay
    Temporary work
    Part time
    Remote work

    ORAU

    Remote
    1 day ago
  • $40 per hour

    A technology company in Massachusetts is seeking an R&D Biologist to join their team to train AI models by evaluating chatbot outputs against complex biology questions. Ideal candidates will hold advanced qualifications in biology or biochemistry. This position allows... 
    Hourly pay
    Full time
    Part time
    Remote work

    DataAnnotation

    Oklahoma City, OK
    4 days ago
  • $14 - $42 per hour

     ...Evaluate AI-generated music and lyrics across a wide range of genres, applying your knowledge of the Hindi music scene and detailed quality standards. Key Responsibilities Assess AI-generated music and rate it against established quality criteria. Compare AI-generated... 
    Hourly pay
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    8 days ago
  • $50 - $70 per hour

     ...Role Overview Help improve frontier AI models by evaluating the quality of real-world professional materials and AI-generated work. You will apply careful judgment across documents, presentations, spreadsheets, and other written content, providing feedback that helps... 
    Hourly pay
    Remote work

    SaidGig

    Canada
    5 days ago
  • $150 per hour

     ...Overview Aerospace engineering professionals apply their domain expertise to evaluate AI-generated outputs, assess technical content, and provide clear, structured feedback that improves models'' understanding of aerospace tasks, terminology, and practices. You will... 
    Hourly pay
    Temporary work
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  •  ...Role Overview Work with a leading AI lab to evaluate outputs from generative music models in German and English. This role focuses on listening, scoring, and annotating AI-generated music and lyrics across genres, using music production and audio engineering vocabulary... 
    Hourly pay
    Part time
    Immediate start
    Remote work
    10 hours per week

    SaidGig

    United States
    a month ago
  • $100 per hour

     ...-matter expertise to improve the performance of large language models on finance tasks. You will work with AI researchers to identify...  ...accelerator focused on advanced AI systems. Key Responsibilities Evaluate LLM performance in finance areas where models struggle,... 
    Hourly pay
    Contract work
    For contractors
    Freelance
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $35 - $62 per hour

     ...Evaluate AI-generated music and lyrics across a wide range of genres, applying Korean music expertise and detailed quality standards in both Korean and English. Key Responsibilities Compare AI-generated lyrics with published songs to identify similarities. Rate... 
    Hourly pay
    For contractors
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    8 days ago
  • $65 - $90 per hour

     ...Role Overview Apply your real-world architecture expertise to evaluate and improve how AI systems understand and reason about architecture. In this flexible, part-time, remote role you will review content for technical accuracy, answer domain-specific questions, and... 
    Hourly pay
    Part time
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    Remote
    19 days ago
  • $17 - $54 per hour

     ...Role Overview Evaluate AI-generated music and lyrics across a wide range of genres, applying detailed quality standards in both French and English. Key Responsibilities Compare AI-generated lyrics with published songs to identify similarities. Rate lyrics for... 
    Hourly pay
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    8 days ago
  •  ...Role Overview Apply advanced mathematical reasoning and computational problem solving to improve and evaluate large language models. You will design rigorous math problems, produce clear, logically structured solutions, verify numerical results with code, and review... 
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    24 days ago
  • $75 per hour

     ...Join a project focused on evaluating AI models in the architecture domain, specifically in visual document understanding and instruction-following. This role involves authoring complex, grounded tasks that include a clear ground-truth output and an objective rubric.... 
    Hourly pay
    Remote work

    SaidGig

    Canada
    3 days ago
  • $65 - $90 per hour

     ...Role Overview Provide expert chemical engineering knowledge to evaluate and improve AI systems, ensuring domain accuracy and practical usefulness. You will review AI-generated content, answer technical questions, and share real-world practices, tools, and standards used... 
    Hourly pay
    Part time
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    Remote
    19 days ago
  • $75 per hour

     ...Records Managers apply professional records, archives, and library expertise to evaluate AI-generated outputs, create field-accurate prompts, and provide structured feedback that improves model performance on records-management tasks. Candidates can include Archivists,... 
    Hourly pay
    Temporary work
    Part time
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $75 per hour

     ...Role Overview Records Managers apply archival, library, and information management expertise to evaluate AI model outputs related to records, collections, and information services. You will use your professional judgment to assess model-generated content, create field... 
    Hourly pay
    Temporary work
    Part time
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $75 per hour

     ...apply geospatial imaging, survey, and photogrammetric expertise to evaluate AI-generated maps and geospatial content, verify spatial accuracy, and provide clear, structured feedback that improves model outputs. No prior AI experience is required. Key Responsibilities... 
    Hourly pay
    Temporary work
    Part time
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $60 per hour

     ...and Life Science Professionals to join their Expert Network to evaluate AI-generated science. This role allows you to work from home with...  ...a competitive pay rate of up to $60 per hour for reviewing model responses, validating technical claims, and critiquing experimental... 
    Remote job
    Hourly pay
    Work from home
    Flexible hours

    Prolific

    Dallas, TX
    2 days ago
  • $60 per hour

    Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing... 
    Remote job
    Hourly pay
    Work from home
    Flexible hours

    Prolific

    San Jose, CA
    3 days ago
  • Overview Lynker Corporation is seeking a Sea Ice Model Evaluation and Applications Scientist to support the Ocean and Tsunami Center (OTC) and the U.S. National Ice Center (USNIC), to support operational and research activities involving numerical sea ice forecast guidance... 
    Temporary work
    Seasonal work
    Local area
    Remote work
    Flexible hours

    Lynker Corporation

    Suitland, MD
    4 days ago
  • Mercor is seeking a Generalist who can operate in English and Punjabi. This contract, remote position focuses on evaluating AI outputs and supporting model evaluation tasks. You will conduct fact-checking, assess reasoning, clarity, tone and completeness, and provide actionable... 
    Remote job
    Contract work

    United States Digital Space LLC

    New York, NY
    23 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Factuality Evaluation Model Evaluation Specialist [Remote]. Be the first to apply!