Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Generalist Writer - AI Response Evaluation and Annotation [Remote]

AuraOne Human Data

Remote
  • Remote job

Generalist Writer - AI Response Evaluation and Annotation is a remote evaluation track for reviewing generalist writer evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain.

Why this role matters

AI data reviewers help turn generalist writer evaluation outputs into auditable labels, rationales, and regression cases for AuraOne Human Data.

Responsibilities

  • Evaluate generalist writer evaluation model outputs against a versioned rubric and assign severity tags for Generalist Writer - AI Response Evaluation and Annotation assignments.
  • Compare paired responses and pick the stronger answer with a written rationale.
  • Label hallucinations, instruction-following failures, and unsafe content with structured tags.
  • Capture ambiguous prompts and route them back to the program team for rubric updates.
  • Maintain reviewer-quality scores by calibrating against gold-standard examples each week.
  • Document recurring failure modes so the modeling team can target them in the next training run.

Qualifications

  • Prior evaluation, annotation, or human-rater experience on generalist writer evaluation or adjacent content for Generalist Writer - AI Response Evaluation and Annotation work.
  • Comfort applying multi-page rubrics consistently across long batches.
  • Clear written reasoning that names the issue and the rubric clause being applied.
  • Strong attention to detail and the ability to flag when a prompt itself is the problem.
  • Reliable async availability for at least 10 hours per week.

Example tasks

  • Compare two generalist writer evaluation model responses to the same prompt and pick the stronger one with rationale.
  • Tag an unsafe response with the correct policy category and severity.
  • Audit a 50-row batch for rubric consistency and report drift to the program lead.
  • Propose a rubric clarification after spotting a recurring failure mode.

Nice to have

  • Background in linguistics, content moderation, or trust & safety review.
  • Experience with inter-rater agreement metrics and calibration cycles.
  • Domain expertise that lets you spot subject-matter errors automated checks miss.

Skills

  • Model output evaluation
  • Rubric-based annotation
  • Severity tagging
  • Inter-rater calibration
  • Generalist Writer evaluation

Work model

Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.

Compensation

Hourly rate confirmed after the interview process.

Application process

Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Generalist Writer - AI Response Evaluation and Annotation [Remote] in Remote vacancy
  • $40 - $50 per hour

     ...Apply your writing and analytical skills to evaluate written content and AI-generated responses, helping improve the accuracy, reliability, and overall quality of advanced AI systems. Key Responsibilities Review written content and AI-generated responses for accuracy... 
    Suggested
    Hourly pay
    For contractors
    Immediate start
    Remote work

    SaidGig

    Remote
    20 days ago
  • $80 per hour

     ...ethically shape the future of AI. What We Do The Mindrift...  ...design realistic and structured evaluation scenarios for LLM-based...  ...agents make decisions. Typical Responsibilities Create structured test...  ...testing, data analysis, or NLP annotation Good understanding of test... 
    Suggested
    Part time
    Freelance
    Remote work
    Flexible hours

    Mindrift

    New York, NY
    5 days ago
  •  ...enterprises who are building AI systems to power magical experiences...  ...we build. Each one of us is responsible for contributing to...  ...detail. Preference-Based Tasks: Evaluate and complete tasks, assessing...  ...and writing samples. Virtual Annotation Test - This assignment will... 
    Suggested
    Hourly pay
    16 hours
    Contract work
    Temporary work
    Part time
    For contractors
    Remote work

    Cohere

    Richmond, VA
    4 days ago
  •  ...exam-style multiple-choice questions on Norwegian law for an AI evaluation dataset. This remote role draws on your professional legal...  ...clear, rigorous assessment content in Norwegian. Key Responsibilities Draft original Norwegian-law multiple-choice questions based... 
    Suggested
    Hourly pay
    Remote work
    Flexible hours

    SaidGig

    United States
    a month ago
  • $27.78 - $33.96 per hour

     ...Create original, exam-style multiple-choice questions that evaluate understanding of Ukrainian law. Each item will be authored from...  ...worked solution demonstrating the correct reasoning. Key Responsibilities Draft original exam-style questions on Ukrainian law based... 
    Suggested
    Hourly pay
    Remote work
    Flexible hours

    SaidGig

    United States
    23 days ago
  •  ...exam-style multiple-choice questions on Finnish law for an AI evaluation dataset. This remote, flexible-hours opportunity draws on your...  ...to produce rigorous assessment content in Finnish. Key Responsibilities Draft original Finnish law questions based on your own professional... 
    Hourly pay
    Remote work
    Flexible hours

    SaidGig

    United States
    a month ago
  • $39.69 - $48.51 per hour

     ...questions on Chinese accounting, auditing, and tax for an AI evaluation dataset. This role draws on your professional expertise to...  ...produce rigorous Chinese (Simplified) finance content. Key Responsibilities Draft original exam-style questions based on your professional... 
    Hourly pay
    Remote work
    Flexible hours

    SaidGig

    United States
    a month ago
  •  ...style multiple-choice questions on Danish law to support an AI evaluation dataset. You will apply your professional legal expertise...  ...develop rigorous questions in native-level Danish. Key Responsibilities Write original Danish law multiple-choice questions based... 
    Hourly pay
    Remote work
    Flexible hours

    SaidGig

    United States
    a month ago
  • $39.69 - $48.51 per hour

     ...knowledge multiple-choice questions in Chinese (Simplified) for an AI evaluation dataset. Each item must be produced from your own expertise...  ...worked solution. Work remotely with flexible hours. Key Responsibilities Draft original exam-style multiple-choice questions in... 
    Hourly pay
    Remote work
    Flexible hours

    SaidGig

    United States
    a month ago
  •  ...script, focused on Taiwanese or Hong Kong law to populate an AI evaluation dataset. Each question must be grounded in your...  ...remote, hourly engagement with flexible scheduling. Key Responsibilities Draft original multiple-choice questions on Taiwanese law... 
    Hourly pay
    Remote work
    Flexible hours

    SaidGig

    United States
    a month ago
  • $24.26 - $29.65 per hour

     ...Arabic general-knowledge multiple-choice questions that support AI evaluation. This flexible, remote hourly role draws on your subject-...  ...expertise to produce rigorous exam-style content. Key Responsibilities Draft original general-knowledge questions in Arabic.... 
    Hourly pay
    Remote work
    Flexible hours

    SaidGig

    United States
    a month ago
  •  ...Role Overview Create original, exam-style multiple-choice questions on Swedish law for an AI evaluation dataset, drawing on your professional legal expertise. Key Responsibilities Write Swedish law questions with ten answer options each. Provide a worked... 
    Hourly pay
    Remote work
    Flexible hours

    SaidGig

    United States
    a month ago
  • $15.79 - $19.29 per hour

     ...-style multiple-choice questions on Thai law to support an AI evaluation dataset. This role draws on your professional legal expertise...  ..., well-explained assessment content in Thai. Key Responsibilities Draft original Thai law questions based on your own professional... 
    Hourly pay
    Remote work
    Flexible hours

    SaidGig

    United States
    a month ago
  •  ...Overview Create original, exam-style multiple-choice questions on Hungarian law for an AI evaluation dataset, drawing on your professional legal expertise. Key Responsibilities Draft Hungarian law questions with ten answer options each. Provide a worked solution... 
    Hourly pay
    Remote work
    Flexible hours

    SaidGig

    United States
    a month ago
  •  ...Hong Kong accounting, auditing, and tax in Chinese (Traditional) for an AI evaluation dataset. This remote role offers flexible hours and draws on your own professional expertise. Key Responsibilities Draft original exam-style questions from your own professional... 
    Remote work
    Flexible hours

    SaidGig

    United States
    a month ago
  • $38.59 - $47.17 per hour

     ...Create original, exam-style multiple-choice questions covering Danish accounting, auditing, and tax to support an AI evaluation dataset. Key Responsibilities Develop questions using your own professional expertise. Provide ten answer options and a worked solution... 
    Hourly pay
    Remote work
    Flexible hours

    SaidGig

    United States
    23 days ago
  • $24.26 - $29.65 per hour

     ...choice questions about the law of your jurisdiction for an AI evaluation dataset. The work relies on your professional legal expertise...  ...options and a worked solution for each question. Key Responsibilities Draft exam-style multiple-choice questions in Arabic, grounded... 
    Hourly pay
    Remote work
    Flexible hours

    SaidGig

    United States
    23 days ago
  • $24.26 - $29.65 per hour

     ...questions in Arabic on accounting, auditing, and tax for an AI evaluation dataset. Questions should reflect rules and practice in...  ...include a worked solution explaining the correct answer. Key Responsibilities Draft original exam-style multiple-choice questions in... 
    Hourly pay
    Contract work
    Remote work
    Flexible hours

    SaidGig

    United States
    23 days ago
  • $20 - $40 per hour

     ...expertise to help improve next-generation AI systems through rigorous, real-world...  ...the central qualification. Key Responsibilities Review, evaluate, and produce solutions for advanced...  ...actionable feedback. Create and annotate high-quality, domain-specific... 
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    a month ago
  • $80 per hour

    A leading AI innovation firm is seeking an Evaluation Scenario Designer to create structured test cases for LLM-based agents. This part-time role allows for flexible project work, focusing on defining gold-standard behaviors and analyzing agent performance. Candidates should... 
    Remote job
    Part time
    Flexible hours

    Mindrift

    Wisconsin
    3 days ago
  • $38.59 - $47.17 per hour

     ...exam-style multiple-choice questions on Swedish law for an AI evaluation dataset. Each item must reflect your professional legal expertise...  ...a worked solution that explains the correct answer. Key Responsibilities Draft exam-style multiple-choice questions covering... 
    Hourly pay
    Remote work
    Flexible hours

    SaidGig

    United States
    3 days ago
  • $39.69 - $48.51 per hour

     ...choice questions in Simplified Chinese that evaluate understanding of Chinese law. Work from...  ...legal expertise to produce items for an AI evaluation dataset. Each item must...  ...hourly role with flexible scheduling. Key Responsibilities Draft original multiple-choice... 
    Hourly pay
    Remote work
    Flexible hours

    SaidGig

    United States
    1 day ago
  • $39.69 - $48.51 per hour

     ...covering Chinese accounting, auditing, and tax, to be used in an AI evaluation dataset. Work independently, drawing on your professional...  ...to produce high-quality items and worked solutions. Key Responsibilities Draft original multiple-choice questions focused on... 
    Hourly pay
    Remote work
    Flexible hours

    SaidGig

    United States
    1 day ago
  • $48.51 - $59.29 per hour

     ...exam-style multiple-choice questions on Norwegian law for an AI evaluation dataset. You will use your professional legal experience to...  ...a worked solution explaining the correct answer. Key Responsibilities Draft original exam-style multiple-choice questions focused... 
    Hourly pay
    Remote work
    Flexible hours

    SaidGig

    United States
    1 day ago
  • $15.79 - $19.29 per hour

     ...-style multiple-choice questions on Thai law to support an AI evaluation dataset. Each item you produce will include ten answer options...  .... This role is remote with flexible scheduling. Key Responsibilities Draft exam-style questions based on your professional expertise... 
    Hourly pay
    Remote work
    Flexible hours

    SaidGig

    United States
    2 days ago
  • $38.59 - $47.17 per hour

     ...style multiple-choice questions on Finnish law to support an AI evaluation dataset. Each item must reflect your professional legal...  ...answer options, and be accompanied by a worked solution. Key Responsibilities Draft original exam-style multiple-choice questions... 
    Hourly pay
    Remote work
    Flexible hours

    SaidGig

    United States
    1 day ago
  • $38.59 - $47.17 per hour

     ...style multiple choice questions on Danish law to support an AI evaluation dataset. The work is remote and offers flexible hours,...  ...professional practice or academic background in Denmark. Key Responsibilities Draft original exam-style questions based on your... 
    Hourly pay
    Remote work
    Flexible hours

    SaidGig

    United States
    1 day ago
  • $24.26 - $29.65 per hour

     ...general-knowledge multiple-choice questions in Arabic for an AI evaluation dataset. You will draft exam-style questions from your own...  ...including ten answer options and a worked solution. Key Responsibilities Create original general-knowledge multiple-choice... 
    Remote work
    Flexible hours

    SaidGig

    United States
    3 days ago
  • $39.69 - $48.51 per hour

     ...on Taiwanese or Hong Kong law, in Chinese (Traditional), for an AI evaluation dataset. This remote opportunity offers flexible hours and draws on your professional legal expertise. Key Responsibilities Draft original exam-style questions from your own professional... 
    Remote work
    Flexible hours

    SaidGig

    United States
    4 days ago
  • $24.26 - $29.65 per hour

     ...accounting, auditing, and tax topics relevant to your jurisdiction. Your work will support an AI evaluation dataset and draw directly on your professional expertise. Key Responsibilities Write original multiple-choice questions covering accounting, auditing, and tax.... 
    Hourly pay
    Local area
    Remote work
    Flexible hours

    SaidGig

    United States
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Generalist Writer - AI Response Evaluation and Annotation [Remote]. Be the first to apply!