Generalist Writer - AI Response Evaluation and Annotation [Remote]
AuraOne Human Data
- Remote job
Generalist Writer - AI Response Evaluation and Annotation is a remote evaluation track for reviewing generalist writer evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain.
Why this role matters
AI data reviewers help turn generalist writer evaluation outputs into auditable labels, rationales, and regression cases for AuraOne Human Data.
Responsibilities
- Evaluate generalist writer evaluation model outputs against a versioned rubric and assign severity tags for Generalist Writer - AI Response Evaluation and Annotation assignments.
- Compare paired responses and pick the stronger answer with a written rationale.
- Label hallucinations, instruction-following failures, and unsafe content with structured tags.
- Capture ambiguous prompts and route them back to the program team for rubric updates.
- Maintain reviewer-quality scores by calibrating against gold-standard examples each week.
- Document recurring failure modes so the modeling team can target them in the next training run.
Qualifications
- Prior evaluation, annotation, or human-rater experience on generalist writer evaluation or adjacent content for Generalist Writer - AI Response Evaluation and Annotation work.
- Comfort applying multi-page rubrics consistently across long batches.
- Clear written reasoning that names the issue and the rubric clause being applied.
- Strong attention to detail and the ability to flag when a prompt itself is the problem.
- Reliable async availability for at least 10 hours per week.
Example tasks
- Compare two generalist writer evaluation model responses to the same prompt and pick the stronger one with rationale.
- Tag an unsafe response with the correct policy category and severity.
- Audit a 50-row batch for rubric consistency and report drift to the program lead.
- Propose a rubric clarification after spotting a recurring failure mode.
Nice to have
- Background in linguistics, content moderation, or trust & safety review.
- Experience with inter-rater agreement metrics and calibration cycles.
- Domain expertise that lets you spot subject-matter errors automated checks miss.
Skills
- Model output evaluation
- Rubric-based annotation
- Severity tagging
- Inter-rater calibration
- Generalist Writer evaluation
Work model
Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.
Compensation
Hourly rate confirmed after the interview process.
Application process
Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.
$40 - $50 per hour
...Apply your writing and analytical skills to evaluate written content and AI-generated responses, helping improve the accuracy, reliability, and overall quality of advanced AI systems. Key Responsibilities Review written content and AI-generated responses for accuracy...SuggestedHourly payFor contractorsImmediate startRemote work$80 per hour
...ethically shape the future of AI. What We Do The Mindrift... ...design realistic and structured evaluation scenarios for LLM-based... ...agents make decisions. Typical Responsibilities Create structured test... ...testing, data analysis, or NLP annotation Good understanding of test...SuggestedPart timeFreelanceRemote workFlexible hours- ...enterprises who are building AI systems to power magical experiences... ...we build. Each one of us is responsible for contributing to... ...detail. Preference-Based Tasks: Evaluate and complete tasks, assessing... ...and writing samples. Virtual Annotation Test - This assignment will...SuggestedHourly pay16 hoursContract workTemporary workPart timeFor contractorsRemote work
- ...exam-style multiple-choice questions on Norwegian law for an AI evaluation dataset. This remote role draws on your professional legal... ...clear, rigorous assessment content in Norwegian. Key Responsibilities Draft original Norwegian-law multiple-choice questions based...SuggestedHourly payRemote workFlexible hours
$27.78 - $33.96 per hour
...Create original, exam-style multiple-choice questions that evaluate understanding of Ukrainian law. Each item will be authored from... ...worked solution demonstrating the correct reasoning. Key Responsibilities Draft original exam-style questions on Ukrainian law based...SuggestedHourly payRemote workFlexible hours- ...exam-style multiple-choice questions on Finnish law for an AI evaluation dataset. This remote, flexible-hours opportunity draws on your... ...to produce rigorous assessment content in Finnish. Key Responsibilities Draft original Finnish law questions based on your own professional...Hourly payRemote workFlexible hours
$39.69 - $48.51 per hour
...questions on Chinese accounting, auditing, and tax for an AI evaluation dataset. This role draws on your professional expertise to... ...produce rigorous Chinese (Simplified) finance content. Key Responsibilities Draft original exam-style questions based on your professional...Hourly payRemote workFlexible hours- ...style multiple-choice questions on Danish law to support an AI evaluation dataset. You will apply your professional legal expertise... ...develop rigorous questions in native-level Danish. Key Responsibilities Write original Danish law multiple-choice questions based...Hourly payRemote workFlexible hours
$39.69 - $48.51 per hour
...knowledge multiple-choice questions in Chinese (Simplified) for an AI evaluation dataset. Each item must be produced from your own expertise... ...worked solution. Work remotely with flexible hours. Key Responsibilities Draft original exam-style multiple-choice questions in...Hourly payRemote workFlexible hours- ...script, focused on Taiwanese or Hong Kong law to populate an AI evaluation dataset. Each question must be grounded in your... ...remote, hourly engagement with flexible scheduling. Key Responsibilities Draft original multiple-choice questions on Taiwanese law...Hourly payRemote workFlexible hours
$24.26 - $29.65 per hour
...Arabic general-knowledge multiple-choice questions that support AI evaluation. This flexible, remote hourly role draws on your subject-... ...expertise to produce rigorous exam-style content. Key Responsibilities Draft original general-knowledge questions in Arabic....Hourly payRemote workFlexible hours- ...Role Overview Create original, exam-style multiple-choice questions on Swedish law for an AI evaluation dataset, drawing on your professional legal expertise. Key Responsibilities Write Swedish law questions with ten answer options each. Provide a worked...Hourly payRemote workFlexible hours
$15.79 - $19.29 per hour
...-style multiple-choice questions on Thai law to support an AI evaluation dataset. This role draws on your professional legal expertise... ..., well-explained assessment content in Thai. Key Responsibilities Draft original Thai law questions based on your own professional...Hourly payRemote workFlexible hours- ...Overview Create original, exam-style multiple-choice questions on Hungarian law for an AI evaluation dataset, drawing on your professional legal expertise. Key Responsibilities Draft Hungarian law questions with ten answer options each. Provide a worked solution...Hourly payRemote workFlexible hours
- ...Hong Kong accounting, auditing, and tax in Chinese (Traditional) for an AI evaluation dataset. This remote role offers flexible hours and draws on your own professional expertise. Key Responsibilities Draft original exam-style questions from your own professional...Remote workFlexible hours
$38.59 - $47.17 per hour
...Create original, exam-style multiple-choice questions covering Danish accounting, auditing, and tax to support an AI evaluation dataset. Key Responsibilities Develop questions using your own professional expertise. Provide ten answer options and a worked solution...Hourly payRemote workFlexible hours$24.26 - $29.65 per hour
...choice questions about the law of your jurisdiction for an AI evaluation dataset. The work relies on your professional legal expertise... ...options and a worked solution for each question. Key Responsibilities Draft exam-style multiple-choice questions in Arabic, grounded...Hourly payRemote workFlexible hours$24.26 - $29.65 per hour
...questions in Arabic on accounting, auditing, and tax for an AI evaluation dataset. Questions should reflect rules and practice in... ...include a worked solution explaining the correct answer. Key Responsibilities Draft original exam-style multiple-choice questions in...Hourly payContract workRemote workFlexible hours$20 - $40 per hour
...expertise to help improve next-generation AI systems through rigorous, real-world... ...the central qualification. Key Responsibilities Review, evaluate, and produce solutions for advanced... ...actionable feedback. Create and annotate high-quality, domain-specific...Hourly payContract workRemote work$80 per hour
A leading AI innovation firm is seeking an Evaluation Scenario Designer to create structured test cases for LLM-based agents. This part-time role allows for flexible project work, focusing on defining gold-standard behaviors and analyzing agent performance. Candidates should...Remote jobPart timeFlexible hours$38.59 - $47.17 per hour
...exam-style multiple-choice questions on Swedish law for an AI evaluation dataset. Each item must reflect your professional legal expertise... ...a worked solution that explains the correct answer. Key Responsibilities Draft exam-style multiple-choice questions covering...Hourly payRemote workFlexible hours$39.69 - $48.51 per hour
...choice questions in Simplified Chinese that evaluate understanding of Chinese law. Work from... ...legal expertise to produce items for an AI evaluation dataset. Each item must... ...hourly role with flexible scheduling. Key Responsibilities Draft original multiple-choice...Hourly payRemote workFlexible hours$39.69 - $48.51 per hour
...covering Chinese accounting, auditing, and tax, to be used in an AI evaluation dataset. Work independently, drawing on your professional... ...to produce high-quality items and worked solutions. Key Responsibilities Draft original multiple-choice questions focused on...Hourly payRemote workFlexible hours$48.51 - $59.29 per hour
...exam-style multiple-choice questions on Norwegian law for an AI evaluation dataset. You will use your professional legal experience to... ...a worked solution explaining the correct answer. Key Responsibilities Draft original exam-style multiple-choice questions focused...Hourly payRemote workFlexible hours$15.79 - $19.29 per hour
...-style multiple-choice questions on Thai law to support an AI evaluation dataset. Each item you produce will include ten answer options... .... This role is remote with flexible scheduling. Key Responsibilities Draft exam-style questions based on your professional expertise...Hourly payRemote workFlexible hours$38.59 - $47.17 per hour
...style multiple-choice questions on Finnish law to support an AI evaluation dataset. Each item must reflect your professional legal... ...answer options, and be accompanied by a worked solution. Key Responsibilities Draft original exam-style multiple-choice questions...Hourly payRemote workFlexible hours$38.59 - $47.17 per hour
...style multiple choice questions on Danish law to support an AI evaluation dataset. The work is remote and offers flexible hours,... ...professional practice or academic background in Denmark. Key Responsibilities Draft original exam-style questions based on your...Hourly payRemote workFlexible hours$24.26 - $29.65 per hour
...general-knowledge multiple-choice questions in Arabic for an AI evaluation dataset. You will draft exam-style questions from your own... ...including ten answer options and a worked solution. Key Responsibilities Create original general-knowledge multiple-choice...Remote workFlexible hours$39.69 - $48.51 per hour
...on Taiwanese or Hong Kong law, in Chinese (Traditional), for an AI evaluation dataset. This remote opportunity offers flexible hours and draws on your professional legal expertise. Key Responsibilities Draft original exam-style questions from your own professional...Remote workFlexible hours$24.26 - $29.65 per hour
...accounting, auditing, and tax topics relevant to your jurisdiction. Your work will support an AI evaluation dataset and draw directly on your professional expertise. Key Responsibilities Write original multiple-choice questions covering accounting, auditing, and tax....Hourly payLocal areaRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Generalist Writer - AI Response Evaluation and Annotation [Remote]. Be the first to apply!


