Conversational AI Evaluator
AI Trainer Jobs
About the role Conversational AI Evaluator is a remote evaluation track for reviewing conversational ai evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain. AI data reviewers help turn conversational ai evaluation outputs into auditable labels, rationales, and regression cases for AuraOne Human Data. Review frontier model outputs. Judge benchmark failures and calibrate other evaluators. Responsibilities Evaluate conversational ai evaluation model outputs against a versioned rubric and assign severity tags for Conversational AI Evaluator assignments. Compare paired responses and pick the stronger answer with a written rationale. Label hallucinations, instruction-following failures, and unsafe content with structured tags. Capture ambiguous prompts and route them back to the program team for rubric updates. Maintain reviewer-quality scores by calibrating against gold-standard examples each week. Role details Track Evaluation & annotation Work model Remote · Independent specialist contractor Compensation Hourly rate confirmed after the interview process. Eligible from US What you should bring Prior evaluation, annotation, or human-rater experience on conversational ai evaluation or adjacent content for Conversational AI Evaluator work. Comfort applying multi-page rubrics consistently across long batches. Clear written reasoning that names the issue and the rubric clause being applied. Strong attention to detail and the ability to flag when a prompt itself is the problem. Reliable async availability for at least 10 hours per week. Role signals Example tasks Compare two conversational ai evaluation model responses to the same prompt and pick the stronger one with rationale. Tag an unsafe response with the correct policy category and severity. Audit a 50-row batch for rubric consistency and report drift to the program lead. Propose a rubric clarification after spotting a recurring failure mode. Useful experience Background in linguistics, content moderation, or trust & safety review. Experience with inter-rater agreement metrics and calibration cycles. Domain expertise that lets you spot subject-matter errors automated checks miss. Compensation and schedule Hourly rate confirmed after the interview process. Expected arrangement: contractor , with program-defined task volume and review pacing. Placement depends on current program demand and reviewer confirmation. Skills used in matching Model output evaluation Rubric-based annotation Severity tagging Inter-rater calibration Conversational AI evaluation Voice, language and multimodal AI evaluation Rubric writing Expert review #J-18808-Ljbffr AI Trainer Jobs
$20 - $160 per hour
...connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...Dorsey . Position: Generalist Annotator — Health AI Conversation Quality Evaluation Type: Contract Compensation: $20–$160/hour...SuggestedContract workSummer workRemote work- AI Trainer Jobs is seeking a remote Conversational AI Evaluator to review prompts and model outputs against a evolving rubric. You will compare paired responses, label errors, and provide structured feedback to support retraining efforts. The role emphasizes clear reasoning...SuggestedRemote jobHourly payFor contractors10 hours per week
$20 - $160 per hour
...creative and technical talent with leading AI research labs. Headquartered in San... ...DorseyPosition: Generalist Annotator — Health AI Conversation Quality EvaluationType:... ...PreferredExperience in annotation, human-feedback, evaluation, or content review.Background in...SuggestedSummer work- Innodata Inc. is seeking detail-oriented Voice Specialists to evaluate AI models via real-time voice conversations. You will compare two models, record your interactions, and provide structured evaluations across five criteria to determine the stronger overall conversational...Suggested
$65 - $70 per hour
Trust and Safety Policy AI Evaluator is a remote red-team track for stress-testing AI systems against adversarial prompts. Reviewers... ...violated. Score model defenses across single-turn and multi-turn conversations. Triage emerging attack vectors and route them to the safety...SuggestedHourly payFor contractorsRemote work10 hours per week- AI Trainer Jobs is seeking a Text-to-Speech Conversation Evaluator for a remote contractor position. You will review prompts and responses against AuraOne's rubric, label issues, and provide structured feedback to help retrain models. The role emphasizes careful reasoning...Remote jobPart timeFor contractors10 hours per weekFlexible hours
- AI Trainer Jobs in the United States seeks individuals with strong baseball knowledge to evaluate AI assistants during MLB postseason games. You will ask questions live, compare two AI products, capture conversations, and rate responses for usefulness and accuracy. Applicants...
- ...Obsidian is hiring expert Evaluators in real estate, hospitality, and events to review AI-generated work products for accuracy and quality. This is a remote position requiring strong expertise in the relevant domains. Ideal candidates should have over 5 years of professional...Work at officeRemote work
$20 - $30 per hour
A remote-focused technology firm is looking for an individual proficient in evaluating large language model outputs. The role involves assessing AI systems, reviewing workflows, and providing actionable feedback to enhance product quality. Ideal candidates will have strong...Remote jobHourly pay- Obsidian is seeking expert Evaluators for a remote, hourly role focused on reviewing AI-generated work products for quality and accuracy. The successful candidates will leverage their subject matter expertise to assess various documents, spreadsheets, and slide decks....Remote jobHourly payWork at office
- About the role JavaScript and TypeScript AI Evaluator is a remote evaluation track for reviewing javascript and typescript ai evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured...Hourly payFor contractorsRemote work10 hours per week
- Cincinnatus LLC is recruiting a senior character animator to help shape evaluation standards for an AI-driven performance transfer model. The role focuses on facial animation, performance capture, and translating complex visuals into precise data captions, with part-time...Part time
- Archangel Health AI is seeking Clinical AI Evaluators to review AI-generated clinical outputs, benchmark diagnostic reasoning, and refine responses to real-world medical queries. You will perform clinical accuracy auditing, RLHF ranking, error and harm identification,...Remote workFlexible hoursShift work
- YO AI Labs is seeking Content Writers & Editors to contribute to a customer project focused on text quality and AI evaluation. You will apply your writing and editing expertise to help train and evaluate next-generation AI systems through high-quality input. No prior AI...Remote job
- AI Trainer Jobs seeks a remote independent contractor to review AI outputs for fintech operations evaluation. You will assess workflow adherence, tone, and escalation logic, assigning severity tags and documenting next steps for model training. The role requires experience...Remote jobHourly payFor contractorsFlexible hours
- Prolific is seeking fluent Norwegian speakers to act as evaluators for AI training, performing side-by-side assessments of text and voice snippets to judge naturalness and authenticity. You will listen to audio clips and rate how naturally the AI speaks, providing detailed...Remote jobWork from homeFlexible hours
- Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards. Start date: immediate; Duration: up to...Immediate startFlexible hours
- Obsidian is seeking expert Evaluators in Humanities / arts / culture to review AI-generated work products for accuracy, rigor, and domain quality. You will apply subject-matter expertise to grade outputs and provide structured feedback. This is a remote, hourly engagement...Remote jobHourly payWork at office
- AI Trainer Jobs is seeking licensed clinicians to remotely evaluate AI outputs in internal medicine clinical reviews. You will assess differential diagnoses, dosing logic, and guideline adherence, and document corrected reasoning for model training. Ideal candidates hold...Remote jobHourly pay10 hours per week
- Feedinkoo is seeking a Graphic Designer to help train and improve AI models focused on UI/UX design, visuals, and user experiences. You will critique AI outputs, provide feedback, and guide model improvements to better reflect aesthetics and usability for designers. The...Remote work
- Turing is a leading AI company enabling the rapid deployment of advanced AI systems. We are seeking a contractor to evaluate chatbot responses across real-world small business scenarios, crafting prompts, and comparing outputs. This role emphasizes structured feedback,...Remote jobFor contractorsFreelance
- AuraOne is seeking a Chemicals Safety Risk Evaluator in a remote, US-eligible red-team track to stress-test AI systems against adversarial prompts. Reviewers craft attack scenarios, document failures, and tie each successful jailbreak to the violated policy clause so the...Remote jobFor contractors
$30 per hour
...Location: Remote Commitment: 10-40 hours/week Role Responsibilities Evaluate outputs from large language models and autonomous agent systems... ...stakeholders. Requirements Strong experience in LLM evaluation, AI output analysis, QA/testing, UX research, or similar analytical...Remote jobHourly payContract work- Mercor is seeking expert Evaluators in Clinical / biomedical / pharma to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. This is a remote, hourly engagement. You will apply deep subject-matter...Remote jobHourly pay
- Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Malayalam and English. Key responsibilities...
- AI Trainer Jobs is seeking a Grocery Shopper Expert contractor for remote work in the US. You will review AI outputs across grocery shopper operations, grade workflow correctness, policy adherence, and stakeholder fit, and document next steps for model training. Experience...Remote jobFor contractors10 hours per week
$24 per hour
Prolific is seeking Fluent Tamil speakers to serve as AI evaluators on a remote, freelance basis. You will assess how well AI models capture emotions and cultural nuance in Tamil, with pay up to $24/hr for one-hour tasks, and the option for shorter work sessions. Requirements...Remote jobFreelance$14.5 per hour
Join to apply for the AI Web Search Evaluator role at Welo Data Welo Data works with technology companies to provide datasets that are high-quality, ethically sourced, relevant, diverse, and scalable to supercharge their AI models. As a Welocalize brand, Welo Data leverages...Hourly payPart timeImmediate startRemote workWork from home10 hours per weekFlexible hours$14.5 per hour
A technology company is seeking an AI Web Search Evaluator to enhance the quality of search engine results. This flexible, remote, part-time role focuses on analyzing search performance and providing feedback to improve algorithms. Ideal candidates will have strong analytical...Remote jobHourly payPart timeFlexible hours- AuraOne is seeking a Children Safety Content AI Evaluator for a remote, contractor-based role. You will evaluate model outputs, compare responses, and label content with structured tags according to our quality rubric. Ideal candidates have prior evaluation/annotation...Remote jobFor contractors10 hours per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Conversational AI Evaluator. Be the first to apply!


