Speaker Diarization Conversation Evaluator [Remote]
AuraOne Human Data
- Remote job
Speaker Diarization Conversation Evaluator is a remote evaluation track for reviewing speaker diarization conversation evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain.
Why this role matters
AI data reviewers help turn speaker diarization conversation evaluation outputs into auditable labels, rationales, and regression cases for AuraOne Human Data.
Responsibilities
- Evaluate speaker diarization conversation evaluation model outputs against a versioned rubric and assign severity tags for Speaker Diarization Conversation Evaluator assignments.
- Compare paired responses and pick the stronger answer with a written rationale.
- Label hallucinations, instruction-following failures, and unsafe content with structured tags.
- Capture ambiguous prompts and route them back to the program team for rubric updates.
- Maintain reviewer-quality scores by calibrating against gold-standard examples each week.
- Document recurring failure modes so the modeling team can target them in the next training run.
Qualifications
- Prior evaluation, annotation, or human-rater experience on speaker diarization conversation evaluation or adjacent content for Speaker Diarization Conversation Evaluator work.
- Comfort applying multi-page rubrics consistently across long batches.
- Clear written reasoning that names the issue and the rubric clause being applied.
- Strong attention to detail and the ability to flag when a prompt itself is the problem.
- Reliable async availability for at least 10 hours per week.
Example tasks
- Compare two speaker diarization conversation evaluation model responses to the same prompt and pick the stronger one with rationale.
- Tag an unsafe response with the correct policy category and severity.
- Audit a 50-row batch for rubric consistency and report drift to the program lead.
- Propose a rubric clarification after spotting a recurring failure mode.
Nice to have
- Background in linguistics, content moderation, or trust & safety review.
- Experience with inter-rater agreement metrics and calibration cycles.
- Domain expertise that lets you spot subject-matter errors automated checks miss.
Skills
- Model output evaluation
- Rubric-based annotation
- Severity tagging
- Inter-rater calibration
- Speaker Diarization Conversation evaluation
- Speech evaluation
- Voice QA
- Audio review
- Speaker
- Diarization
Work model
Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.
Compensation
Hourly rate confirmed after the interview process.
Application process
Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.
- ...We are seeking native German speakers to help evaluate and improve the next generation of AI-powered voice assistants in collaboration with... ...required application form. Participate in a 25-minute conversational interview discussing your background, experience, and interest...SuggestedContract workFor contractorsH1bImmediate startRemote workFlexible hours
- ...counseling, communication, qualitative research, HCI, conflict resolution, or related advisory disciplines to support an AI conversation-evaluation project. The role involves reviewing conversations between people and AI systems, assessing whether the AI gathered enough...SuggestedContract work
$20 per hour
...external tools. Generate high-quality human evaluation data by identifying response strengths,... ...model responses align with expected conversational behavior and system guidelines. Work... ...Qualifications Must-Have Bachelor's degree Native speaker in Urdu Significant experience using...SuggestedRemote jobContract workPart timeSummer work- ...the Admissions Manager and admission staff, as well as nursing and other internal and external staff to facilitate the referral conversion. Qualifications Education and Training: Degree in a health-related field from an accredited college or university preferred...SuggestedFull timeLocal area
- ...Beauty is seeking Turkish-speaking annotators for a contract role evaluating AI-generated content. You will review content for coherence,... ...Turkish. The ideal candidates are native or fluent Turkish speakers with strong cultural insight, attention to detail, and the...SuggestedContract workTemporary workImmediate startRemote work
- ...BAM Ventures is seeking Swedish-speaking remote annotators to evaluate AI-generated content, ensuring that it's coherent and aligns with... ...to enhance performance. Candidates should be native Swedish speakers with a keen attention to detail and a strong grasp of cultural...Remote work
- ...Tutor Conversation AI Evaluator is a remote review track for evaluating AI outputs across tutor conversation ai specialist operations workflows. Reviewers grade workflow correctness, policy adherence, and stakeholder fit; flag operational risk; and document the right next...Remote jobHourly payFor contractorsWork experience placement10 hours per week
$10 per hour
Indonesian - AI Product Evaluator The Project Productive Playhouse is building a talent pool of Indonesian speakers for an upcoming project testing and evaluating leading AI... ...chatbots or language models through structured conversations, using assigned goals and prompts. Voice...Hourly payFor contractorsFreelanceRemote workWorldwideFlexible hours- A virtual AI evaluation firm is seeking individuals to review and evaluate AI-generated responses in therapeutic conversations. The ideal candidate will possess strong written communication and analytical skills, as well as a keen attention to detail for assessing tone...Remote jobImmediate start
- Crossing Hurdles is hiring a remote AI evaluation contractor to perform Vietnamese voice assessments and rate AI responses for quality and accuracy. You will participate in short voice conversations with AI models, follow scenarios, and provide objective feedback to help...Remote jobFor contractorsFlexible hours
$45 - $55 per hour
This is a non-engineering content-policy evaluation role. Applicants must demonstrate relevant depth in violent fiction or media, military... ...a model. You will read user requests, model responses, and conversation history, then decide which policy category applies and...Remote jobMonday to FridayShift work$20 - $26 per hour
Prolific is seeking fluent Kannada speakers to act as evaluators. You will assess how naturally and authentically AI captures Kannada speech, by listening to audio clips and comparing text and voice. This fast-paced project pays $20-26 per hour and may require about one...Remote jobHourly payWork from homeFlexible hours- ...re looking for contract, remote Turkish-speaking annotators to evaluate AI-generated content with a sharp eye for detail and cultural nuance... ...over time What We’re Looking For Native or fluent Turkish speaker Strong understanding of Turkish language and cultural context Ability...Contract workTemporary workFreelanceImmediate startRemote work
- Productive Playhouse is building a talent pool of Armenian speakers to test and evaluate leading AI chatbots. This freelance, project-based opportunity lets you choose tasks, set your own hours, and work with other clients as needed. Open to freelancers outside the U.S....Remote jobFreelanceFlexible hours
- Handshake is seeking an AI Policy Specialist on the Violence & Fiction team to evaluate user requests and model responses within a full conversation context, distinguishing violence in fiction from real-world uplift and ensuring precise policy application. You will read...Remote job
- ...’re looking for contract, remote French-speaking annotators to evaluate AI-generated content with a sharp eye for detail and cultural nuance... ...over time What We’re Looking For Native or fluent French speaker Strong understanding of language and cultural context Ability...Contract workTemporary workFreelanceImmediate startRemote work
$20 - $26 per hour
Prolific seeks fluent Kannada speakers to act as evaluators for AI language data. You will assess text and voice segments, rate naturalness, and help identify cultural nuances. A PayPal account is required to receive payments for tasks priced at $20-$26 per hour. Responsibilities...Remote jobHourly payWork from homeFlexible hours$20 - $26 per hour
Prolific is looking for an AI Trainer who is a fluent Lithuanian speaker to evaluate audio models used in AI development. The role involves reviewing audio snippets, providing feedback on voice quality, and ensuring cultural relevance in pronunciation. Candidates must possess...Remote jobHourly payWork from homeFlexible hours- ...data and language projects, the hourly contractor AI Trainer and Evaluator will work remotely to generate content, annotate data, and... ...and cultural appropriateness Required qualifications Native speakers of American English born and raised in the United States Excellent...Hourly payFor contractorsRemote work
$30 per hour
...Prolific is seeking Advanced Dutch Speakers in Chicago, IL to train AI models. You will complete AI tasks and assess AI performance... ...with competitive rates and direct payment through PayPal. Pass the evaluation, and you can start within 15 minutes. #J-18808-Ljbffr...Remote workWork from home- DataAnnotation is seeking analytical, detail-oriented contractors to develop prompts, test AI chatbots, and evaluate outputs. You will craft diverse conversations, write high-quality responses, and assess model performance for accuracy and style. This independent role...Remote jobFor contractorsSelf employmentFlexible hours
$20 - $26 per hour
Prolific is seeking fluent Kannada speakers to act as evaluators for AI language models. You will compare text and audio clips to assess naturalness and authenticity, rating pronunciation and tone. The role is remote and pays $20-$26 per hour for roughly one hour blocks...Remote jobHourly pay- Work from Home | Internet Analyst | Social Media Evaluator At Appen, we work with 8 out of the top 10 global technology companies in the... ...big? Appen constantly seeks language professionals and speakers across the globe for different language-related projects. If you...Part timeWork from homeWorldwideFlexible hours
- ...tables, and other content to support AI training datasets. LLM evaluation: reviewing AI-generated responses for accuracy, reasoning... ...accurate across outputs. This position exclusively seeks native speakers of American English born and raised in the United States....Hourly payFor contractorsRemote workFlexible hours
- **Job Title: AI Trainer || Image Quality Evaluator || English** **Location**: Remote | Work from Home **Employment Type:** Project-based... ...accuracy and consistency. This role is ideal for English speakers (Current resident of South Korea) who are detail-oriented and...Contract workRemote workWork from homeMonday to FridayDay shift
- ...changes, and cost-to-cure items. Conduct field inspections and evaluate property characteristics, improvements, access, utilities,... ...skills. Ability to manage difficult or sensitive conversations calmly and respectfully. Ability to collaborate with multidisciplinary...Immediate startRemote workRelocation
- Title: Forensic EvaluatorState Role Title: Psych III/Psychology Assoc IIIHiring Range: $123,326.00 - $161,762.00Pay Band: 6Agency: Dept Behavioral Health/DevelopLocation: Va Center for Behavioral RehabAgency Website: Type: General Public - GJob DutiesThis position has ...Work experience placementWork at officeLocal areaRemote work
$75 per hour
...Medical Coder AI Content Evaluator (Remote) This role focuses on applying medical coding and healthcare operations expertise to assess... ..., software installations, and routine file-sharing and file-conversion processes. Ability to work independently in a remote setting...Remote jobContract workTemporary workWork at officeFlexible hours- Centraprise is looking for detail-oriented and motivated AI Data Annotators proficient in Canadian French to enhance AI-powered conversational systems. This role involves reviewing AI-generated content for accuracy and quality, with comprehensive training provided for...Remote jobFlexible hours
- ...A global data evaluation company is looking for a Lyric Translation Reviewer to evaluate machine-translated song lyrics. The role requires reviewing for accuracy and fluency in both English and Spanish, focusing on meaning and cultural nuances. Ideal candidates are bilingual...FreelanceRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Speaker Diarization Conversation Evaluator [Remote]. Be the first to apply!




