Factuality Evaluation Model Evaluation Specialist [Remote]
AuraOne Human Data
- Remote job
Factuality Evaluation Model Evaluation Specialist is a remote evaluation track for reviewing factuality evaluation model evaluation evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain.
Why this role matters
AI data reviewers help turn factuality evaluation model evaluation evaluation outputs into auditable labels, rationales, and regression cases for AuraOne Human Data.
Responsibilities
- Evaluate factuality evaluation model evaluation evaluation model outputs against a versioned rubric and assign severity tags for Factuality Evaluation Model Evaluation Specialist assignments.
- Compare paired responses and pick the stronger answer with a written rationale.
- Label hallucinations, instruction-following failures, and unsafe content with structured tags.
- Capture ambiguous prompts and route them back to the program team for rubric updates.
- Maintain reviewer-quality scores by calibrating against gold-standard examples each week.
- Document recurring failure modes so the modeling team can target them in the next training run.
Qualifications
- Prior evaluation, annotation, or human-rater experience on factuality evaluation model evaluation evaluation or adjacent content for Factuality Evaluation Model Evaluation Specialist work.
- Comfort applying multi-page rubrics consistently across long batches.
- Clear written reasoning that names the issue and the rubric clause being applied.
- Strong attention to detail and the ability to flag when a prompt itself is the problem.
- Reliable async availability for at least 10 hours per week.
Example tasks
- Compare two factuality evaluation model evaluation evaluation model responses to the same prompt and pick the stronger one with rationale.
- Tag an unsafe response with the correct policy category and severity.
- Audit a 50-row batch for rubric consistency and report drift to the program lead.
- Propose a rubric clarification after spotting a recurring failure mode.
Nice to have
- Background in linguistics, content moderation, or trust & safety review.
- Experience with inter-rater agreement metrics and calibration cycles.
- Domain expertise that lets you spot subject-matter errors automated checks miss.
Skills
- Model output evaluation
- Rubric-based annotation
- Severity tagging
- Inter-rater calibration
- Factuality Evaluation Model Evaluation evaluation
- Frontier evaluation
- Rubric calibration
- Failure analysis
- Factuality
- Evaluation
Work model
Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.
Compensation
Hourly rate confirmed after the interview process.
Application process
Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.
- Dorado is seeking a Physics Specialist to contribute deep scientific expertise to AI model evaluation. You will craft and assess challenging physics problems, probe model reasoning at the frontier, and help identify where models fail under rigorous scientific scrutiny....SuggestedRemote job
- ...Causal Reasoning Model Evaluation Specialist is a remote review track for evaluating AI outputs across causal reasoning model evaluation research review reasoning, calculations, and research workflows. Reviewers grade derivations and assumptions, reproduce key results,...SuggestedRemote jobHourly payFor contractors10 hours per week
$40 - $50 per hour
...Role Overview Apply your linguistics expertise to evaluate large language model outputs and help train next-generation AI systems. You will analyze model-human conversations for linguistic precision, assess adherence to instructions, and deliver clear, actionable feedback...SuggestedRemote jobHourly payFor contractorsImmediate start$15 - $20 per hour
...external tools. ~Generate high-quality human evaluation data by identifying response strengths, areas for improvement, and factual inaccuracies. ~Assess reasoning quality,... ..., and completeness of responses. ~Ensure model responses align with expected conversational...SuggestedPart timeSummer work$15 - $20 per hour
...external tools . Generate high-quality human evaluation data by identifying response strengths, areas for improvement, and factual inaccuracies. Assess reasoning quality,... ..., and completeness of responses. Ensure model responses align with expected conversational...SuggestedContract workSummer workRemote work- ...first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world... ...City, Montreal, Seoul, Germany and Paris. Join us!Why this role?Evaluation is critical to making progress in scaling intelligence. As...Full timeWork at officeLocal areaRemote workHome office
$60 - $90 per hour
...Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey . Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract Compensation: $60–$90/hour Location: Remote Commitment:...Full timeContract workSummer workRemote work- ...is seeking Biology Experts and Life Science Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities...Hourly payRemote workFlexible hours
$120 per hour
...engagement you will apply that operational experience to evaluate outputs from AI models, assess field-specific content, and provide clear, structured... ...traffic control scenarios and terminology. Identify factual errors, safety risks, ambiguous language, and operational...Hourly payTemporary workPart timeImmediate startRemote workFlexible hours- ...Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant...Remote workFlexible hours
$89 - $95 per hour
Role Description ORAU is seeking a fully remote Senior Advisor – Payment Model Evaluation to support the Centers for Medicare and Medicaid Services Innovation Center (CMMI) as an ORAU employee. This is a part-time, temporary role expected to last 8 months or longer....Hourly payTemporary workPart timeRemote work$40 per hour
A technology company in Massachusetts is seeking an R&D Biologist to join their team to train AI models by evaluating chatbot outputs against complex biology questions. Ideal candidates will hold advanced qualifications in biology or biochemistry. This position allows...Hourly payFull timePart timeRemote work$14 - $42 per hour
...Evaluate AI-generated music and lyrics across a wide range of genres, applying your knowledge of the Hindi music scene and detailed quality standards. Key Responsibilities Assess AI-generated music and rate it against established quality criteria. Compare AI-generated...Hourly payImmediate startRemote workFlexible hours$50 - $70 per hour
...Role Overview Help improve frontier AI models by evaluating the quality of real-world professional materials and AI-generated work. You will apply careful judgment across documents, presentations, spreadsheets, and other written content, providing feedback that helps...Hourly payRemote work$150 per hour
...Overview Aerospace engineering professionals apply their domain expertise to evaluate AI-generated outputs, assess technical content, and provide clear, structured feedback that improves models'' understanding of aerospace tasks, terminology, and practices. You will...Hourly payTemporary workPart timeRemote workFlexible hours- ...Role Overview Work with a leading AI lab to evaluate outputs from generative music models in German and English. This role focuses on listening, scoring, and annotating AI-generated music and lyrics across genres, using music production and audio engineering vocabulary...Hourly payPart timeImmediate startRemote work10 hours per week
$100 per hour
...-matter expertise to improve the performance of large language models on finance tasks. You will work with AI researchers to identify... ...accelerator focused on advanced AI systems. Key Responsibilities Evaluate LLM performance in finance areas where models struggle,...Hourly payContract workFor contractorsFreelanceRemote work10 hours per weekFlexible hours$35 - $62 per hour
...Evaluate AI-generated music and lyrics across a wide range of genres, applying Korean music expertise and detailed quality standards in both Korean and English. Key Responsibilities Compare AI-generated lyrics with published songs to identify similarities. Rate...Hourly payFor contractorsImmediate startRemote workFlexible hours$65 - $90 per hour
...Role Overview Apply your real-world architecture expertise to evaluate and improve how AI systems understand and reason about architecture. In this flexible, part-time, remote role you will review content for technical accuracy, answer domain-specific questions, and...Hourly payPart timeRemote work10 hours per weekFlexible hours$17 - $54 per hour
...Role Overview Evaluate AI-generated music and lyrics across a wide range of genres, applying detailed quality standards in both French and English. Key Responsibilities Compare AI-generated lyrics with published songs to identify similarities. Rate lyrics for...Hourly payImmediate startRemote workFlexible hours- ...Role Overview Apply advanced mathematical reasoning and computational problem solving to improve and evaluate large language models. You will design rigorous math problems, produce clear, logically structured solutions, verify numerical results with code, and review...Contract workFor contractorsFreelanceRemote work
$75 per hour
...Join a project focused on evaluating AI models in the architecture domain, specifically in visual document understanding and instruction-following. This role involves authoring complex, grounded tasks that include a clear ground-truth output and an objective rubric....Hourly payRemote work$65 - $90 per hour
...Role Overview Provide expert chemical engineering knowledge to evaluate and improve AI systems, ensuring domain accuracy and practical usefulness. You will review AI-generated content, answer technical questions, and share real-world practices, tools, and standards used...Hourly payPart timeRemote work10 hours per weekFlexible hours$75 per hour
...Records Managers apply professional records, archives, and library expertise to evaluate AI-generated outputs, create field-accurate prompts, and provide structured feedback that improves model performance on records-management tasks. Candidates can include Archivists,...Hourly payTemporary workPart timeImmediate startRemote workFlexible hours$75 per hour
...Role Overview Records Managers apply archival, library, and information management expertise to evaluate AI model outputs related to records, collections, and information services. You will use your professional judgment to assess model-generated content, create field...Hourly payTemporary workPart timeImmediate startRemote workFlexible hours$75 per hour
...apply geospatial imaging, survey, and photogrammetric expertise to evaluate AI-generated maps and geospatial content, verify spatial accuracy, and provide clear, structured feedback that improves model outputs. No prior AI experience is required. Key Responsibilities...Hourly payTemporary workPart timeRemote work$60 per hour
...and Life Science Professionals to join their Expert Network to evaluate AI-generated science. This role allows you to work from home with... ...a competitive pay rate of up to $60 per hour for reviewing model responses, validating technical claims, and critiquing experimental...Remote jobHourly payWork from homeFlexible hours$60 per hour
Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing...Remote jobHourly payWork from homeFlexible hours- Overview Lynker Corporation is seeking a Sea Ice Model Evaluation and Applications Scientist to support the Ocean and Tsunami Center (OTC) and the U.S. National Ice Center (USNIC), to support operational and research activities involving numerical sea ice forecast guidance...Temporary workSeasonal workLocal areaRemote workFlexible hours
- Mercor is seeking a Generalist who can operate in English and Punjabi. This contract, remote position focuses on evaluating AI outputs and supporting model evaluation tasks. You will conduct fact-checking, assess reasoning, clarity, tone and completeness, and provide actionable...Remote jobContract work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Factuality Evaluation Model Evaluation Specialist [Remote]. Be the first to apply!
- junk removal specialist Remote
- correspondence specialist Remote
- measurement specialist Remote
- disclosure specialist Remote
- partnership specialist Remote
- absence management specialist Remote
- continuous improvement specialist Remote
- loss prevention specialist Remote
- hospitality specialist Remote
- learning management system specialist Remote







