Hallucination Failure Model Evaluation Specialist [Remote]
AuraOne Human Data
- Remote job
Hallucination Failure Model Evaluation Specialist is a remote evaluation track for reviewing hallucination failure model evaluation evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain.
Why this role matters
AI data reviewers help turn hallucination failure model evaluation evaluation outputs into auditable labels, rationales, and regression cases for AuraOne Human Data.
Responsibilities
- Evaluate hallucination failure model evaluation evaluation model outputs against a versioned rubric and assign severity tags for Hallucination Failure Model Evaluation Specialist assignments.
- Compare paired responses and pick the stronger answer with a written rationale.
- Label hallucinations, instruction-following failures, and unsafe content with structured tags.
- Capture ambiguous prompts and route them back to the program team for rubric updates.
- Maintain reviewer-quality scores by calibrating against gold-standard examples each week.
- Document recurring failure modes so the modeling team can target them in the next training run.
Qualifications
- Prior evaluation, annotation, or human-rater experience on hallucination failure model evaluation evaluation or adjacent content for Hallucination Failure Model Evaluation Specialist work.
- Comfort applying multi-page rubrics consistently across long batches.
- Clear written reasoning that names the issue and the rubric clause being applied.
- Strong attention to detail and the ability to flag when a prompt itself is the problem.
- Reliable async availability for at least 10 hours per week.
Example tasks
- Compare two hallucination failure model evaluation evaluation model responses to the same prompt and pick the stronger one with rationale.
- Tag an unsafe response with the correct policy category and severity.
- Audit a 50-row batch for rubric consistency and report drift to the program lead.
- Propose a rubric clarification after spotting a recurring failure mode.
Nice to have
- Background in linguistics, content moderation, or trust & safety review.
- Experience with inter-rater agreement metrics and calibration cycles.
- Domain expertise that lets you spot subject-matter errors automated checks miss.
Skills
- Model output evaluation
- Rubric-based annotation
- Severity tagging
- Inter-rater calibration
- Hallucination Failure Model Evaluation evaluation
- Frontier evaluation
- Rubric calibration
- Failure analysis
- Hallucination
- Failure
Work model
Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.
Compensation
Hourly rate confirmed after the interview process.
Application process
Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.
$20 - $60 per hour
...AI systems. You will design realistic Fortune 500 style scenarios and interact iteratively with an advanced language model to create, edit, and evaluate Office Open XML files, with a focus on .pptx deliverables. No prior AI experience is required, your domain knowledge...SuggestedHourly payContract workFor contractorsWork at officeRemote work$20 - $60 per hour
...Help train next-generation AI systems by creating rigorous, real-world evaluations that test how well advanced models learn, reason, and perform. This remote contract opportunity is open to recent graduates, advanced-degree holders, and professionals from any background...SuggestedHourly payContract workFor contractorsRemote work- ...CI Failure Evaluation Specialist is a remote review track for evaluating AI outputs across ci failure evaluation specialist operations workflows... ...operational risk; and document the right next step so the modeling team can train on it. Why this role matters CI...SuggestedRemote jobHourly payFor contractorsWork experience placement10 hours per week
$70 - $90 per hour
...accelerator kernel development tasks that support the training and evaluation of advanced AI models. This role focuses on assessing task quality, numerical... ...errors, shape or stride mismatches, and autotuning failures. Experience with at least three kernel task types:...SuggestedHourly payRemote work- ...that happens: every product built on a model is bounded by what it costs to run, so the... ...for: ~ This role builds the evaluation and decision systems that make agentic inference... ...automated scoring, regression testing, and failure analysis In addition to hands-on LLM...SuggestedFull time
$60 - $90 per hour
...rigorous, real-world materials science and engineering tasks that evaluate an AI model’s expert reasoning. You will create prompts, supporting data... ...task. Run tasks against the model, analyze successes and failures, and refine tasks until they meaningfully discriminate...Hourly payWork at officeRemote work- ...Frontier Model Misuse Red Team Specialist is a remote red-team track for stress-testing AI systems against... ...attack scenarios, document the failure mode, and pair each successful jailbreak... ...Why this role matters Adversarial evaluation is how AuraOne hardens AI models...Remote jobHourly payFor contractors10 hours per week
- ...Job DescriptionProSidian Seeks a Infrastructure Project Evaluation and Coordination Specialist [DOE0139138] for Program Support on a Exempt W2: No... ...well together Humility - exhibits grace in success and failure while doing meaningful work where skills have impact and...Full timeContract workTemporary workFor contractorsFor subcontractorWork at officeRemote workFlexible hours
$295k
...first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world... ...City, Montreal, Seoul, Germany and Paris. Join us!Why this role?Evaluation is critical to making progress in scaling intelligence. As...Full timeWork at officeLocal areaRemote workHome office$60 - $90 per hour
...Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey . Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract Compensation: $60–$90/hour Location: Remote Commitment:...Full timeContract workSummer workRemote work- ...leading AI lab network. The role involves probing frontier models, designing robust evaluation tasks, and documenting findings for reproducibility.... ...with researchers to strengthen benchmark tasks and defenses against failure modes in AI systems. #J-18808-Ljbffr MercorFull timeRemote work
$70 - $80 per hour
...Role Overview Apply advanced drug safety expertise to help improve AI systems through rigorous, real-world evaluation of pharmacovigilance documentation and data. This remote contract role focuses on the quality, accuracy, and regulatory alignment of complex safety reports...Hourly payContract workRemote work- ...Role Overview Use your investment and finance expertise to evaluate and improve AI model performance on financial reasoning, valuation, markets, and real-world investment scenarios. Key Responsibilities Assess AI model outputs on valuation, financial modeling, markets...For contractorsRemote work
$70 - $90 per hour
...Role Overview Help evaluate Neuron Kernel Interface development tasks that support the training and evaluation of advanced AI models. You will assess kernel quality, numerical correctness, CUDA-to-NKI migration fidelity, and whether implementations are well suited to...Hourly payRemote work$100 per hour
...Role Overview Apply your finance expertise to help improve AI models across complex financial problem-solving areas, including capital... ...No prior AI experience is required. Key Responsibilities Evaluate language models in finance domains where performance needs...Hourly payContract workRemote work10 hours per weekFlexible hours- ...Operations Research Model Evaluator is a remote review track for evaluating AI outputs across... ...under scrutiny. AuraOne uses scientific specialists to grade outputs the way a peer reviewer... ...AI-assisted research tooling and its failure modes. Multilingual fluency for non-...Remote jobHourly payFor contractors10 hours per week
$60 - $70 per hour
...safety, alignment, and overall quality of frontier AI model outputs on complex, policy sensitive, and ambiguous "grey area" topics. Work through structured evaluations to identify unsafe behavior, reasoning failures, hallucinations, and policy violations, and provide clear...Hourly payRemote work- ...As the Manager of Model Validation & Verification (VnV) for Behavior Autonomy, you will lead an engineering and data science team responsible for evaluating, benchmarking, and validating the machine learning models and behavioral algorithms that drive our autonomous vehicle...Full timeTemporary workRelocation package
$20 per hour
...public sources and external tools. Generate high-quality human evaluation data by identifying response strengths, areas for improvement,... ...quality, clarity, tone, and completeness of responses. Ensure model responses align with expected conversational behavior and system...Remote jobContract workPart timeSummer work$20 per hour
SupportFinity™ is seeking an Editorial Proofreader to evaluate AI models and improve their quality through expert writing and editing skills. This role can be part‑time or full‑time, allowing for a flexible schedule and project selection. Applicants must be fluent in English...Remote jobHourly payFull timePart timeFlexible hours- SupportFinity™ is looking for an Editorial Proofreader to join our team to train AI models. In this role, you will measure AI chatbot progress, evaluate logic, and solve problems to enhance model quality. Applicants should have a strong command of English and experience...Remote jobHourly payFull timePart timeFlexible hours
- Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant...Remote jobFlexible hours
- ...is seeking Biology Experts and Life Science Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities...Remote jobHourly payFlexible hours
$20 per hour
SupportFinity™ is looking for an Editorial Proofreader to join our team for AI model training. In this remote role, you'll evaluate AI chatbots and enhance model quality. Candidates should have fluency in English and strong editing skills. This position can be full‑time...Remote jobHourly payFull timeContract workPart time$20 per hour
SupportFinity™ in Maine is seeking an Editorial Proofreader to join their team focused on training AI models. The position involves evaluating AI chatbot outputs and improving model quality through expert editing and writing skills. This flexible role allows you to work...Remote jobHourly payFlexible hours- ...Refusal Preference Reward Model Evaluator is a remote red-team track for stress-testing AI systems... ...craft attack scenarios, document the failure mode, and pair each successful... ...Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR...Remote jobHourly payFor contractors10 hours per week
$60 per hour
Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing...Remote jobHourly payWork from homeFlexible hours- ...Medical Document OCR Model Evaluator is a remote clinical-review track for evaluating AI outputs... ...clinical decision support and its failure modes. Bilingual clinical experience... ...Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR...Remote jobHourly payFor contractors10 hours per week
$50 - $100 per hour
...Apply your software engineering expertise to help train and evaluate next-generation AI systems through real-world coding tasks, technical... ...remote contract role focuses on code generation workflows and model evaluation; prior AI experience is not required. Key...Hourly payContract workFor contractorsRemote work- ...Role Title: Research Engineer - Code Generation & Model Evaluation Role Type: Contractor Location: Remote micro1 is engaging Research Engineers to participate in a project focused on code generation and model evaluation for a customer's initiative. In this role...Temporary workFor contractorsRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Hallucination Failure Model Evaluation Specialist [Remote]. Be the first to apply!
- letter of credit specialist Remote
- candidate experience specialist Remote
- gaming specialist Remote
- channel specialist Remote
- information technology specialist Remote
- leasing specialist Remote
- settlement specialist Remote
- infectious disease specialist Remote
- intake specialist Remote
- measurement specialist Remote



