STEM Explanation AI Evaluator [Remote]
AuraOne Human Data
- Remote job
STEM Explanation AI Evaluator is a remote review track for evaluating AI outputs across stem research reasoning, calculations, and research workflows. Reviewers grade derivations and assumptions, reproduce key results, and document the correct method so the modeling team can train on it.
Why this role matters
STEM research models live or die on whether their derivations actually hold up under scrutiny. AuraOne uses scientific specialists to grade outputs the way a peer reviewer would — checking assumptions, reproducing key steps, and capturing the right method alongside the wrong one.
Responsibilities
- Review AI outputs against current stem research methods, conventions, and prior work for STEM Explanation AI Evaluator assignments.
- Reproduce or sanity-check key derivations, calculations, or experimental claims.
- Flag dimensional, methodological, and citation errors with structured severity tags.
- Capture the corrected reasoning or worked example so the modeling team can train on it.
- Adjudicate disputed answers against textbooks, papers, or community standards.
- Maintain reviewer-quality scores in inter-rater calibration cycles.
Qualifications
- Graduate-level training or equivalent applied experience in stem research or a closely related field for STEM Explanation AI Evaluator work.
- Hands-on experience publishing, teaching, or advising on the topic at a professional level.
- Comfort applying multi-page rubrics consistently across long batches.
- Clear written reasoning that cites methods, papers, or worked examples.
- Reliable async availability for at least 10 hours per week.
Example tasks
- Reproduce an stem research derivation from a model output and flag any algebraic or dimensional errors.
- Grade a model's literature summary against the cited papers and rate the citation quality.
- Adjudicate a disputed answer between two reviewers using textbook methods.
- Audit a 25-row batch for rubric consistency and report drift to the program lead.
Nice to have
- PhD, postdoc, or industry research experience in the topic area.
- Prior work reviewing AI-assisted research tooling and its failure modes.
- Multilingual fluency for non-English papers and corpora.
Skills
- Scientific reasoning
- Method validation
- Citation review
- Quantitative analysis
- STEM research
- Learning design
- Assessment review
- Pedagogy
- STEM
- Explanation
Work model
Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.
Compensation
Hourly rate confirmed after the interview process.
Application process
Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.
- About the role PhD Mathematics AI Evaluator is a remote review track for evaluating AI outputs across mathematics reasoning, calculations,... ...example so the modeling team can train on it. Role details Track STEM research review Work model Remote · Independent specialist...SuggestedHourly payFor contractorsRemote work10 hours per week
$40 - $100 per hour
About OpenTrain OpenTrain AI is the hiring and contracting organization for this opportunity... ...work About AI Training and Scientific Evaluation AI training is the human side of... ...when models handle calculations, technical explanations, uncertainty, and limits of inference....SuggestedHourly payContract workPart timeFor contractorsRemote work- A leading AI research accelerator is hiring a position focused on contributing to projects that evaluate and enhance AI systems. You will design community service scenarios, write structured explanations, and evaluate AI accuracy. The ideal candidate will have 4+ years...SuggestedRemote jobFull timeFor contractors
$80 per hour
prolificacademicltd seeks senior AI/ML engineers to join an expert network contributing... ...to large language model training and evaluation. Roles are task-based, with participants... ...auditing training code, evaluating explanations, and providing structured feedback for RLHF...SuggestedRemote jobHourly payFlexible hours- CNTXT AI is seeking a remote contractor to evaluate AI-generated financial content and develop test cases that probe analytical reasoning. You will help... ...AI models handle financial information with clear explanations and rigorous checks. Responsibilities include assessing...SuggestedRemote jobFor contractors
- YO AI Labs seeks PhD and academic experts to support AI research projects remotely. You will evaluate model outputs across technical and humanities topics, using your subject-matter... .... Strong written English and clear explanation of complex concepts are essential. #J-...Remote jobFor contractors
$14.5 per hour
A technology company is seeking an AI Web Search Evaluator to enhance the quality of search engine results. This flexible, remote, part-time role focuses on analyzing search performance and providing feedback to improve algorithms. Ideal candidates will have strong analytical...Hourly payPart timeRemote workFlexible hours- ...Weekday 1 is seeking expert Evaluators to review AI-generated real estate, hospitality and events outputs for accuracy, rigor and domain quality. You will apply deep expertise to grade documents, spreadsheets and slide decks. Requirements include 5+ years in Real estate...Hourly payWeekly payContract workWork at officeRemote workWeekday work
- ...Seeking a full-time Remote AI Research Evaluator with a PhD in Quantitative Finance to assess and enhance AI models' capabilities in financial reasoning and quantitative analysis through flexible, contract-based work. Key responsibilities Assessing the factuality and...Full timeContract workRemote workFlexible hours
- ...About the role Swahili Evaluation AI Evaluator is a remote evaluation track for reviewing swahili generalist evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback...Hourly payFor contractorsRemote work10 hours per week
$20 per hour
A tech company specializing in AI is hiring a Digital Web Designer. In this remote role, you will evaluate AI-generated designs and provide feedback to enhance the model’s understanding of aesthetics. An ideal candidate will have a strong background in UI/UX design and...Remote workFlexible hours- ...Obsidian is hiring expert Evaluators in Real estate, hospitality, and events to review AI-generated work for accuracy, rigor, and domain quality. This remote position requires deep expertise and involves grading outputs like documents and presentations. Applicants must...Work at officeRemote work
- ...Obsidian is hiring expert Evaluators in real estate, hospitality, and events to review AI-generated work products for accuracy and quality. This is a remote position requiring strong expertise in the relevant domains. Ideal candidates should have over 5 years of professional...Work at officeRemote work
- ...Alignerr is seeking Creative Writing Evaluators to assess AI-generated stories, essays, poetry, and other writings to shape how AI learns compelling prose. This fully remote, flexible contract roles welcomes avid readers and writers with no publishing credits required...Contract workRemote workFlexible hours
$14.5 per hour
...AI Web Search Evaluator Welo Data works with technology companies to provide datasets that are high-quality, ethically sourced, relevant, diverse, and scalable to supercharge their AI models. As a Welocalize brand, Welo Data leverages over 25 years of experience in...Bi-weekly payHourly payPart timeImmediate startRemote workWork from homeFlexible hours$11.5 per hour
...position as an Online Task Contributor. In this role, you will evaluate and provide feedback on content to enhance search engine results... ...11.50 hourly, based on task completion, with a supportive community of contributors involved in AI advancements. #J-18808-Ljbffr...Hourly payPart timeRemote work- ...MERIT Beauty is seeking Turkish-speaking annotators for a contract role evaluating AI-generated content. You will review content for coherence, consistency, and alignment with real-world expectations in Turkish. The ideal candidates are native or fluent Turkish speakers...Contract workTemporary workImmediate startRemote work
$14.5 per hour
...AI Web Search Evaluator As a Web Search Evaluator, you will play a key role in improving the quality of search engine results, ensuring users find the most relevant and useful information. Your work will directly impact the development of AI algorithms, making search...Hourly payPart timeCurrently hiringImmediate startRemote workWork from home10 hours per weekFlexible hours- ...About the role We are hiring expert Evaluators in Data analysis / quantitative readouts to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise...Hourly payWork at officeRemote work
- Obsidian is hiring expert Evaluators in Healthcare operations for a remote, hourly engagement. You will review AI-generated work products for accuracy, rigor, and domain quality, leveraging your extensive subject-matter expertise in the field. The ideal candidate has over...Remote jobHourly payWork at office
- Mercor is seeking experts in Spreadsheet QA and workbook maintenance to review AI-generated documents, spreadsheets, and slide decks for accuracy and quality. This is a remote, hourly engagement. Ideal candidates have 5+ years in Spreadsheet QA, fluent English, and strong...Remote jobHourly payWork at office
$65 - $70 per hour
Trust and Safety Policy AI Evaluator is a remote red-team track for stress-testing AI systems against adversarial prompts. Reviewers craft attack scenarios, document the failure mode, and pair each successful jailbreak with the rubric clause it violated so the safety team...Hourly payFor contractorsRemote work10 hours per week- Dialect Code-Switching AI Evaluator is a remote evaluation track for reviewing dialect code switching ai evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the...Hourly payFor contractorsRemote work10 hours per week
- Rex.zone is seeking a Senior AI data annotator to perform data labeling and evaluation for NLP tasks, RLHF assessments, and prompt QA to improve training data quality and model performance. This is a US-based remote, full-time role aligned with Miami talent demand. You...Remote jobFull time
- AI Trainer Jobs seeks a remote independent contractor to review AI outputs for fintech operations evaluation. You will assess workflow adherence, tone, and escalation logic, assigning severity tags and documenting next steps for model training. The role requires experience...Remote jobHourly payFor contractorsFlexible hours
- Archangel Health AI is seeking Clinical AI Evaluators to review AI-generated clinical outputs, benchmark diagnostic reasoning, and refine responses to real-world medical queries. You will perform clinical accuracy auditing, RLHF ranking, error and harm identification,...Remote workFlexible hoursShift work
- Prolific is seeking fluent Norwegian speakers to act as evaluators for AI training, performing side-by-side assessments of text and voice snippets to judge naturalness and authenticity. You will listen to audio clips and rate how naturally the AI speaks, providing detailed...Remote jobWork from homeFlexible hours
- OpenTrain AI, Inc. seeks a senior dermatology reviewer to provide expert clinical judgment on complex dermatology cases and evaluate longitudinal data for AI systems. You will document commentary according to guidelines and participate in consensus processes, with occasional...Remote jobPart timeFor contractors
- AuraOne is seeking a Chemicals Safety Risk Evaluator in a remote, US-eligible red-team track to stress-test AI systems against adversarial prompts. Reviewers craft attack scenarios, document failures, and tie each successful jailbreak to the violated policy clause so the...Remote jobFor contractors
$30 per hour
...Location: Remote Commitment: 10-40 hours/week Role Responsibilities Evaluate outputs from large language models and autonomous agent systems... ...stakeholders. Requirements Strong experience in LLM evaluation, AI output analysis, QA/testing, UX research, or similar analytical...Remote jobHourly payContract work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to STEM Explanation AI Evaluator [Remote]. Be the first to apply!

