STEM Explanation AI Evaluator [Remote]
AuraOne Human Data
- Remote job
STEM Explanation AI Evaluator is a remote review track for evaluating AI outputs across stem research reasoning, calculations, and research workflows. Reviewers grade derivations and assumptions, reproduce key results, and document the correct method so the modeling team can train on it.
Why this role matters
STEM research models live or die on whether their derivations actually hold up under scrutiny. AuraOne uses scientific specialists to grade outputs the way a peer reviewer would — checking assumptions, reproducing key steps, and capturing the right method alongside the wrong one.
Responsibilities
- Review AI outputs against current stem research methods, conventions, and prior work for STEM Explanation AI Evaluator assignments.
- Reproduce or sanity-check key derivations, calculations, or experimental claims.
- Flag dimensional, methodological, and citation errors with structured severity tags.
- Capture the corrected reasoning or worked example so the modeling team can train on it.
- Adjudicate disputed answers against textbooks, papers, or community standards.
- Maintain reviewer-quality scores in inter-rater calibration cycles.
Qualifications
- Graduate-level training or equivalent applied experience in stem research or a closely related field for STEM Explanation AI Evaluator work.
- Hands-on experience publishing, teaching, or advising on the topic at a professional level.
- Comfort applying multi-page rubrics consistently across long batches.
- Clear written reasoning that cites methods, papers, or worked examples.
- Reliable async availability for at least 10 hours per week.
Example tasks
- Reproduce a stem research derivation from a model output and flag any algebraic or dimensional errors.
- Grade a model's literature summary against the cited papers and rate the citation quality.
- Adjudicate a disputed answer between two reviewers using textbook methods.
- Audit a 25-row batch for rubric consistency and report drift to the program lead.
Nice to have
- PhD, postdoc, or industry research experience in the topic area.
- Prior work reviewing AI-assisted research tooling and its failure modes.
- Multilingual fluency for non-English papers and corpora.
Skills
- Scientific reasoning
- Method validation
- Citation review
- Quantitative analysis
- STEM research
- Learning design
- Assessment review
- Pedagogy
- STEM
- Explanation
Work model
Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.
Compensation
Hourly rate confirmed after the interview process.
Application process
Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.
- Turing is seeking graduate students or professionals for a remote role in evaluating AI-generated research reports. Responsibilities include reading, annotating, and scoring reports on a 1-5 scale, alongside providing written justifications. Candidates must possess strong...SuggestedRemote job
- A leading AI research accelerator is hiring a position focused on contributing to projects that evaluate and enhance AI systems. You will design community service scenarios, write structured explanations, and evaluate AI accuracy. The ideal candidate will have 4+ years...SuggestedRemote jobFull timeFor contractors
$190 per hour
...creative and technical talent with leading AI research labs. Headquartered in San... ...:20+ hours/week Role Responsibilities Evaluate AI systems on complex personal workflows... ...personal plugins/connectors. Write clear explanations of AI successes and failures to improve...SuggestedRemote jobHourly payFor contractorsSummer workTrial period- CNTXT AI is seeking a remote contractor to evaluate AI-generated financial content and develop test cases that probe analytical reasoning. You will help... ...AI models handle financial information with clear explanations and rigorous checks. Responsibilities include assessing...SuggestedRemote jobFor contractors
- ...remote, hourly contractor role supporting AI data and language projects on a project-... ...to support AI training datasets. LLM evaluation: reviewing AI-generated responses for... ...gaps, methodological errors, and unclear explanations even when language is fluent....SuggestedHourly payFor contractorsRemote workFlexible hours
$20 - $80 per hour
...Role Overview Train and evaluate next-generation AI systems by scoring model outputs, annotating real-world content, and delivering clear, actionable... ...knowledge across fields such as finance, healthcare, STEM engineering, and more, producing the annotations and feedback...Hourly payFor contractorsRemote work$20 per hour
A tech company specializing in AI is hiring a Digital Web Designer. In this remote role, you will evaluate AI-generated designs and provide feedback to enhance the model’s understanding of aesthetics. An ideal candidate will have a strong background in UI/UX design and...Remote workFlexible hours$14.5 per hour
...ethically sourced, relevant, diverse, and scalable to supercharge their AI models. As a Welocalize brand, Welo Data leverages over 25 years... ...online search experiences? Join our team as a Web Search Evaluator and help shape the future of search engines from the comfort of...Bi-weekly payHourly payPart timeImmediate startRemote workWork from homeFlexible hours- ...A leading AI company is seeking detail-oriented linguists with native Turkish fluency for a remote freelance opportunity. This role will focus on AI-related projects that involve prompt evaluation and multimedia content understanding. Ideal candidates will have a strong...FreelanceRemote work
- ...BAM Ventures is seeking Swedish-speaking remote annotators to evaluate AI-generated content, ensuring that it's coherent and aligns with real-world expectations. Your role will involve reviewing outputs, identifying deviations, and providing structured feedback to enhance...Remote work
- ...Alignerr is seeking a Search Quality Evaluator to assess search engine results, AI-generated answers, and content recommendations. This fully remote, flexible contract role values strong critical thinking and quality instincts, with no technical background required beyond...Contract workRemote workFlexible hours
- ...MERIT Beauty is seeking Turkish-speaking annotators for a contract role evaluating AI-generated content. You will review content for coherence, consistency, and alignment with real-world expectations in Turkish. The ideal candidates are native or fluent Turkish speakers...Contract workTemporary workImmediate startRemote work
$14.5 per hour
...ethically sourced, relevant, diverse, and scalable to supercharge their AI models. As a Welocalize brand, Welo Data leverages over 25 years... ...performance and provide insights on relevance and quality. Evaluate and rate the effectiveness of search engine results to ensure...Hourly payPart timeImmediate startRemote workWork from home10 hours per weekFlexible hours$14.5 per hour
A technology company is seeking an AI Web Search Evaluator to enhance the quality of search engine results. This flexible, remote, part-time role focuses on analyzing search performance and providing feedback to improve algorithms. Ideal candidates will have strong analytical...Hourly payPart timeRemote workFlexible hours$15 per hour
Productive Playhouse is building a pool of Spanish (Spain) speakers for evaluating leading AI chatbots. You will work as an independent contractor, choosing tasks, setting your own hours, and collaborating with clients globally. Pay is $15.00 USD per hour, with flexible...Remote jobHourly payFor contractorsFreelanceFlexible hours$20 - $26 per hour
Prolific is seeking fluent Kannada speakers to act as evaluators. You will assess how naturally and authentically AI captures Kannada speech, by listening to audio clips and comparing text and voice. This fast-paced project pays $20-26 per hour and may require about one...Remote jobHourly payWork from homeFlexible hours- About the role We are hiring expert Evaluators in Compliance / regulatory response with financial-services AI to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject‑...Hourly payWork at officeRemote work
$100 - $125 per hour
...hour fast-track onboarding required Key Responsibilities Translate real-world insurance workflows into structured tasks for AI systems Evaluate AI-generated outputs for accuracy, logical reasoning, and business relevance Work on use cases including underwriting,...Remote jobHourly payFor contractorsFreelanceWork at officeImmediate startFlexible hours- A virtual AI evaluation firm is seeking individuals to review and evaluate AI-generated responses in therapeutic conversations. The ideal candidate will possess strong written communication and analytical skills, as well as a keen attention to detail for assessing tone...Remote jobImmediate start
- Dorado is seeking an AI Language Quality Evaluator fluent in Greek and English for an ongoing, task-based project. This remote freelance role involves reviewing translated and AI-flagged content to judge accuracy, classify issues, and suggest corrected translations. You...Remote jobFor contractorsFreelanceFlexible hours
- Turing is seeking an experienced professional to review AI training issues and research environments for frontier AI labs. The ideal candidate will have a Master's/PhD or 4+ years experience in relevant engineering fields. This role emphasizes feedback skills, structured...Remote jobContract workFor contractors
$8 per hour
Dorado is seeking a Speech AI Evaluation Specialist to help improve AI-generated content in Thai or Chinese Simplified. This freelance, part-time role is remote from Thailand, offering 10+ hours weekly with an immediate start and a rate of 8 USD/hour. You will engage in...Remote jobPart timeFreelanceImmediate start$18 per hour
Dorado is seeking a Speech AI Evaluation Specialist to help improve AI-generated content in German. This freelance, part-time role offers a flexible remote schedule with work-from-home options in Germany. Start immediately with 10+ hours per week and a rate of 18 USD per...Remote jobHourly payPart timeFreelanceImmediate startWork from home10 hours per weekFlexible hours- Obsidian is hiring expert Evaluators in real estate, hospitality, and events to review AI-generated work products for accuracy and quality. This is a remote position requiring strong expertise in the relevant domains. Ideal candidates should have over 5 years of professional...Remote jobWork at office
$80 - $120 per hour
Mercor in New York, NY, is looking for a Legal contracts / diligence / redlines Evaluator to assess AI-generated artifacts. This role requires 5+ years of relevant experience and native or professional fluency in English. The ideal candidate will evaluate quality rubrics...Remote job- Unknown is seeking a Polish Bilingual Expert (contractor) to evaluate Polish audio content for AI training. The role emphasizes nativeness, fluency, and linguistic quality, with feedback delivered in English. Strong Polish language skills and clear written and verbal English...Remote jobFor contractors
- TELUS Digital AI Community invites freelance content evaluators in the United States to join a remote independent contractor role. You will help improve AI-powered search by reviewing online content, including text, images, videos, webpages, and search results, and provide...Remote jobFor contractorsFreelance
- About OpenTrain OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and... ...States English-language work 20+ hours per week About AI Response Evaluation AI training is the human side of building artificial...Part timeFor contractorsRemote work
- A leading tech company is seeking English AI Search Evaluators for a remote position in the United States. This part-time role offers flexible hours, allowing employees to set their own schedules. Candidates will provide ratings on search engine results based on project...Remote jobPart timeFlexible hours
- micro1 is seeking an AI Image & Video Evaluation Specialist (remote, contractor) to generate and compare AI-generated visuals across platforms. You will assess realism, composition, lighting, color, anatomy, and text rendering, delivering structured analyses. No formal...Remote jobFor contractors
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to STEM Explanation AI Evaluator [Remote]. Be the first to apply!


