Statistics and Probability AI Evaluator [Remote]
AuraOne Human Data
- Remote job
Statistics and Probability AI Evaluator is a remote review track for evaluating AI outputs across statistics reasoning, calculations, and research workflows. Reviewers grade derivations and assumptions, reproduce key results, and document the correct method so the modeling team can train on it.
Why this role matters
Statistics models live or die on whether their derivations actually hold up under scrutiny. AuraOne uses scientific specialists to grade outputs the way a peer reviewer would — checking assumptions, reproducing key steps, and capturing the right method alongside the wrong one.
Responsibilities
- Review AI outputs against current statistics methods, conventions, and prior work for Statistics and Probability AI Evaluator assignments.
- Reproduce or sanity-check key derivations, calculations, or experimental claims.
- Flag dimensional, methodological, and citation errors with structured severity tags.
- Capture the corrected reasoning or worked example so the modeling team can train on it.
- Adjudicate disputed answers against textbooks, papers, or community standards.
- Maintain reviewer-quality scores in inter-rater calibration cycles.
Qualifications
- Graduate-level training or equivalent applied experience in statistics or a closely related field for Statistics and Probability AI Evaluator work.
- Hands-on experience publishing, teaching, or advising on the topic at a professional level.
- Comfort applying multi-page rubrics consistently across long batches.
- Clear written reasoning that cites methods, papers, or worked examples.
- Reliable async availability for at least 10 hours per week.
Example tasks
- Reproduce a statistics derivation from a model output and flag any algebraic or dimensional errors.
- Grade a model's literature summary against the cited papers and rate the citation quality.
- Adjudicate a disputed answer between two reviewers using textbook methods.
- Audit a 25-row batch for rubric consistency and report drift to the program lead.
Nice to have
- PhD, postdoc, or industry research experience in the topic area.
- Prior work reviewing AI-assisted research tooling and its failure modes.
- Multilingual fluency for non-English papers and corpora.
Skills
- Scientific reasoning
- Method validation
- Citation review
- Quantitative analysis
- Statistics
- Science and advanced mathematics
- AI evaluation
- Rubric writing
- Expert review
Work model
Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.
Compensation
Hourly rate confirmed after the interview process.
Application process
Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.
$20 per hour
A tech company specializing in AI is hiring a Digital Web Designer. In this remote role, you will evaluate AI-generated designs and provide feedback to enhance the model’s understanding of aesthetics. An ideal candidate will have a strong background in UI/UX design and...SuggestedRemote workFlexible hours$11.5 per hour
...position as an Online Task Contributor. In this role, you will evaluate and provide feedback on content to enhance search engine results... ...at $11.50 hourly, based on task completion, with a supportive community of contributors involved in AI advancements. #J-18808-LjbffrSuggestedHourly payPart timeRemote work$80 - $120 per hour
...Cybersecurity / IT GRC Evaluator $80-$120 per hour Hourly contract Remote About the Role We are hiring expert Evaluators in Cybersecurity / IT GRC to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor...SuggestedHourly payContract workFor contractorsWork at officeRemote work- ...Seeking a full-time Remote AI Research Evaluator with a PhD in Quantitative Finance to assess and enhance AI models' capabilities in financial reasoning and quantitative analysis through flexible, contract-based work. Key responsibilities Assessing the factuality and...SuggestedFull timeContract workRemote workFlexible hours
- ...Obsidian is hiring expert Evaluators in real estate, hospitality, and events to review AI-generated work products for accuracy and quality. This is a remote position requiring strong expertise in the relevant domains. Ideal candidates should have over 5 years of professional...SuggestedWork at officeRemote work
$14.5 per hour
...Join to apply for the AI Web Search Evaluator role at Welo Data Welo Data works with technology companies to provide datasets that are high-quality, ethically sourced, relevant, diverse, and scalable to supercharge their AI models. As a Welocalize brand, Welo Data...Hourly payPart timeImmediate startRemote workWork from home10 hours per weekFlexible hours- Prolific in Edison, New Jersey is looking for fluent Gujarati speakers to assist as evaluators in AI training. This involves assessing the naturalness and authenticity of AI-generated Gujarati speech through audio and text evaluations. Candidates will enjoy flexible hours...Remote jobHourly payFreelanceFlexible hours
- Rex.zone is seeking a Senior AI data annotator to perform data labeling and evaluation for NLP tasks, RLHF assessments, and prompt QA to improve training data quality and model performance. This is a US-based remote, full-time role aligned with Miami talent demand. You...Remote jobFull time
- Mercor is seeking expert Evaluators in Legal/compliance to review AI-generated documents, spreadsheets, and slide decks for accuracy and quality. This is a remote, hourly engagement requiring deep subject-matter expertise and careful judgment. You will assess artifacts...Remote jobHourly payWork at office
- Turing is seeking detail-oriented AI Analysts based in the United States for a Google Wallet evaluation project. This role allows you to engage with advanced AI tools while contributing to the future of AI. You will evaluate model responses, review output quality, and provide...Remote jobFull timeContract work
- AuraOne is seeking Italian Music & Lyrics Expert - AI Evaluation (Remote) to review prompts and model outputs against AuraOne's quality rubric. You will label edge cases and write structured feedback to help retrain models. As a CONTRACTOR, you will compare paired responses...Remote jobFor contractors10 hours per week
- Obsidian is hiring expert Evaluators in Investment analysis / valuation / credit to review AI-generated work products for accuracy and quality. This remote, hourly position requires deep subject-matter expertise and professional fluency in English to provide structured...Hourly payWork at officeRemote work
- Turing is seeking an experienced professional to review AI training issues and research environments for frontier AI labs. The ideal candidate will have a Master's/PhD or 4+ years experience in relevant engineering fields. This role emphasizes feedback skills, structured...Remote jobContract workFor contractors
$8 per hour
Dorado is seeking a Speech AI Evaluation Specialist to help improve AI-generated content in Thai or Chinese Simplified. This freelance, part-time role is remote from Thailand, offering 10+ hours weekly with an immediate start and a rate of 8 USD/hour. You will engage in...Remote jobPart timeFreelanceImmediate start- Dorado is seeking Speech AI Evaluation Specialists based in Thailand to help improve Vietnamese-language AI content. This is a remote, freelance role with Bangkok as the working location and a flexible, part-time schedule. You will participate in short voice conversations...Remote jobPart timeFreelanceImmediate start10 hours per weekFlexible hours
- Dorado is seeking an AI Language Quality Evaluator fluent in Greek and English for an ongoing, task-based project. This remote freelance role involves reviewing translated and AI-flagged content to judge accuracy, classify issues, and suggest corrected translations. You...Remote jobFor contractorsFreelanceFlexible hours
- A virtual AI evaluation firm is seeking individuals to review and evaluate AI-generated responses in therapeutic conversations. The ideal candidate will possess strong written communication and analytical skills, as well as a keen attention to detail for assessing tone...Remote jobImmediate start
- YO AI Labs is seeking Dutch bilingual experts to remotely evaluate AI-generated Dutch speech for nativeness and quality, providing detailed written feedback and documented findings. The role requires native Dutch, strong English, and high attention to detail to ensure linguistic...Remote job
- About the role We are hiring expert Evaluators in Compliance / regulatory response with financial-services AI to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject‑...Hourly payWork at officeRemote work
$100 - $125 per hour
...hour fast-track onboarding required Key Responsibilities Translate real-world insurance workflows into structured tasks for AI systems Evaluate AI-generated outputs for accuracy, logical reasoning, and business relevance Work on use cases including underwriting,...Remote jobHourly payFor contractorsFreelanceWork at officeImmediate startFlexible hours- A leading AI research firm is looking for a contractor in internal or emergency medicine to help enhance language models by providing expert feedback based on real-world medical scenarios. Candidates should possess an MD and have relevant experience, including board certification...Remote jobFor contractors
- About OpenTrain OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and... ...States English-language work 20+ hours per week About AI Response Evaluation AI training is the human side of building artificial...Part timeFor contractorsRemote work
- Productive Playhouse is building a talent pool of Armenian speakers to test and evaluate leading AI chatbots. This freelance, project-based opportunity lets you choose tasks, set your own hours, and work with other clients as needed. Open to freelancers outside the U.S....Remote jobFreelanceFlexible hours
$14.5 per hour
A technology company is seeking an AI Web Search Evaluator to enhance the quality of search engine results. This flexible, remote, part-time role focuses on analyzing search performance and providing feedback to improve algorithms. Ideal candidates will have strong analytical...Remote jobHourly payPart timeFlexible hours- Productive Playhouse is building an Albanian-speaking evaluator pool for AI chatbot testing. As an AI Evaluator, you’ll interact with AI models to assess capabilities, safety, and usefulness, providing data to improve models. Tasks come in batches with flexible hours and...Remote jobFreelanceFlexible hours
$80 - $120 per hour
Mercor in New York, NY, is looking for a Legal contracts / diligence / redlines Evaluator to assess AI-generated artifacts. This role requires 5+ years of relevant experience and native or professional fluency in English. The ideal candidate will evaluate quality rubrics...Remote job$80 - $120 per hour
Mercor is looking for a Biology / environmental science Evaluator to assess AI-generated artifacts based on quality rubrics. This position requires evaluating documents for errors and collaborating with AI teams to improve model performance. The ideal candidate should...Remote jobHourly payContract workWork at office- A leading AI research accelerator is hiring a position focused on contributing to projects that evaluate and enhance AI systems. You will design community service scenarios, write structured explanations, and evaluate AI accuracy. The ideal candidate will have 4+ years...Remote jobFull timeFor contractors
$18 per hour
Dorado is seeking a Speech AI Evaluation Specialist to help improve AI-generated content in German. This freelance, part-time role offers a flexible remote schedule with work-from-home options in Germany. Start immediately with 10+ hours per week and a rate of 18 USD per...Remote jobHourly payPart timeFreelanceImmediate startWork from home10 hours per weekFlexible hours- YO AI Labs is seeking experienced Specialized Industry Experts to evaluate AI-powered workflows across Government Contracting, Grants, Education, and related sectors. Remote work is supported; your industry expertise and professional judgment matter most. You will test...Remote job
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Statistics and Probability AI Evaluator [Remote]. Be the first to apply!

