Deep Research Task Evaluator [Remote]
AuraOne Human Data
- Remote job
Deep Research Task Evaluator is a remote review track for evaluating AI outputs across deep research task research review reasoning, calculations, and research workflows. Reviewers grade derivations and assumptions, reproduce key results, and document the correct method so the modeling team can train on it.
Why this role matters
Deep Research Task research review models live or die on whether their derivations actually hold up under scrutiny. AuraOne uses scientific specialists to grade outputs the way a peer reviewer would — checking assumptions, reproducing key steps, and capturing the right method alongside the wrong one.
Responsibilities
- Review AI outputs against current deep research task research review methods, conventions, and prior work for Deep Research Task Evaluator assignments.
- Reproduce or sanity-check key derivations, calculations, or experimental claims.
- Flag dimensional, methodological, and citation errors with structured severity tags.
- Capture the corrected reasoning or worked example so the modeling team can train on it.
- Adjudicate disputed answers against textbooks, papers, or community standards.
- Maintain reviewer-quality scores in inter-rater calibration cycles.
Qualifications
- Graduate-level training or equivalent applied experience in deep research task research review or a closely related field for Deep Research Task Evaluator work.
- Hands-on experience publishing, teaching, or advising on the topic at a professional level.
- Comfort applying multi-page rubrics consistently across long batches.
- Clear written reasoning that cites methods, papers, or worked examples.
- Reliable async availability for at least 10 hours per week.
Example tasks
- Reproduce a deep research task research review derivation from a model output and flag any algebraic or dimensional errors.
- Grade a model's literature summary against the cited papers and rate the citation quality.
- Adjudicate a disputed answer between two reviewers using textbook methods.
- Audit a 25-row batch for rubric consistency and report drift to the program lead.
Nice to have
- PhD, postdoc, or industry research experience in the topic area.
- Prior work reviewing AI-assisted research tooling and its failure modes.
- Multilingual fluency for non-English papers and corpora.
Skills
- Scientific reasoning
- Method validation
- Citation review
- Quantitative analysis
- Deep Research Task research review
- Web research
- Source grounding
- Browser automation
- Deep
- Research
Work model
Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.
Compensation
Hourly rate confirmed after the interview process.
Application process
Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.
$40 - $100 per hour
...work About AI Training and Scientific Evaluation AI training is the human side of... ...realistic scenarios, calculations, and research-style reasoning tasks. Focus on scientific evaluation of AI... ...planetary science, or space physics Deep knowledge of relevant fundamentals, including...SuggestedHourly payContract workPart timeFor contractorsRemote work- ...Seeking a full-time Remote AI Research Evaluator with a PhD in Quantitative Finance to assess and enhance AI models' capabilities in financial... ...Mathematics, or a closely related quantitative field Deep familiarity with quantitative methods such as stochastic modeling...SuggestedFull timeContract workRemote workFlexible hours
- Productive Playhouse seeks a Portugal Portuguese AI Product Evaluator for an upcoming project testing and evaluating leading AI chatbots... .... This is an independent contractor engagement with batches of tasks; you set your hours and may work with other clients. Transitions...SuggestedRemote jobFor contractorsFreelanceFlexible hours
- Mercor seeks expert Evaluators in User/customer research and feedback synthesis to review AI-generated artifacts (documents, spreadsheets, and slide decks... ...for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. This is a...SuggestedRemote jobHourly pay
- ...adjudicate high-level physics arguments and guide evaluation of complex scientific claims. You will apply deep subject-matter expertise to judge reasoning and provide... .... The ideal candidate has a strong independent research record, leadership, and the ability to articulate...SuggestedRemote jobFor contractors
- Mercor is seeking expert Evaluators in user and customer research to review AI-generated work products for accuracy, rigor, and domain quality. This remote, hourly role requires applying deep subject-matter expertise to grade outputs and provide clear, structured feedback...Remote jobHourly pay
$16 - $20 per hour
...JOB DESCRIPTION AND POSITION REQUIREMENTS The research team at Penn State Ross and Carol Nese College of Nursing is hiring part‑time research evaluators for projects focused on dementia care in assisted living settings. The research evaluator will assist with in‑person...Hourly payPart timeSummer work- ...by our partnership with EQT. Website: Linkedin Job Title: Evaluator - Political Science Location: Remote (USA) Job Type: Contract... ...from quality checks. Stay Current: Stay updated with the latest research, guidelines, and advancements in your area of expertise,...Contract workRemote workWorldwide
- ...Principal Software Engineer - GitHub Platform & Enterprise Principal Software Engineer - GitHub Platform & Enterprise Software Engineer, Research Developer Productivity Staff Developer Advocate, GitHub Security Lab Software Engineer, Developer Productivity Staff Software...Remote jobFor contractors
- ...Surgical Robotics Task Evaluator is a remote review track for evaluating robotics review outputs against real-world physical-AI constraints. Reviewers grade trajectory choices, perception assumptions, and safety envelopes; flag failure modes; and document the corrected...Remote jobHourly payFor contractors10 hours per week
- ...Search Query Planning Task Evaluator is a remote evaluation track for reviewing search query planning task evaluation prompts and responses... ...calibration Search Query Planning Task evaluation Web research Source grounding Browser automation Search Query...Remote jobHourly payFor contractors10 hours per week
- ...Government Policy and Legislative Research Evaluator is a remote review track for evaluating AI outputs in policy review workflows. Reviewers... ...async availability for at least 10 hours per week. Example tasks Verify the citations in a model's policy review memo and...Remote jobHourly payFor contractors10 hours per week
- ...Robot Policy Compliance Task Evaluator is a remote red-team track for stress-testing AI systems against adversarial prompts. Reviewers craft... ...Demonstrated experience red-teaming AI systems, security research, or adversarial ML work for Robot Policy Compliance Task Evaluator...Remote jobHourly payFor contractors10 hours per week
- ...Equity Research AI Evaluator is a remote review track for evaluating AI outputs across finance and risk workflows. Reviewers grade calculations... ...availability for at least 10 hours per week. Example tasks Re-perform a finance and risk calculation produced by a model...Remote jobHourly payFor contractorsWork experience placement10 hours per week
$36 - $40 per hour
...Health is looking for a Clinical Evaluator - PRN to join our growing team... ...Utilize independent judgment, deep clinical expertise,... ...literature, regulations, and research pertaining to long-term care,... ...assessment activities. Execute tasks with strong attention to detail...Hourly payContract workReliefWork at officeLocal areaRemote workWork from homeHome office$23.2 per hour
...help us improve Roblox systems. As a Human Evaluator you will have the opportunity to provide... ...Demonstrated patience for repetitive tasks and attention to detail Effective time management... ...moderation or data annotation. A deep passion for gaming trends and the Roblox...Hourly payFull timeContract workWork experience placementWork at officeLocal areaRemote workMonday to Friday- ...Biomedical Research AI Evaluator is a remote clinical-review track for evaluating AI outputs that touch biomedical. Reviewers grade differential... ...availability for at least 10 hours per week. Example tasks Grade a model's differential-diagnosis reasoning on a biomedical...Remote jobHourly payFor contractors10 hours per week
- ...Professor/PI to provide high-level scientific expertise for AI-system evaluation. Remote contractor role focusing on evaluating physics... ...senior physics professionals. The role emphasizes leadership in research judgment, awareness of uncertainties, and clear communication...Remote jobFor contractors
$100 - $150 per hour
...technical talent with leading AI research labs. Headquartered in San... ...Responsibilities Design precise, task-specific grading criteria for... ...into large enterprises. Deep fluency in outbound prospecting... ...experience with AI training, evaluation, or human-data projects. #J-1...Summer workRemote work- ...expertise for next‑gen AI systems. Remote contractor role focusing on evaluating complex physics arguments and providing authoritative judgments. The ideal candidate has a strong independent research record and leadership in science. No prior AI experience is required;...Remote jobFor contractors
- ...a project focused on advancing next-generation AI systems. Remote involvement, contractor role, and strong research leadership are essential. You will evaluate complex physics arguments, calibrate expert judgments, and draft rigorous written evaluations for senior researchers...Remote jobFor contractors
$40 per hour
...flexible position pays $40+/h and requires deep expertise in neuroscience, along with advanced academic credentials. Your main tasks include reviewing scientific papers and... ...Join us to influence how AI accurately summarizes complex research. #J-18808-Ljbffr ProlificRemote jobFlexible hours- ...- Fri approx. 60 hours per month for up to 4 years. Remote. Department of Children, Youth and Families Qualitative Research Evaluator Position Summary The Rhode Island Behavioral Health System of Care (RISOC) Qualitative Research Evaluator Contractor...Temporary workFor contractorsLocal areaRemote work
$52k - $57k
...POSITION TITLE: Program Evaluator REPORTS TO: Director of Quality BROAD FUNCTION: Collects... ...with both quantitative and qualitative research methods. Proficiency with applicable software... ...and self-direction to complete assigned tasks and identify tools and approaches for...Full timeSummer workLocal area$20 per hour
...is hiring a Digital Web Designer. In this remote role, you will evaluate AI-generated designs and provide feedback to enhance the model’... .... The position offers flexible hours and pays $20+ USD/hr for general projects and $40+ for design-focused tasks. #J-18808-LjbffrRemote workFlexible hours- YO AI Labs seeks Medical Evaluation Specialists to craft and validate challenging medical questions for evaluating AI systems. This remote... ...AI outputs for clinical accuracy and reasoning. Medical knowledge and research skills are essential. #J-18808-Ljbffr YO AI LabsFor contractorsRemote work
- ...About the role We are hiring expert Evaluators in Program management / implementation planning to review and assess AI-generated work products... ...decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. This is a remote,...Hourly payWork at officeRemote work
$80 - $120 per hour
...Compensation: $80 - $120 per hour We are hiring expert Evaluators in Special education / IEP to review and assess AI-generated... ...decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. This is a remote,...Hourly payContract workFor contractorsWork at officeRemote work$80 - $120 per hour
...Special Education Iep Evaluator This role is for one of our clients. Compensation: $80 - $120 per hour. We are hiring expert Evaluators... ...decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. This is a remote,...Hourly payRemote work$80 - $120 per hour
...Compensation: $80 - $120 per hour We are hiring expert Evaluators in Nonprofit / philanthropy / community programs to review and... ...decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. This is a remote,...Hourly payContract workFor contractorsWork at officeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Deep Research Task Evaluator [Remote]. Be the first to apply!


