Deep Research Task Evaluator [Remote]
AuraOne Human Data
- Remote job
Deep Research Task Evaluator is a remote review track for evaluating AI outputs across deep research task research review reasoning, calculations, and research workflows. Reviewers grade derivations and assumptions, reproduce key results, and document the correct method so the modeling team can train on it.
Why this role matters
Deep Research Task research review models live or die on whether their derivations actually hold up under scrutiny. AuraOne uses scientific specialists to grade outputs the way a peer reviewer would — checking assumptions, reproducing key steps, and capturing the right method alongside the wrong one.
Responsibilities
- Review AI outputs against current deep research task research review methods, conventions, and prior work for Deep Research Task Evaluator assignments.
- Reproduce or sanity-check key derivations, calculations, or experimental claims.
- Flag dimensional, methodological, and citation errors with structured severity tags.
- Capture the corrected reasoning or worked example so the modeling team can train on it.
- Adjudicate disputed answers against textbooks, papers, or community standards.
- Maintain reviewer-quality scores in inter-rater calibration cycles.
Qualifications
- Graduate-level training or equivalent applied experience in deep research task research review or a closely related field for Deep Research Task Evaluator work.
- Hands-on experience publishing, teaching, or advising on the topic at a professional level.
- Comfort applying multi-page rubrics consistently across long batches.
- Clear written reasoning that cites methods, papers, or worked examples.
- Reliable async availability for at least 10 hours per week.
Example tasks
- Reproduce a deep research task research review derivation from a model output and flag any algebraic or dimensional errors.
- Grade a model's literature summary against the cited papers and rate the citation quality.
- Adjudicate a disputed answer between two reviewers using textbook methods.
- Audit a 25-row batch for rubric consistency and report drift to the program lead.
Nice to have
- PhD, postdoc, or industry research experience in the topic area.
- Prior work reviewing AI-assisted research tooling and its failure modes.
- Multilingual fluency for non-English papers and corpora.
Skills
- Scientific reasoning
- Method validation
- Citation review
- Quantitative analysis
- Deep Research Task research review
- Web research
- Source grounding
- Browser automation
- Deep
- Research
Work model
Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.
Compensation
Hourly rate confirmed after the interview process.
Application process
Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.
- ...level expert in Quantitative Finance, the full-time Remote AI Research Evaluator will assess AI-generated financial content, craft relevant... ...Financial Mathematics, or a closely related quantitative field Deep familiarity with quantitative methods such as stochastic...SuggestedFull timeContract workRemote workFlexible hours
$16 - $20 per hour
...JOB DESCRIPTION AND POSITION REQUIREMENTS The research team at Penn State Ross and Carol Nese College of Nursing is hiring part‑time research evaluators for projects focused on dementia care in assisted living settings. The research evaluator will assist with in‑person...SuggestedHourly payPart timeSummer work- ...Browser Automation Task Evaluator is a remote evaluation track for reviewing browser automation task evaluation prompts and responses against... ...calibration Browser Automation Task evaluation Web research Source grounding Browser automation Browser Automation...SuggestedRemote jobHourly payFor contractors10 hours per week
- ...Government Policy and Legislative Research Evaluator is a remote review track for evaluating AI outputs in policy review workflows. Reviewers... ...async availability for at least 10 hours per week. Example tasks Verify the citations in a model's policy review memo and...SuggestedRemote jobHourly payFor contractors10 hours per week
$50 per hour
...Overview Lead transcription, annotation, and evaluation of Hindi audio and video to help train... ...assurance and review cycles to validate task design, rubrics, and outputs before... ...annotation, localization, evaluation, or research workflows in Hindi. Familiarity with...SuggestedHourly payTemporary workRemote work10 hours per week- ...Biomedical Research AI Evaluator is a remote clinical-review track for evaluating AI outputs that touch biomedical. Reviewers grade differential... ...availability for at least 10 hours per week. Example tasks Grade a model's differential-diagnosis reasoning on a biomedical...Remote jobHourly payFor contractors10 hours per week
- ...by our partnership with EQT. Website: Linkedin Job Title: Evaluator - Political Science Location: Remote (USA) Job Type: Contract... ...from quality checks. Stay Current: Stay updated with the latest research, guidelines, and advancements in your area of expertise,...Contract workRemote workWorldwide
- ...Equity Research AI Evaluator is a remote review track for evaluating AI outputs across finance and risk workflows. Reviewers grade calculations... ...availability for at least 10 hours per week. Example tasks Re-perform a finance and risk calculation produced by a model...Remote jobHourly payFor contractorsWork experience placement10 hours per week
- ...Dual-Use Research Risk Evaluator is a remote red-team track for stress-testing AI systems against adversarial prompts. Reviewers craft attack scenarios... ...availability for at least 10 hours per week. Example tasks Construct a 5-turn adversarial conversation that bypasses...Remote jobHourly payFor contractors10 hours per week
$80 - $120 per hour
...creative and technical talent with leading AI research labs. Headquartered in San Francisco, our... ...customer research and feedback synthesis Evaluator Type: Contract Compensation: $... ...to enhance AI work products. Apply deep subject-matter expertise to grade outputs...Contract workSummer workWork at officeRemote work$40 per hour
A leading data services company in the United States is seeking a Research Scientist (Chemistry) to enhance AI models by evaluating their performance with complex chemistry questions. The role allows for flexible scheduling and project selection. Ideal candidates will have...Hourly payContract workRemote workFlexible hours- ...engineering executives on product decisions, process flows, and compensation plans YOU BRING Strong analytical, problem-solving, and research skills High proficiency in using technology, including software, apps, and databases Ability to work independently and rely...Full timeTraineeshipRemote workLong distance
$80 - $120 per hour
Role Description ~Evaluate AI-generated artifacts against domain-specific quality rubrics. ~Identify factual, aesthetic, and presentation... ...feedback to improve AI-generated work products. ~Apply deep subject-matter expertise in Brand, creative direction, and marketing...Part timeWork at officeRemote work$60 per hour
A leading AI research firm is seeking experienced quantitative professionals to evaluate AI-generated analysis and solve technical problems. This fully remote position offers flexible scheduling and competitive pay of up to $60 per hour. Ideal candidates have 2+ years of...Hourly payRemote workFlexible hours$20 per hour
...company specializing in AI is looking for a Digital Web Designer to evaluate AI-generated designs and help train models for better aesthetic... ...pay rates start at $20/hr for general projects and $40/hr for design-focused tasks, with possible bonuses. #J-18808-LjbffrRemote work$20 per hour
...is hiring a Digital Web Designer. In this remote role, you will evaluate AI-generated designs and provide feedback to enhance the model’... .... The position offers flexible hours and pays $20+ USD/hr for general projects and $40+ for design-focused tasks. #J-18808-LjbffrRemote workFlexible hours$20 per hour
...company focused on AI and design is seeking a Digital Designer to evaluate AI-generated designs and enhance AI's understanding of user-... ...from $20/hr for general projects and $40/hr for design-specific tasks. Candidates should possess strong design backgrounds and...Remote work$20 per hour
...firm in the United States is seeking a Digital Web Designer to evaluate AI-generated designs and enhance the model's understanding of aesthetics... ...and involves both general AI training and design-focused tasks, with compensation starting at $20/hr. Candidates should have a...Remote work$20 per hour
...looking for an Experience Designer to improve AI model outputs by evaluating design work, including interfaces and visuals. This role offers... ..., with compensation starting at $20+ USD/hr for general AI tasks and $40+ USD/hr for design-focused work. Applicants must be fluent...Contract workRemote workFlexible hours$20 per hour
...Candidates should possess a strong design background and fluency in English. Compensation starts at $20+ USD/hr for general AI projects and $40+ USD/hr for design-focused tasks. This is a remote, independent contract position for applicants in the United States. #J-18808-LjbffrContract workRemote work$20 per hour
A leading data annotation firm is seeking a Web Designer to evaluate AI-generated designs and enhance their understanding of design principles... ...at $20/hr for general projects and $40/hr for design-specific tasks, with bonus opportunities for exceptional work. #J-18808-LjbffrRemote work- ...Prolific is seeking fluent Thai speakers to act as evaluators for AI language models. You will assess how naturally Thai is spoken in AI... ...ensuring cultural and contextual accuracy. This is a remote, paid task with flexible hours and competitive pay. Applicants should have...Remote workFlexible hours
$60 per hour
...team, you'll work closely with state-of-the-art AI models on tasks like evaluating AI-generated quantitative analysis, solving technical... ...science, astrophysics, economics, biostatistics, operations research, or any other quantitative field, if you think rigorously about...Hourly payFull timeRemote workFlexible hours$20 per hour
...outputs and help refine design quality. Applicants should have a background in UI/UX and be fluent in English. This is a remote position with project flexibility and competitive pay starting at $20/hr for general projects and $40/hr for design-focused tasks. #J-18808-LjbffrRemote work$25 - $30 per hour
...to help train AI chatbots. You will manage various tasks, write high-quality responses, and evaluate AI models based on guidelines. This role offers flexibility... ...5–$30+. Candidates should be fluent in English and possess strong writing and research skills. #J-18808-Ljbffr...Hourly payContract workFor contractorsRemote workWork from home$80 - $120 per hour
Role Description ~Evaluate AI-generated artifacts against domain-specific quality rubrics.... ...feedback to improve AI model outputs. ~Apply deep subject-matter expertise to grade outputs... ...and technical talent with leading AI research labs. Headquartered in San Francisco, our...Summer workWork at office$60 per hour
...practice. That's where you come in. As a member of DataAnnotation's team, you'll work closely with state-of-the-art AI models on tasks like evaluating AI-generated security content, solving technical security problems, and providing feedback that directly shapes how these...Remote jobHourly payFull timeFlexible hours- ...creative and technical talent with leading AI research labs. They are seeking a Market research / competitive intelligence Evaluator. This remote position offers a chance to... ...provide structured feedback while leveraging deep expertise in market research. The ideal candidate...Remote jobWork at officeFlexible hours
$5,083 per month
...will begin on February 11, 2026. This position specializes in evaluation functions under the general supervision of the University Registrar... ...and degree audit programs to assure timely completion of tasks. Accurately enter and update student information in administrative...Full timeWork at officeRemote work- ...Digita l are currently hiring for a Personalized Internet Ads Evaluator role! This is a freelance, independent contractor position that... ...application to be installed on your smartphone to complete certain tasks. Assessment In order to be hired into the program, you’ll be...Part timeFor contractorsFreelanceCurrently hiringImmediate startWork from homeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Deep Research Task Evaluator [Remote]. Be the first to apply!



