Theorem Proving Model Evaluator [Remote]
AuraOne Human Data
- Remote job
Theorem Proving Model Evaluator is a remote review track for evaluating AI outputs across theorem proving model research review reasoning, calculations, and research workflows. Reviewers grade derivations and assumptions, reproduce key results, and document the correct method so the modeling team can train on it.
Why this role matters
Theorem Proving Model research review models live or die on whether their derivations actually hold up under scrutiny. AuraOne uses scientific specialists to grade outputs the way a peer reviewer would — checking assumptions, reproducing key steps, and capturing the right method alongside the wrong one.
Responsibilities
- Review AI outputs against current theorem proving model research review methods, conventions, and prior work for Theorem Proving Model Evaluator assignments.
- Reproduce or sanity-check key derivations, calculations, or experimental claims.
- Flag dimensional, methodological, and citation errors with structured severity tags.
- Capture the corrected reasoning or worked example so the modeling team can train on it.
- Adjudicate disputed answers against textbooks, papers, or community standards.
- Maintain reviewer-quality scores in inter-rater calibration cycles.
Qualifications
- Graduate-level training or equivalent applied experience in theorem proving model research review or a closely related field for Theorem Proving Model Evaluator work.
- Hands-on experience publishing, teaching, or advising on the topic at a professional level.
- Comfort applying multi-page rubrics consistently across long batches.
- Clear written reasoning that cites methods, papers, or worked examples.
- Reliable async availability for at least 10 hours per week.
Example tasks
- Reproduce a theorem proving model research review derivation from a model output and flag any algebraic or dimensional errors.
- Grade a model's literature summary against the cited papers and rate the citation quality.
- Adjudicate a disputed answer between two reviewers using textbook methods.
- Audit a 25-row batch for rubric consistency and report drift to the program lead.
Nice to have
- PhD, postdoc, or industry research experience in the topic area.
- Prior work reviewing AI-assisted research tooling and its failure modes.
- Multilingual fluency for non-English papers and corpora.
Skills
- Scientific reasoning
- Method validation
- Citation review
- Quantitative analysis
- Theorem Proving Model research review
- Formal reasoning
- Proof review
- Theorem
- Proving
Work model
Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.
Compensation
Hourly rate confirmed after the interview process.
Application process
Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.
$20 per hour
...public sources and external tools. Generate high-quality human evaluation data by identifying response strengths, areas for improvement,... ...quality, clarity, tone, and completeness of responses. Ensure model responses align with expected conversational behavior and system...SuggestedRemote jobContract workPart timeSummer work$65 per hour
Prolific is seeking registered nurses in Las Vegas, NV, to assist in training AI models. You will review AI responses, rate their accuracy and safety, and write feedback to enhance AI learning. Candidates must be verified registered nurses, have recent clinical experience...SuggestedHourly paySelf employmentWork from homeFlexible hours- ...Geometry Reasoning Model Evaluator is a remote evaluation track for reviewing geometry reasoning model evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling...SuggestedRemote jobHourly payFor contractors10 hours per week
- ...Frontier Model Misuse Risk Evaluator is a remote red-team track for stress-testing AI systems against adversarial prompts. Reviewers craft attack scenarios, document the failure mode, and pair each successful jailbreak with the rubric clause it violated so the safety team...SuggestedRemote jobHourly payFor contractors10 hours per week
- ...Quant Finance Reasoning Model Evaluator is a remote review track for evaluating AI outputs across finance and risk workflows. Reviewers grade calculations, narrative reasoning, and policy adherence; flag compliance and reconciliation issues; and document the correct treatment...SuggestedRemote jobHourly payFor contractorsWork experience placement10 hours per week
- ...Policy Preference Reward Model Evaluator is a remote review track for evaluating AI outputs in policy review workflows. Reviewers grade citation accuracy, statutory reasoning, and policy adherence; flag risk; and document the corrected analysis so the modeling team can...Remote jobHourly payFor contractors10 hours per week
- ...reasoning and computational problem solving to improve and evaluate large language models. You will design rigorous math problems, produce clear,... ...practical Python implementation, and the use of formal theorem proving to push the reasoning capabilities of state of the art...Contract workFor contractorsFreelanceRemote work
- ...where that happens: every product built on a model is bounded by what it costs to run, so... ...for: ~ This role builds the evaluation and decision systems that make agentic inference... ...what an evaluation result does and does not prove Ability to analyze per-request and per...Full time
- ...position is M-F during standard business hours with a hybrid work model (4 days in-office, 1 day working from home). It is only... ...internal analysts and external consultants on validation activities. Evaluate model performance monitoring and complete annual model reviews....Full timeWork at officeWork from homeRelocation
- ...work.Work You’ll DoAs a Project - Program Analytics & Insights Evaluator IIon the project, you will be responsible for:Supporting the design... ...programsDeveloping and refining evaluation frameworks, logic models, research questions, and evaluation plans aligned with program...Local area
- Overview At PAM Health, we care for chronically and critically ill patients who require extended hospital care. PAM Health has over 80 hospital locations and employs over 11,000 people across the country. Our teams work together to deliver the highest level of compassionate...Full timeLocal area
- ...Risk assessment experience required• Excellent written communication skills• Thorough knowledge of psychopathology and its treatment, models of behavior change and management, the psychotherapeutic process, and models of personality development.• Knowledge of and...Work experience placementWork at officeLocal areaRemote work
- HealthAlliance Hospital · Psych Emer RoomKingston, NYAllied Health Prof/TechnicalPer DiemAll ShiftsvariedThe Mental Health Evaluator (MHE) is a critical role dedicated to delivering comprehensive, clinically appropriate, culturally competent, and trauma-informed patient...
- ...We are hiring for: Family Model Provider Type: Family Model Provider (TN) - Independent Contractor If you are a positive and personable individual looking for a satisfying and fun opportunity to make a real difference in the lives of people with intellectual...Full timeContract workFor contractorsLive inWork from home
- ...a caring, compassionate, and reliable person? Here is an opportunity to make a lasting impact on someone’s life – become an Family Model Provider !! This is a chance to help make a difference by opening your home to an individual with developmental and/or physical disabilities...Daily paid
- ...programs and deliver value to our business. At Applied Intuition, You Will: Conduct research on pretraining world-action foundation model with various world modalities including vision and physics associated with ego actions, serving the purpose for both robot action...Full timeFor contractorsFor subcontractorCasual workInternshipWork at officeImmediate startRemote workDay shift
- ...Luxury Brand Evaluator Turn your passion for luxury into a career opportunity. Explore the world of premium brands and make a lasting impact in fashion, beauty, jewelry, or automobiles. Join CXG, the global leader in customer experience, and work alongside iconic names...Remote workWorldwideFlexible hours
$52k - $57k
...Job Description Job Description POSITION TITLE: Program Evaluator REPORTS TO: Director of Quality BROAD FUNCTION: Collects, analyzes, and reports on data to evaluate effectiveness of program services. I. CORE VALUES: # CULTURAL PROFICIENCY: Articulates...Full timeSummer workLocal area- ...affordable health care to underserved communities in the Mississippi Delta and Southwest Georgia region. Our research focuses on program evaluation and provision of care in rural communities. Role Description The Program Evaluator is a full-time remote role open to...Full timeRemote work
- Are you passionate about helping others? Are you a caring, compassionate, and reliable person? Here is an opportunity to make a lasting impact on someone’s life – become an Family ModelProvider ! This is a chance to help make a difference by opening your home to an individual...Daily paid
- ...In conjuncion with Talent Model Recruiters, we are seeking new and experienced models for the apparel and fashion industry to display clothing and merchandise in commercials, advertisements, and/or fashion shows. Promote products and services in online ads, social media...Part timeFlexible hours
$16 - $20 per hour
...AND POSITION REQUIREMENTS The research team at Penn State Ross and Carol Nese College of Nursing is hiring part-time research evaluators for projects focused on dementia care in assisted living settings. The research evaluator will assist with in-person...Hourly payPart timeFor contractorsSummer workRemote work$9 - $30 per hour
ALTA Language Services, Inc. is looking for a remote Testing Evaluator in Atlanta, GA. This part-time position involves assisting with language proficiency examinations and evaluating tests according to set criteria. Candidates should have strong communication skills, professionalism...Part timeRemote work$20 per hour
A tech company specializing in AI is hiring a Digital Web Designer. In this remote role, you will evaluate AI-generated designs and provide feedback to enhance the model’s understanding of aesthetics. An ideal candidate will have a strong background in UI/UX design and...Remote workFlexible hours$80 - $120 per hour
...Special Education Iep Evaluator This role is for one of our clients. Compensation: $80 - $120 per hour. We are hiring expert Evaluators in special education / IEP to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy...Hourly payRemote work$79.4k - $119.1k
...This position represents Auto Spec Control department in a development team environment as Project Lead on a mix of complex Full Model Change (FMC) and Minor Model Change (MMC) developments. Creates, promotes, and manages critical milestones for successful package delivery...Full timeTemporary workWork experience placementWork at officeRemote workRelocation package$80 - $120 per hour
...This role is for one of our clients Compensation: $80 - $120 per hour We are hiring expert Evaluators in Special education / IEP to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality...Hourly payContract workFor contractorsWork at officeRemote work- ...Seeking a full-time Remote AI Research Evaluator with a PhD in Quantitative Finance to assess and enhance AI models' capabilities in financial reasoning and quantitative analysis through flexible, contract-based work. Key responsibilities Assessing the factuality and...Full timeContract workRemote workFlexible hours
- ...A leading research accelerator is seeking a contractor to evaluate North American teen humor. The role involves reviewing short-form content, rating based on cultural relevance, and explaining humor dynamics clearly. Ideal candidates are 18 to 19 years old, familiar with...Contract workFor contractorsFreelanceRemote workFlexible hours
- ...Internet Modeling is a premier adult modeling agency recruiting and hiring webcam models for high paying webcam jobs. We are one of the oldest and most experienced online modeling agencies, representing webcam models since 1998. Our agency recruits for the largest...Immediate start
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Theorem Proving Model Evaluator [Remote]. Be the first to apply!




