Theorem Proving Model Evaluator [Remote]
AuraOne Human Data
- Remote job
Theorem Proving Model Evaluator is a remote review track for evaluating AI outputs across theorem proving model research review reasoning, calculations, and research workflows. Reviewers grade derivations and assumptions, reproduce key results, and document the correct method so the modeling team can train on it.
Why this role matters
Theorem Proving Model research review models live or die on whether their derivations actually hold up under scrutiny. AuraOne uses scientific specialists to grade outputs the way a peer reviewer would — checking assumptions, reproducing key steps, and capturing the right method alongside the wrong one.
Responsibilities
- Review AI outputs against current theorem proving model research review methods, conventions, and prior work for Theorem Proving Model Evaluator assignments.
- Reproduce or sanity-check key derivations, calculations, or experimental claims.
- Flag dimensional, methodological, and citation errors with structured severity tags.
- Capture the corrected reasoning or worked example so the modeling team can train on it.
- Adjudicate disputed answers against textbooks, papers, or community standards.
- Maintain reviewer-quality scores in inter-rater calibration cycles.
Qualifications
- Graduate-level training or equivalent applied experience in theorem proving model research review or a closely related field for Theorem Proving Model Evaluator work.
- Hands-on experience publishing, teaching, or advising on the topic at a professional level.
- Comfort applying multi-page rubrics consistently across long batches.
- Clear written reasoning that cites methods, papers, or worked examples.
- Reliable async availability for at least 10 hours per week.
Example tasks
- Reproduce a theorem proving model research review derivation from a model output and flag any algebraic or dimensional errors.
- Grade a model's literature summary against the cited papers and rate the citation quality.
- Adjudicate a disputed answer between two reviewers using textbook methods.
- Audit a 25-row batch for rubric consistency and report drift to the program lead.
Nice to have
- PhD, postdoc, or industry research experience in the topic area.
- Prior work reviewing AI-assisted research tooling and its failure modes.
- Multilingual fluency for non-English papers and corpora.
Skills
- Scientific reasoning
- Method validation
- Citation review
- Quantitative analysis
- Theorem Proving Model research review
- Formal reasoning
- Proof review
- Theorem
- Proving
Work model
Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.
Compensation
Hourly rate confirmed after the interview process.
Application process
Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.
$20 per hour
...public sources and external tools. Generate high-quality human evaluation data by identifying response strengths, areas for improvement,... ...quality, clarity, tone, and completeness of responses. Ensure model responses align with expected conversational behavior and system...SuggestedRemote jobContract workPart timeSummer work$65 per hour
Prolific is seeking registered nurses in Las Vegas, NV, to assist in training AI models. You will review AI responses, rate their accuracy and safety, and write feedback to enhance AI learning. Candidates must be verified registered nurses, have recent clinical experience...SuggestedHourly paySelf employmentWork from homeFlexible hours- Turing is seeking a Software Engineering Evaluator to create cutting‑edge datasets for training, benchmarking, and advancing large language models, collaborating closely with researchers. This includes curating code examples, providing precise solutions, and making corrections...SuggestedRemote job
- ...Image-Text Grounding Model Evaluator is a remote evaluation track for reviewing image text grounding model evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the...SuggestedRemote jobHourly payFor contractors10 hours per week
- ...Refusal Preference Reward Model Evaluator is a remote red-team track for stress-testing AI systems against adversarial prompts. Reviewers craft attack scenarios, document the failure mode, and pair each successful jailbreak with the rubric clause it violated so the safety...SuggestedRemote jobHourly payFor contractors10 hours per week
- ...Quant Finance Reasoning Model Evaluator is a remote review track for evaluating AI outputs across finance and risk workflows. Reviewers grade calculations, narrative reasoning, and policy adherence; flag compliance and reconciliation issues; and document the correct treatment...Remote jobHourly payFor contractorsWork experience placement10 hours per week
- ...Medical Document OCR Model Evaluator is a remote clinical-review track for evaluating AI outputs that touch clinical review. Reviewers grade differential reasoning, dosing logic, and guideline adherence; flag patient-safety issues; and document the corrected clinical reasoning...Remote jobHourly payFor contractors10 hours per week
- ...Policy Preference Reward Model Evaluator is a remote review track for evaluating AI outputs in policy review workflows. Reviewers grade citation accuracy, statutory reasoning, and policy adherence; flag risk; and document the corrected analysis so the modeling team can...Remote jobHourly payFor contractors10 hours per week
- ...reasoning and computational problem solving to improve and evaluate large language models. You will design rigorous math problems, produce clear,... ...practical Python implementation, and the use of formal theorem proving to push the reasoning capabilities of state of the art...Contract workFor contractorsFreelanceRemote work
$228.7k - $343.1k
...financial crime at enormous scale, and one bad model can mean millions in credit losses,... ...finding what looks right and is wrong, then proving it well enough to hold up under... ...team validate at scale, so you critically evaluate what it produces and own the evaluation that...Remote jobFull timeLocal areaShift work$71.6k - $89.4k
...Bring the Future! Job Purpose Coordinates and executes New Model project parts/White body/complete vehicle shipping and... ...coordinating with each problem PICs. Support NM E-PM for New Model Evaluation preparation - Support documentation creation - Support room...Full timeTemporary workWork experience placementRemote workRelocation package- ...Team seeks a systems engineering leader to serve as an engine line Model Leader. This role provides technical and programmatic leadership... ...S. export-controlled information. Employment is contingent upon proving U.S. Person status (U.S. citizen, lawful permanent resident,...Permanent employmentFull timeRemote workRelocation package
$79.4k - $119.1k
...This position represents Auto Spec Control department in a development team environment as Project Lead on a mix of complex Full Model Change (FMC) and Minor Model Change (MMC) developments. Creates, promotes, and manages critical milestones for successful package delivery...Full timeTemporary workWork experience placementWork at officeRemote workRelocation package- ...where that happens: every product built on a model is bounded by what it costs to run, so... ...for: ~ This role builds the evaluation and decision systems that make agentic inference... ...what an evaluation result does and does not prove Ability to analyze per-request and per...Full time
- ...The Mental Health Evaluator (MHE) is a critical role dedicated to delivering comprehensive, clinically appropriate, culturally competent, and trauma-informed patient behavioral health interventions in the hospital Emergency Department. Responsibilities include conducting...
- ...Risk assessment experience required• Excellent written communication skills• Thorough knowledge of psychopathology and its treatment, models of behavior change and management, the psychotherapeutic process, and models of personality development.• Knowledge of and...Work experience placementWork at officeLocal areaRemote work
- Overview At PAM Health, we care for chronically and critically ill patients who require extended hospital care. PAM Health has over 80 hospital locations and employs over 11,000 people across the country. Our teams work together to deliver the highest level of compassionate...Full timeLocal area
- Job Description Our client in MI is looking for Model-N admins . experience in integrating with Model N and configuring Model N. Additional InformationFull time
- ...setup, and the ability to follow scripts and prompts with attention to quality. You will record expressive Italian-language audio, evaluate AI-generated speech for naturalness and pronunciation, and provide actionable feedback to improve outputs. Flexible schedule,...Remote jobHourly payContract workFlexible hours
- ...our platform every day. This role leads the next step in how our model-based pipeline scales: generation, validation, and quality gates... ...safety standards, enough to define what the platform must prove and to align with safety owners. A track record of evolving how...Local areaRemote work
- ...wholesale partners, and sells online through its website, which serves customers across the United States and globally at MEN’S FIT MODEL: This is a part-time opportunity to be a part of the design process and to influence the comfort and fit of our clothing. You will...Part timeFlexible hours
- ...you a caring, compassionate, and reliable person? Here is an opportunity to make a lasting impact on someone’s life – become a Family Model Provider ! This is a chance to help make a difference by opening your home to an individual with developmental and/or physical...Daily paid
$35 per hour
...Filly Flair is a fast paced upbeat online boutique. We are looking for part-time models who love to be in front of the camera! If you LOVE taking photos, fashion and social media this position is for you! This position will take product photos daily, lifestyle photos...Full timePart timeLive in$141.7k - $268.3k
In this position...The New Model Launch Manager holds a pivotal role in strategically overseeing and directing the successful launch of multiple new vehicle or powertrain models within manufacturing plants. The primary goal is to ensure these launches consistently meet...Immediate startFlexible hours$185k - $400k
...real-time, multimodal generation and intelligent agentic platforms. We are seeking accomplished Research Scientists in Foundation Models with expertise in pre-training and mid-training large-scale multimodal foundation models to advance our mission of making agentic,...Remote work$60k - $100k
The ETF and Model Portfolio Business Development Associate will support the ETF and Model Portfolio Specialist Team and broader business... ...role.Strong analytical skills, including the ability to evaluate market trends, sales data, product positioning, platform opportunities...Full timeTemporary workLocal areaHome officeFlexible hours- ...Turing is seeking graduate students or professionals for a remote role in evaluating AI-generated research reports. Responsibilities include reading, annotating, and scoring reports on a 1-5 scale, alongside providing written justifications. Candidates must possess strong...Remote work
- ...function secures Carnival Corp’s adoption of AI by ensuring GenAI, model, and agentic systems are deployed, validated, and operated... ...drive remediation.Design, build, and run model validation and evaluation harnesses to test AI models and pipelines for robustness, data...Full timePart timeWork at officeLocal areaWork from homeRelocationMonday to Thursday
- ...first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world... ...City, Montreal, Seoul, Germany and Paris. Join us!Why this role?Evaluation is critical to making progress in scaling intelligence. As...Full timeWork at officeLocal areaRemote workHome office
$160k - $327k
...Opportunities by outlining clear career pathways for promotions and role advancements and constructive feedback through regular evaluations.• Work-Life Balance Support• Strong and inclusive organizational culture• Access to cutting edge technology for best practices and...Temporary workRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Theorem Proving Model Evaluator [Remote]. Be the first to apply!





