Lean Theorem Model Evaluator [Remote]
AuraOne Human Data
- Remote job
Lean Theorem Model Evaluator is a remote review track for evaluating AI outputs across lean theorem model research review reasoning, calculations, and research workflows. Reviewers grade derivations and assumptions, reproduce key results, and document the correct method so the modeling team can train on it.
Why this role matters
Lean Theorem Model research review models live or die on whether their derivations actually hold up under scrutiny. AuraOne uses scientific specialists to grade outputs the way a peer reviewer would — checking assumptions, reproducing key steps, and capturing the right method alongside the wrong one.
Responsibilities
- Review AI outputs against current lean theorem model research review methods, conventions, and prior work for Lean Theorem Model Evaluator assignments.
- Reproduce or sanity-check key derivations, calculations, or experimental claims.
- Flag dimensional, methodological, and citation errors with structured severity tags.
- Capture the corrected reasoning or worked example so the modeling team can train on it.
- Adjudicate disputed answers against textbooks, papers, or community standards.
- Maintain reviewer-quality scores in inter-rater calibration cycles.
Qualifications
- Graduate-level training or equivalent applied experience in lean theorem model research review or a closely related field for Lean Theorem Model Evaluator work.
- Hands-on experience publishing, teaching, or advising on the topic at a professional level.
- Comfort applying multi-page rubrics consistently across long batches.
- Clear written reasoning that cites methods, papers, or worked examples.
- Reliable async availability for at least 10 hours per week.
Example tasks
- Reproduce a lean theorem model research review derivation from a model output and flag any algebraic or dimensional errors.
- Grade a model's literature summary against the cited papers and rate the citation quality.
- Adjudicate a disputed answer between two reviewers using textbook methods.
- Audit a 25-row batch for rubric consistency and report drift to the program lead.
Nice to have
- PhD, postdoc, or industry research experience in the topic area.
- Prior work reviewing AI-assisted research tooling and its failure modes.
- Multilingual fluency for non-English papers and corpora.
Skills
- Scientific reasoning
- Method validation
- Citation review
- Quantitative analysis
- Lean Theorem Model research review
- Formal reasoning
- Proof review
- Lean
- Theorem
Work model
Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.
Compensation
Hourly rate confirmed after the interview process.
Application process
Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.
$20 per hour
...public sources and external tools. Generate high-quality human evaluation data by identifying response strengths, areas for improvement,... ...quality, clarity, tone, and completeness of responses. Ensure model responses align with expected conversational behavior and system...SuggestedRemote jobContract workPart timeSummer work$65 per hour
Prolific is seeking registered nurses in Las Vegas, NV, to assist in training AI models. You will review AI responses, rate their accuracy and safety, and write feedback to enhance AI learning. Candidates must be verified registered nurses, have recent clinical experience...SuggestedHourly paySelf employmentWork from homeFlexible hours- ...Visual Question Answering Model Evaluator is a remote evaluation track for reviewing visual question answering model evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured...SuggestedRemote jobHourly payFor contractors10 hours per week
- ...Medical Document OCR Model Evaluator is a remote clinical-review track for evaluating AI outputs that touch clinical review. Reviewers grade differential reasoning, dosing logic, and guideline adherence; flag patient-safety issues; and document the corrected clinical reasoning...SuggestedRemote jobHourly payFor contractors10 hours per week
- ...Refusal Preference Reward Model Evaluator is a remote red-team track for stress-testing AI systems against adversarial prompts. Reviewers craft attack scenarios, document the failure mode, and pair each successful jailbreak with the rubric clause it violated so the safety...SuggestedRemote jobHourly payFor contractors10 hours per week
- ...Policy Preference Reward Model Evaluator is a remote review track for evaluating AI outputs in policy review workflows. Reviewers grade citation accuracy, statutory reasoning, and policy adherence; flag risk; and document the corrected analysis so the modeling team can...Remote jobHourly payFor contractors10 hours per week
- ...Quant Finance Reasoning Model Evaluator is a remote review track for evaluating AI outputs across finance and risk workflows. Reviewers grade calculations, narrative reasoning, and policy adherence; flag compliance and reconciliation issues; and document the correct treatment...Remote jobHourly payFor contractorsWork experience placement10 hours per week
- ...computational problem solving to improve and evaluate large language models. You will design rigorous math... ..., and the use of formal theorem proving to push the reasoning capabilities... ...Work on theorem prover tasks using Lean, translating problems and proofs into...Contract workFor contractorsFreelanceRemote work
$228.7k - $343.1k
...for financial crime at enormous scale, and one bad model can mean millions in credit losses, suspicious activity... ...applies to AI. We build the tooling that lets a lean team validate at scale, so you critically evaluate what it produces and own the evaluation that confirms...Remote jobFull timeLocal areaShift work- ...demanding AI workloads. We provide high-performance GPU compute and Model API services, enabling AI companies, research labs, and... ...helping shape our go-to-market strategy. As a core member of a lean startup team, you’ll combine frontline sales execution with strategic...Remote jobFull timeFlexible hours
$110.6k - $178k
Principal Model Based Design Engineer - HybridThis is a hybrid position that requires onsite work three per week at our Towson, MD campus... ...a wealth of state-of-the-art learning resources, including our Lean Academy and online university (where you can get certificates...Full timeLocal area- ...The Mental Health Evaluator (MHE) is a critical role dedicated to delivering comprehensive, clinically appropriate, culturally competent, and trauma-informed patient behavioral health interventions in the hospital Emergency Department. Responsibilities include conducting...
- ...Engineering problems. ~ This is an opportunity for experienced Model Based System Engineers who are highly motivated, innovative and... ...creativity, and energy to every challenge. Join us at Wingbrace and lean into the future. Wingbrace LLC is an equal opportunity...Full timeWork at officeRemote work
- ...work.Work You’ll DoAs a Project - Program Analytics & Insights Evaluator IIon the project, you will be responsible for:Supporting the design... ...programsDeveloping and refining evaluation frameworks, logic models, research questions, and evaluation plans aligned with program...Local area
- ...Risk assessment experience required• Excellent written communication skills• Thorough knowledge of psychopathology and its treatment, models of behavior change and management, the psychotherapeutic process, and models of personality development.• Knowledge of and...Work experience placementWork at officeLocal areaRemote work
- Overview At PAM Health, we care for chronically and critically ill patients who require extended hospital care. PAM Health has over 80 hospital locations and employs over 11,000 people across the country. Our teams work together to deliver the highest level of compassionate...Full timeLocal area
- ...Job Title: Corporate Vice President – Model Validation and AI Governance Location: New York, NY (hybrid 3 days in office) Job... ...party solutions, independently challenging model methodologies, evaluation approaches, controls, and monitoring strategies, with a particular...Work at office
$50 per hour
...Role Overview This role focuses on improving and evaluating large language models through advanced mathematical reasoning, clear written... ...libraries and verifying numerical answers. Work on theorem-prover tasks in Lean, translating problems and proofs into formal...Contract workFor contractorsFreelanceRemote work$110.6k - $178k
Lead Engineer--Model-Based Design, Systems Verification & ValidationTowson, MDCome make... ...decomposition, interface definition, architecture evaluation, and early design verification.Define and... ...-art learning resources, including our Lean Academy and online university (where you...Full timeLocal area- Job Description Our client in MI is looking for Model-N admins . experience in integrating with Model N and configuring Model N. Additional InformationFull time
$79.4k - $119.1k
...This position represents Auto Spec Control department in a development team environment as Project Lead on a mix of complex Full Model Change (FMC) and Minor Model Change (MMC) developments. Creates, promotes, and manages critical milestones for successful package delivery...Full timeTemporary workWork experience placementWork at officeRemote workRelocation package$16 - $20 per hour
...Research Evaluator Position The research team at Penn State Ross and Carol Nese College of Nursing is hiring part-time research evaluators for projects focused on dementia care in assisted living settings. The research evaluator will assist with in-person recruitment...Hourly payPart timeSummer work$15 per hour
As a Red Bull Student Marketeer, you are part of the most dynamic and empowered brand and product ambassador program in the world. Reporting to the local Field Marketing Specialist (FMS), you will learn Red Bull’s target group and are responsible for driving the brand ...Hourly payPart timeLocal areaFlexible hoursAfternoon shift- ...In conjuncion with Talent Model Recruiters, we are seeking new and experienced models for the apparel and fashion industry to display clothing and merchandise in commercials, advertisements, and/or fashion shows. Promote products and services in online ads, social media...Part timeFlexible hours
- ...Obsidian is hiring expert Evaluators in real estate, hospitality, and events to review AI-generated work products for accuracy and quality. This is a remote position requiring strong expertise in the relevant domains. Ideal candidates should have over 5 years of professional...Work at officeRemote work
$2,300 per month
...AdvanceCare Family Model Professional If you love helping people, you'll love this job! AdvanceCare FMPs offer a family environment for clients who are unable to stay at home or choose not to go to a facility. We need you to provide a safe, home environment for our...Live inRemote work$182.4k - $273.6k
..., and technology leaders to deliver scalable, production ready models and AI driven decision systems that support complex risks, bespoke... ...the portfolio. Reinforce shared expectations for quality, evaluation rigor, and production readiness.Provide portfolio-level technical...Temporary workWork at officeRemote workShift work3 days per week$122.3k - $183.5k
...reliance and success.About the RoleAre you passionate about sharing leading practices and helping others avoid pitfalls? The Support Model & Governance Workday Success Plan Delivery team works directly with customers through strategic advisory consulting services,...Full timeWork at officeRemote workWork from homeHome officeFlexible hours$135k - $165k
...OverviewWintrust Corporate Risk Management is seeking a highly motivated Model Risk Vice President to join our Model Risk Management (MRM)... ..., k-means clustering, etc.). Familiarity with performance evaluation metrics and techniques for LLMs.PhD or Master’s in Mathematics,...Full timeTemporary workRemote workFlexible hours$115k - $200k
Role Summary/Purpose: The AVP, Acquisition Fraud Strategy and Model Monitoring, is a multi-functional role within credit fraud acquisitions... ...expected. Additional responsibilities include supporting the evaluation of new fraud models, fraud and technology tools, coordinating...Full timeWork experience placementWork from homeVisa sponsorshipWork visaMonday to Friday
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lean Theorem Model Evaluator [Remote]. Be the first to apply!







