Mathematics Expert for AI Model Evaluation
$50 per hourSaidGig
Role Overview
Apply advanced mathematical reasoning, problem-solving, computational thinking, and clear written communication to improve and evaluate large language models. This remote contract role combines challenging mathematics, model evaluation, Python-based computational work, and formal proof tasks in Lean.
Key Responsibilities
- Create original, challenging mathematics problems that test large language model reasoning, including multi-step, abstract, and proof-based problems.
- Solve problems independently and produce detailed, logically structured solutions with clear justifications.
- Review model-generated solutions, identify mathematical errors and missing arguments, and provide precise feedback, annotations, and corrections.
- Help define evaluation benchmarks covering mathematics curricula from early undergraduate through Ph.D.-level topics.
- Design precise, closed-ended prompts for computational projects, write reliable Python code, verify numerical answers, and provide clear rationales.
- Develop and validate Python solutions for computational tasks using approved scientific libraries.
- Complete Lean theorem-prover tasks, including translating mathematical problems and proofs into formal language and confirming that formal proofs compile correctly.
Qualifications
- Strong mathematical foundation at engineering entrance-exam, graduate, or Ph.D. level.
- Research and analytical skills, with the ability to solve complex mathematics problems through a structured, logical approach.
- Ability to explain complex mathematical concepts clearly using simple language, visuals, and examples.
- Ability to provide constructive feedback and detailed annotations.
- Creative and lateral thinking skills.
- Excellent structured communication and collaboration skills in a remote environment.
- Self-motivated, independent work style.
- A desktop or laptop and reliable internet connection.
Eligibility
Candidates pursuing a Master’s, Ph.D., or postdoctoral degree in Mathematics, Applied Mathematics, Statistics, or a related field are eligible and encouraged to apply.
Work Terms
- Fully remote contractor assignment.
- Choose a commitment of 20, 30, or 40 hours per week.
- Commit at least 4 hours per day and maintain 4 hours of overlap with Pacific Time.
- This independent contractor engagement does not include medical benefits or paid leave.
- ...Role Overview Apply advanced mathematical reasoning and computational... ...solving to improve and evaluate large language models. You will design rigorous math... ...Accelerate frontier AI research by contributing high... ...advanced training pipelines and expert researchers who specialize...SuggestedContract workFor contractorsFreelanceRemote work
- ...Prolific is seeking Biology Experts and Life Science Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia....SuggestedHourly payRemote workFlexible hours
- ...Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant...SuggestedRemote workFlexible hours
$60 per hour
Prolific is seeking Biology Experts and Life Science Professionals to join their Expert Network to evaluate AI-generated science. This role allows you to work from home with... ...pay rate of up to $60 per hour for reviewing model responses, validating technical claims, and critiquing...SuggestedRemote jobHourly payWork from homeFlexible hours$60 per hour
Prolific, located in Arizona, is seeking Biology Experts and Life Science Professionals to join their Expert Network. This role involves evaluating and training AI models with real scientific expertise. Successful candidates will review AI-generated content for accuracy...SuggestedHourly pay- Cincinnatus LLC is seeking a Marketing SME to join a leading GenAI team focused on evaluating AI outputs against rubrics and strengthening brand strategy in AI training data. We require 8+ years of marketing experience with top-tier brands, plus hands-on evaluation of LLM...Weekday work
$60 per hour
Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing...Remote jobHourly payWork from homeFlexible hours$60 per hour
Prolific is seeking Chemistry Experts and Chemical Engineers to join their Expert Network. Participants will evaluate AI-generated chemistry through tasks that assess factual accuracy... ..., enabling cutting-edge advancements in AI models. The position requires a strong educational...Hourly pay- ...Physics Specialist to contribute deep scientific expertise to AI model evaluation. You will craft and assess challenging physics problems,... ...electromagnetism, use adversarial prompting to surface errors, and provide expert critique of AI responses while working with project #J-1880...Remote job
$76 per hour
Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. This setup suits researchers and mathematicians seeking flexible, outcome-driven work. Participation is project-based,...Hourly payPermanent employmentTemporary workPart time10 hours per weekFlexible hours$76 per hour
Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. This opportunity... ...employment. Contributors may design mathematics problems, evaluate AI solutions, and improve...Hourly payPermanent employmentTemporary workPart time10 hours per week$17 - $54 per hour
...Music & Lyrics Expert - French | Remote AI Model Evaluation is a remote evaluation track for reviewing french generalist evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured...Remote jobFor contractors10 hours per week$50 per hour
...Overview Work on fine-tuning large language models by designing and solving physics problems... ..., and collaborating with researchers to evaluate and improve model reasoning. This role... ...coursework, and the opportunity to learn AI-assisted analysis techniques. Key Responsibilities...Contract workFor contractorsFreelanceRemote work- Rise Data Labs is seeking advanced Mathematics and Statistics experts to support the training and evaluation of state-of-the-art AI systems. We need subject-matter experts who can apply deep quantitative knowledge to AI evaluation problems, assess AI-generated reasoning...Remote jobContract workImmediate startFlexible hours
$140 per hour
...matter expertise to improve how next generation AI systems learn, reason, and perform. As an AI Domain Expert, you will review AI outputs, craft realistic... ...clear written feedback that helps train and evaluate frontier models. This is a part time contractor role, fully...Hourly payContract workPart timeFor contractorsLocal areaRemote work$60 - $80 per hour
...operations expertise to help develop advanced generative AI models. You will create and evaluate retail-focused tasks, bringing practical judgment and... ...rubrics for retail tasks. Work with fellow subject-matter experts to maintain consistent, accurate training data....Hourly payWeekday work- Kake is seeking Mathematics professionals with Python proficiency to contribute to project-based AI development initiatives focused on enhancing frontier AI models. You will be responsible for identifying suitable mathematical software, developing problem solutions in Python...Remote work
$208k - $300k
...Machine Learning Engineer - Model Evaluations, Public Sector The Public Sector ML team at Scale deploys advanced AI systems—including LLMs, agentic models, and multimodal... ...with operations teams and subject matter experts to produce high-quality evaluation datasets...Full time$60 - $90 per hour
...creative and technical talent with leading AI research labs. Headquartered in San... ...Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract... .... Collaborate with researchers and experts to maintain task consistency and rigor....Full timeContract workSummer workRemote work$84 per hour
...thin-film processes to generate, structure, and evaluate scientific data used to train and evaluate advanced AI models for semiconductor and physical-science applications... ...technical accuracy. Design and solve expert-level problems in ALD and semiconductor processing...Hourly payRemote work10 hours per week$224k - $356.5k
...people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts... ...computing. As a Senior / Principal Deep Learning Engineer — Model Evaluation & AI Systems, you will play a meaningful role in crafting the...Full time- Cincinnatus LLC is recruiting for an Insurance Subject‑Matter Expert to join a leading AI lab's GenAI team in New York. This W-2 role involves evaluating insurance tasks and guiding model development for high‑quality underwriting judgments, with placement at the client...Weekday work
- Cincinnatus LLC is placing finance SMEs at a leading AI lab to critically evaluate AI model outputs for financial tasks. The role focuses on rigorous, rubric-based assessment of model performance and constructing finance-focused evaluation frameworks. Candidates should...Weekday work
$40 per hour
A technology company in Massachusetts is seeking an R&D Biologist to join their team to train AI models by evaluating chatbot outputs against complex biology questions. Ideal candidates will hold advanced qualifications in biology or biochemistry. This position allows full...Hourly payFull timePart timeRemote work$11 - $19 per hour
...Evaluate AI-generated music and lyrics across a broad range of genres, applying your Bengali music expertise to help assess quality, originality, and natural expression. Key Responsibilities Compare AI-generated lyrics with published songs to identify similarities....Hourly payImmediate startRemote workFlexible hours$70 - $110 per hour
...Help advance frontier AI systems by bringing rigorous materials... ...and engineering judgment to the evaluation, design, and improvement of... ...looks like in practice and ensure model outputs can withstand... ...skills and tools. Translate expert, tacit judgment into clear, teachable...Hourly payFull timeLive inRelocationRelocation package$60 - $80 per hour
...Help shape the training and evaluation of foundational large language models by applying real-world expertise in brand... ...brings rigorous marketing judgment to AI tasks, model assessments, and... ...Work with other subject-matter experts to maintain consistent, accurate training...Hourly payWeekday work- ...Role Overview Work with a leading AI lab to evaluate outputs from generative music models in German and English. This role focuses on listening, scoring, and annotating AI-generated music and lyrics across genres, using music production and audio engineering vocabulary...Hourly payPart timeImmediate startRemote work10 hours per week
$100 per hour
...the performance of large language models on finance tasks. You will work with AI researchers to identify model weaknesses... .... Key Responsibilities Evaluate LLM performance in finance areas... ...AI researchers and other finance experts to influence training approaches,...Hourly payContract workFor contractorsFreelanceRemote work10 hours per weekFlexible hours$70 - $110 per hour
...Role Overview Help a leading AI research team improve how advanced AI models reason about real clinical work. In this hybrid, full-time role, you will... ...define high-quality clinical tasks, model answers, and evaluation standards alongside research and program management...Hourly payFull timeFreelanceLive inRelocationRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Mathematics Expert for AI Model Evaluation. Be the first to apply!
- expert systems engineer United States
- fruit expert United States
- technology expert United States
- expert data analyst United States
- subject matter expert work from home United States
- subject matter expert United States
- fulfillment expert United States
- guest service support expert United States
- subject matter expert senior United States
- sql expert United States




