Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Mathematics Expert for AI Model Evaluation

$50 per hour

SaidGig

Role Overview

Apply advanced mathematical reasoning, problem-solving, computational thinking, and clear written communication to improve and evaluate large language models. This remote contract role combines challenging mathematics, model evaluation, Python-based computational work, and formal proof tasks in Lean.

Key Responsibilities

  • Create original, challenging mathematics problems that test large language model reasoning, including multi-step, abstract, and proof-based problems.
  • Solve problems independently and produce detailed, logically structured solutions with clear justifications.
  • Review model-generated solutions, identify mathematical errors and missing arguments, and provide precise feedback, annotations, and corrections.
  • Help define evaluation benchmarks covering mathematics curricula from early undergraduate through Ph.D.-level topics.
  • Design precise, closed-ended prompts for computational projects, write reliable Python code, verify numerical answers, and provide clear rationales.
  • Develop and validate Python solutions for computational tasks using approved scientific libraries.
  • Complete Lean theorem-prover tasks, including translating mathematical problems and proofs into formal language and confirming that formal proofs compile correctly.

Qualifications

  • Strong mathematical foundation at engineering entrance-exam, graduate, or Ph.D. level.
  • Research and analytical skills, with the ability to solve complex mathematics problems through a structured, logical approach.
  • Ability to explain complex mathematical concepts clearly using simple language, visuals, and examples.
  • Ability to provide constructive feedback and detailed annotations.
  • Creative and lateral thinking skills.
  • Excellent structured communication and collaboration skills in a remote environment.
  • Self-motivated, independent work style.
  • A desktop or laptop and reliable internet connection.

Eligibility

Candidates pursuing a Master’s, Ph.D., or postdoctoral degree in Mathematics, Applied Mathematics, Statistics, or a related field are eligible and encouraged to apply.

Work Terms

  • Fully remote contractor assignment.
  • Choose a commitment of 20, 30, or 40 hours per week.
  • Commit at least 4 hours per day and maintain 4 hours of overlap with Pacific Time.
  • This independent contractor engagement does not include medical benefits or paid leave.
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Mathematics Expert for AI Model Evaluation in United States vacancy
  •  ...Role Overview Apply advanced mathematical reasoning and computational...  ...solving to improve and evaluate large language models. You will design rigorous math...  ...Accelerate frontier AI research by contributing high...  ...advanced training pipelines and expert researchers who specialize... 
    Suggested
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    26 days ago
  •  ...Prolific is seeking Biology Experts and Life Science Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia.... 
    Suggested
    Hourly pay
    Remote work
    Flexible hours

    Prolific

    Jacksonville, FL
    4 days ago
  •  ...Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant... 
    Suggested
    Remote work
    Flexible hours

    Prolific

    Charlotte, NC
    1 day ago
  •  ...Overview Apply advanced physics knowledge to help improve and evaluate large language models. You will design rigorous problems, produce clear reasoning...  ...into accessible explanations while contributing to AI research projects. Key Responsibilities Design and solve... 
    Suggested
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $60 per hour

    Prolific is seeking Biology Experts and Life Science Professionals to join their Expert Network to evaluate AI-generated science. This role allows you to work from home with...  ...pay rate of up to $60 per hour for reviewing model responses, validating technical claims, and critiquing... 
    Suggested
    Remote job
    Hourly pay
    Work from home
    Flexible hours

    Prolific

    Dallas, TX
    4 days ago
  • $60 per hour

    Prolific, located in Arizona, is seeking Biology Experts and Life Science Professionals to join their Expert Network. This role involves evaluating and training AI models with real scientific expertise. Successful candidates will review AI-generated content for accuracy... 
    Hourly pay

    Prolific

    Arizona City, AZ
    4 days ago
  • $60 per hour

    Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing... 
    Remote job
    Hourly pay
    Work from home
    Flexible hours

    Prolific

    San Jose, CA
    5 days ago
  • Cincinnatus LLC is seeking a Marketing SME to join a leading GenAI team focused on evaluating AI outputs against rubrics and strengthening brand strategy in AI training data. We require 8+ years of marketing experience with top-tier brands, plus hands-on evaluation of LLM... 
    Weekday work

    Mercor

    San Francisco, CA
    3 days ago
  • $60 per hour

    Prolific is seeking Chemistry Experts and Chemical Engineers to join their Expert Network. Participants will evaluate AI-generated chemistry through tasks that assess factual accuracy...  ..., enabling cutting-edge advancements in AI models. The position requires a strong educational... 
    Hourly pay

    Prolific

    Arizona City, AZ
    3 days ago
  •  ...Physics Specialist to contribute deep scientific expertise to AI model evaluation. You will craft and assess challenging physics problems,...  ...electromagnetism, use adversarial prompting to surface errors, and provide expert critique of AI responses while working with project #J-1880... 
    Remote job

    Dorado

    New York, NY
    4 days ago
  • $76 per hour

    Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. This setup suits researchers and mathematicians seeking flexible, outcome-driven work. Participation is project-based,... 
    Hourly pay
    Permanent employment
    Temporary work
    Part time
    10 hours per week
    Flexible hours

    Mindrift

    Brooklyn, NY
    1 day ago
  • $76 per hour

    Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. This opportunity...  ...employment. Contributors may design mathematics problems, evaluate AI solutions, and improve... 
    Hourly pay
    Permanent employment
    Temporary work
    Part time
    10 hours per week

    Mindrift

    Phoenix, AZ
    2 days ago
  • $17 - $54 per hour

     ...Music & Lyrics Expert - French | Remote AI Model Evaluation is a remote evaluation track for reviewing french generalist evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured... 
    Remote job
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    9 days ago
  • $50 per hour

     ...Overview Work on fine-tuning large language models by designing and solving physics problems...  ..., and collaborating with researchers to evaluate and improve model reasoning. This role...  ...coursework, and the opportunity to learn AI-assisted analysis techniques. Key Responsibilities... 
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    1 day ago
  • Rise Data Labs is seeking advanced Mathematics and Statistics experts to support the training and evaluation of state-of-the-art AI systems. We need subject-matter experts who can apply deep quantitative knowledge to AI evaluation problems, assess AI-generated reasoning... 
    Remote job
    Contract work
    Immediate start
    Flexible hours

    BAM Ventures

    New York, NY
    5 days ago
  • $80 - $90 per hour

     ...physics expertise to help train next-generation AI systems through accurate, scientifically...  .... Key Responsibilities Analyze and evaluate real-world physics data, concepts, and...  ...scientific feedback on complex questions and model outputs to support technical accuracy and... 
    Hourly pay
    Contract work
    For contractors
    Remote work

    SaidGig

    Indiana
    a month ago
  • $60 - $80 per hour

     ...operations expertise to help develop advanced generative AI models. You will create and evaluate retail-focused tasks, bringing practical judgment and...  ...rubrics for retail tasks. Work with fellow subject-matter experts to maintain consistent, accurate training data.... 
    Hourly pay
    Weekday work

    SaidGig

    United States
    11 days ago
  • $140 per hour

     ...matter expertise to improve how next generation AI systems learn, reason, and perform. As an AI Domain Expert, you will review AI outputs, craft realistic...  ...clear written feedback that helps train and evaluate frontier models. This is a part time contractor role, fully... 
    Hourly pay
    Contract work
    Part time
    For contractors
    Local area
    Remote work

    SaidGig

    Indiana
    19 days ago
  • Kake is seeking Mathematics professionals with Python proficiency to contribute to project-based AI development initiatives focused on enhancing frontier AI models. You will be responsible for identifying suitable mathematical software, developing problem solutions in Python... 
    Remote work

    Kake

    Austin, TX
    4 days ago
  • $60 - $90 per hour

     ...creative and technical talent with leading AI research labs. Headquartered in San...  ...Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract...  .... Collaborate with researchers and experts to maintain task consistency and rigor.... 
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    Remote
    1 day ago
  • $208k - $300k

     ...Machine Learning Engineer - Model Evaluations, Public Sector The Public Sector ML team at Scale deploys advanced AI systems—including LLMs, agentic models, and multimodal...  ...with operations teams and subject matter experts to produce high-quality evaluation datasets... 
    Full time

    Scale Ai

    Washington DC
    1 day ago
  • $84 per hour

     ...thin-film processes to generate, structure, and evaluate scientific data used to train and evaluate advanced AI models for semiconductor and physical-science applications...  ...technical accuracy. Design and solve expert-level problems in ALD and semiconductor processing... 
    Hourly pay
    Remote work
    10 hours per week

    SaidGig

    United States
    22 days ago
  • $224k - $356.5k

     ...people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts...  ...computing. As a Senior / Principal Deep Learning Engineer — Model Evaluation & AI Systems, you will play a meaningful role in crafting the... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • Cincinnatus LLC is placing finance SMEs at a leading AI lab to critically evaluate AI model outputs for financial tasks. The role focuses on rigorous, rubric-based assessment of model performance and constructing finance-focused evaluation frameworks. Candidates should... 
    Weekday work

    Mercor Inc

    San Francisco, CA
    2 days ago
  • Cincinnatus LLC is recruiting for an Insurance Subject‑Matter Expert to join a leading AI lab's GenAI team in New York. This W-2 role involves evaluating insurance tasks and guiding model development for high‑quality underwriting judgments, with placement at the client... 
    Weekday work

    Mercor

    New York, NY
    2 days ago
  • $40 per hour

    A technology company in Massachusetts is seeking an R&D Biologist to join their team to train AI models by evaluating chatbot outputs against complex biology questions. Ideal candidates will hold advanced qualifications in biology or biochemistry. This position allows full... 
    Hourly pay
    Full time
    Part time
    Remote work

    DataAnnotation

    Oklahoma City, OK
    1 day ago
  •  ...Role Overview Work with a leading AI lab to evaluate outputs from generative music models in German and English. This role focuses on listening, scoring, and annotating AI-generated music and lyrics across genres, using music production and audio engineering vocabulary... 
    Hourly pay
    Part time
    Immediate start
    Remote work
    10 hours per week

    SaidGig

    United States
    a month ago
  • $100 per hour

     ...the performance of large language models on finance tasks. You will work with AI researchers to identify model weaknesses...  .... Key Responsibilities Evaluate LLM performance in finance areas...  ...AI researchers and other finance experts to influence training approaches,... 
    Hourly pay
    Contract work
    For contractors
    Freelance
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  •  ...Remote Mathematics Expert (AI/LLM) - 34877 Remote Mathematics Expert (AI/LLM) - 34877 1 week ago Be...  ...s leading AI labs to advance frontier model capabilities in thinking, reasoning, coding...  ...to align problem types with evaluation goals, particularly in areas where models... 
    Hourly pay
    Contract work
    Part time
    For contractors
    Freelance
    Internship
    Remote work
    Work from home
    Worldwide
    Afternoon shift

    Turing Inc

    New York, NY
    4 days ago
  • $65 - $90 per hour

     ...deep financial judgment to help improve foundational AI models. In this role, you will create and evaluate finance-focused work that strengthens AI systems''...  ...closely with research, engineering, and subject-matter experts. Key Responsibilities Help research and... 
    Hourly pay
    Weekday work

    SaidGig

    United States
    a month ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Mathematics Expert for AI Model Evaluation. Be the first to apply!