Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Mathematician for AI Model Evaluation

$50 per hour

SaidGig

Role Overview

Apply advanced mathematical reasoning to improve and evaluate large language models. You will create rigorous problems, assess model outputs, develop computational solutions, and explain complex concepts with clarity across topics ranging from early undergraduate mathematics to Ph.D.-level material.

Key Responsibilities

  • Design original, challenging mathematics problems that test multi-step reasoning, abstraction, and proof-based thinking.
  • Solve problems independently and produce detailed, logically structured solutions with clear justifications.
  • Review model-generated solutions, identify errors or missing reasoning, and provide precise feedback, annotations, and corrections.
  • Help define mathematics evaluation benchmarks spanning early undergraduate through Ph.D.-level curricula.
  • Design precise, closed-ended computational prompts; write reliable Python code; verify numerical answers; and provide clear rationales.
  • Develop and validate Python-based solutions using approved scientific libraries.
  • Work on Lean theorem-prover tasks, including translating mathematical problems and proofs into formal language and confirming that formal proofs compile correctly.

Qualifications

  • Strong mathematical foundation, including material at engineering entrance-exam and graduate or Ph.D. program levels.
  • Research, analytical, creative, and lateral-thinking skills.
  • Ability to solve complex mathematical problems using a structured, logical approach.
  • Ability to explain mathematics clearly in simple language, using visuals and examples when helpful.
  • Ability to provide constructive feedback and detailed annotations.
  • Excellent structured communication and remote collaboration skills.
  • Self-motivated, efficient, and able to work independently in a remote environment.
  • Access to a desktop or laptop and a reliable internet connection.

Eligibility

  • Candidates pursuing a Master''s, Ph.D., or postdoctoral degree in Mathematics, Applied Mathematics, Statistics, or a related field are eligible and encouraged to apply.

Work Terms

  • Fully remote independent contractor engagement.
  • Choose a commitment of 20, 30, or 40 hours per week.
  • Minimum commitment is 4 hours per day and 20 hours per week, including 4 hours of overlap with Pacific Time.
  • This engagement does not include medical benefits or paid leave.
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Mathematician for AI Model Evaluation in United States vacancy
  • $60 - $80 per hour

     ...Role Overview Join a Mathematician Expert Network that connects mathematicians with AI labs and companies to shape and evaluate cutting-edge AI in mathematics. Experts contribute domain expertise to model training and evaluation, create real-world tasks and deliverables... 
    Suggested
    Hourly pay
    Contract work
    Immediate start
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • Kake is seeking Mathematics professionals with Python proficiency to contribute to project-based AI development initiatives focused on enhancing frontier AI models. You will be responsible for identifying suitable mathematical software, developing problem solutions in... 
    Suggested
    Remote work

    Kake

    Austin, TX
    2 days ago
  • $80 - $135 per hour

     ...09.26574v3), a frontier research-level physics benchmark. The role produces fully human-verified reference data used to evaluate large language model performance on frontier physics reasoning. Work includes solving CritPt research-level problems end-to-end, auditing other... 
    Suggested
    Hourly pay
    Remote work
    10 hours per week

    SaidGig

    United States
    a month ago
  • A leading AI development firm is seeking a Mathematician (PhD) to join their remote AI training project. In this role, you will evaluate AI-generated mathematical responses, ensuring accuracy and clarity. To qualify, you must have a PhD in Mathematics/Statistics, significant... 
    Suggested
    Weekly pay
    Remote work
    Flexible hours

    SME Careers

    Cambridge, MA
    4 hours ago
  • $60 - $90 per hour

     ...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors...  ...Jack Dorsey . Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract Compensation... 
    Suggested
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    Remote
    1 day ago
  • $208k - $300k

     ...Machine Learning Engineer - Model Evaluations, Public Sector The Public Sector ML team at Scale deploys advanced AI systems—including LLMs, agentic models, and multimodal pipelines—into mission-critical government environments. We build evaluation frameworks that ensure... 
    Full time

    Scale Ai

    Washington DC
    1 day ago
  • $100 per hour

     ...Role Overview Apply deep domain expertise to train and evaluate next-generation AI systems by producing, refining, and validating high-quality,...  ...data. This part-time contractor role focuses on improving model outputs through careful content review, prompt refinement,... 
    Hourly pay
    Part time
    For contractors
    Remote work

    SaidGig

    Indiana
    21 days ago
  • $60 - $90 per hour

     ...Role Overview Help advance frontier AI research by creating rigorous, real-world data analysis evaluations for generative AI models. You will design and complete complex analytical tasks that mirror practical research work, then use your reference analyses to assess where... 
    Hourly pay
    Full time
    Freelance
    Remote work

    SaidGig

    United States
    a month ago
  • $70 per hour

     ...original, executable scientific computing challenges that help evaluate the limits of advanced AI systems. This role focuses on creating rigorous...  ...scientific computing that are difficult for current frontier AI models to solve. Source problem material from published... 
    Hourly pay
    Part time
    For contractors
    Immediate start
    Remote work

    SaidGig

    United States
    2 days ago
  • $224k - $356.5k

     ...people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts...  ...computing. As a Senior / Principal Deep Learning Engineer — Model Evaluation & AI Systems, you will play a meaningful role in crafting the... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • Cincinnatus LLC is placing finance SMEs at a leading AI lab to critically evaluate AI model outputs for financial tasks. The role focuses on rigorous, rubric-based assessment of model performance and constructing finance-focused evaluation frameworks. Candidates should... 
    Weekday work

    Mercor Inc

    San Francisco, CA
    5 days ago
  • Cincinnatus LLC is recruiting for an Insurance Subject‑Matter Expert to join a leading AI lab's GenAI team in New York. This W-2 role involves evaluating insurance tasks and guiding model development for high‑quality underwriting judgments, with placement at the client... 
    Weekday work

    Mercor

    New York, NY
    5 days ago
  • $60 per hour

     ...Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing... 
    Hourly pay
    Remote work
    Work from home
    Flexible hours

    Prolific

    San Jose, CA
    4 days ago
  • $65 - $90 per hour

     ...Apply deep financial judgment to help improve foundational AI models. In this role, you will create and evaluate finance-focused work that strengthens AI systems'' reasoning, analysis, and decision-making capabilities. Role Overview This is a W-2 employment opportunity... 
    Hourly pay
    Weekday work

    SaidGig

    United States
    a month ago
  • $13 - $54 per hour

     ...Evaluate AI-generated music and lyrics across a wide range of genres, using your knowledge of the Spanish (Mexico) music scene to assess quality against detailed standards. This remote role involves working in both Spanish (Mexico) and English. Key Responsibilities... 
    Hourly pay
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    2 days ago
  •  ...reasoning and computational problem solving to improve and evaluate large language models. You will design rigorous math problems, produce clear, logically...  ...How this work supports customers Accelerate frontier AI research by contributing high quality data and evaluation... 
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    29 days ago
  • $70 - $110 per hour

     ...Help advance frontier AI systems by bringing rigorous materials science and engineering judgment to the evaluation, design, and improvement of technical knowledge work. You will...  ...reasoning looks like in practice and ensure model outputs can withstand technical scrutiny.... 
    Hourly pay
    Full time
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    3 days ago
  • $60 - $80 per hour

     ...Help shape the training and evaluation of foundational large language models by applying real-world expertise in brand strategy, growth marketing, and campaign...  .... This role brings rigorous marketing judgment to AI tasks, model assessments, and training data for a leading... 
    Hourly pay
    Weekday work

    SaidGig

    United States
    26 days ago
  • $35 - $62 per hour

     ...Apply your Japanese music expertise to evaluate AI-generated music and lyrics across a wide range of genres. You will assess outputs against detailed quality standards in both Japanese and English. Key Responsibilities Compare AI-generated lyrics with published songs... 
    Hourly pay
    For contractors
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    13 days ago
  • $17 - $42 per hour

     ...Evaluate AI-generated music and lyrics in Hebrew and English, applying your knowledge of the Hebrew music scene and detailed quality standards across a wide range of genres. Key Responsibilities Compare AI-generated lyrics with published songs to identify similarities... 
    Hourly pay
    For contractors
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    13 days ago
  •  ...Overview Apply advanced physics knowledge to help improve and evaluate large language models. You will design rigorous problems, produce clear reasoning...  ...into accessible explanations while contributing to AI research projects. Key Responsibilities Design and solve... 
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $18 - $42 per hour

     ...Role Overview Evaluate generative music AI outputs across a wide range of genres, applying your knowledge of Portuguese-language music and lyrics to detailed quality standards. This role combines critical listening with lyric analysis in Portuguese and English. Key... 
    Hourly pay
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    2 days ago
  • $18 per hour

     ...Role Overview Evaluate AI-generated music and lyrics across a range of genres, applying detailed quality standards in both Thai and English...  ...of Thai music, language, and lyrical expression to help assess model outputs. Key Responsibilities Compare AI-generated lyrics... 
    Hourly pay
    For contractors
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    2 days ago
  • $70 - $110 per hour

     ...Role Overview Help a leading AI research team improve how advanced AI models reason about real clinical work. In this hybrid, full-time role, you will...  ...define high-quality clinical tasks, model answers, and evaluation standards alongside research and program management... 
    Hourly pay
    Full time
    Freelance
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    3 days ago
  • $400 per month

    About the Role Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments. The work focuses on realistic data engineering workflows... 

    Mercor

    New York, NY
    4 days ago
  • $20 - $60 per hour

     ...Help train next-generation AI systems by creating rigorous, real-world evaluations that test how well advanced models learn, reason, and perform. This remote contract opportunity is open to recent graduates, advanced-degree holders, and professionals from any background... 
    Hourly pay
    Contract work
    For contractors
    Remote work

    SaidGig

    Indiana
    2 days ago
  •  ...Evaluate generative music AI across a broad range of genres, applying your knowledge of Indonesian music and lyrics to detailed quality standards. You will work in both Indonesian and English to help assess lyric quality and authenticity. Key Responsibilities Compare... 
    Hourly pay
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    2 days ago
  • Cincinnatus LLC is seeking a Marketing SME to join a leading GenAI team focused on evaluating AI outputs against rubrics and strengthening brand strategy in AI training data. We require 8+ years of marketing experience with top-tier brands, plus hands-on evaluation of LLM... 
    Weekday work

    Mercor

    San Francisco, CA
    1 day ago
  • $70 - $110 per hour

     ...Role Overview Apply your legal practice experience to evaluate and improve AI-generated legal content and workflows. You will assess work grounded in litigation, legal drafting, and client advisory practice. Key Responsibilities Evaluate AI-generated legal memoranda... 
    Hourly pay
    Immediate start
    Remote work

    SaidGig

    United States
    2 days ago
  • $60 per hour

     ...in Arizona, is seeking Biology Experts and Life Science Professionals to join their Expert Network. This role involves evaluating and training AI models with real scientific expertise. Successful candidates will review AI-generated content for accuracy and assist in fact... 
    Hourly pay

    Prolific

    Arizona City, AZ
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Mathematician for AI Model Evaluation. Be the first to apply!