Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Mathematician for AI Model Evaluation

SaidGig

Role Overview

Apply advanced mathematical reasoning and computational problem solving to improve and evaluate large language models. You will design rigorous math problems, produce clear, logically structured solutions, verify numerical results with code, and review model outputs to identify errors and missing arguments. This role blends formal mathematical thinking, practical Python implementation, and the use of formal theorem proving to push the reasoning capabilities of state of the art models.

How this work supports customers
  • Accelerate frontier AI research by contributing high quality data and evaluation material, and by working with advanced training pipelines and expert researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents.
  • Help enterprises move AI from proof of concept to reliable, proprietary intelligence by producing evaluation artifacts and methodologies that measure real world performance and deliver measurable impact on business outcomes.
Key Responsibilities
  • Create original, challenging mathematics problems that probe multi step reasoning, abstraction, and proof skills of large language models.
  • Solve problems independently and write detailed, logically ordered solutions with clear justifications and steps.
  • Review and annotate model generated solutions, identify mathematical mistakes or omitted arguments, and provide precise corrections and feedback.
  • Help define new evaluation benchmarks based on mathematics curricula spanning early undergraduate through PhD level topics.
  • Design exact, closed ended prompts and produce reliable Python solutions for computational tasks using approved scientific libraries, verifying numerical answers.
  • Work on theorem prover tasks using Lean, translating problems and proofs into formal language and confirming formal proofs compile correctly.
  • Contribute to improving your own workflows by learning to leverage AI tools to be a more effective analyst.
Qualifications
  • Strong foundation in mathematics at the level expected for engineering entrance exams and for graduate or PhD level programs.
  • Proven ability to break down complex mathematical concepts into clear, simple explanations, using visuals or examples where helpful.
  • Experience writing reliable Python code for numerical or symbolic computation, and familiarity with scientific libraries.
  • Experience or willingness to work with the Lean theorem prover for formalization and verification of proofs.
  • Excellent structured written communication, ability to provide constructive feedback and detailed annotations.
  • Creative and lateral thinking, strong research and analytical skills, and the ability to work independently in a remote setting.
  • Technical requirements: access to a desktop or laptop with a stable internet connection.
Preferred Qualifications
  • Pursuing or holding a Master’s, PhD, or Postdoctoral degree in Mathematics, Applied Mathematics, Statistics, or a closely related field. Candidates in these programs are eligible and encouraged to apply.
  • Demonstrated ability to analyze and solve complex math problems using structured, logical approaches.
  • Facility explaining math concepts in plain language, with examples and visuals when appropriate.
Work Terms
  • Remote contractor assignment, freelance engagement. This role is contractor status only, there is no medical or paid leave provided.
  • Time commitment options: 20 hours per week, 30 hours per week, or 40 hours per week.
  • Minimum required commitment: at least 4 hours per day and a minimum of 20 hours per week, with 4 hours of overlap with Pacific Standard Time, to enable collaboration.
  • Potential for contract extension based on performance and project needs.
Compensation

Compensation details are not specified in the listing.

Eligibility
  • Applicants pursuing or holding graduate level degrees in mathematics related fields are eligible and encouraged to apply. No additional work authorization details were provided in the source.
Application

Candidates who meet the qualifications, especially those enrolled in or holding Master’s, PhD, or Postdoctoral degrees in relevant fields, are encouraged to apply. The listing does not provide further application steps or platform specific instructions.

Vacancy posted 20 days ago
Similar jobs that could be interesting for youBased on the Mathematician for AI Model Evaluation in United States vacancy
  • $50 per hour

     ...Role Overview This role focuses on improving and evaluating large language models through advanced mathematical reasoning, clear written solutions,...  ...efforts and the application of those advances into reliable AI systems for enterprise use. Key Responsibilities... 
    Suggested
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    5 days ago
  • A technology firm focused on AI training is seeking a Biostatistician to enhance AI models' performance. This remote role involves giving AI chatbots complex mathematical challenges and evaluating their outputs for quality and correctness. Ideal candidates should possess... 
    Suggested
    Hourly pay
    Remote work

    DataAnnotation

    Lincoln, NE
    5 days ago
  • $40 per hour

    A technology firm specializing in AI is seeking a Biostatistician to enhance AI models by evaluating their performance and solving complex mathematical problems. This role offers flexibility as a remote position, allowing you to choose your projects and work according... 
    Suggested
    Hourly pay
    Remote work

    DataAnnotation

    Jackson, MS
    5 days ago
  • $75 per hour

     ...apply geospatial imaging, survey, and photogrammetric expertise to evaluate AI-generated maps and geospatial content, verify spatial accuracy, and provide clear, structured feedback that improves model outputs. No prior AI experience is required. Key Responsibilities... 
    Suggested
    Hourly pay
    Temporary work
    Part time
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $80 - $90 per hour

     ...that will be used to train next-generation AI systems. This contract role is remote and...  ...analysis, and clear explanations that improve model learning and reasoning. No prior...  ...ensuring clarity, rigor, and correctness. Evaluate and refine mathematical content for inclusion... 
    Suggested
    Hourly pay
    Contract work
    For contractors
    Local area
    Remote work
    Worldwide

    SaidGig

    Indiana
    10 days ago
  • $20 - $40 per hour

     ...annotated data and feedback that will help train next-generation AI systems. This is a remote, project-based contractor role...  ...ability, and technical communication will directly shape model learning and evaluation. No prior AI experience is required. Key... 
    Hourly pay
    For contractors
    Remote work

    SaidGig

    Indiana
    a month ago
  •  ...Role Overview Join a Mathematician Expert Network that connects mathematicians with AI labs and companies to shape and evaluate cutting-edge AI in mathematics. Experts contribute domain expertise to model training and evaluation, create real-world tasks and deliverables... 
    Hourly pay
    Contract work
    Immediate start
    Remote work

    SaidGig

    United States
    more than 2 months ago
  •  ...computational problems that test whether advanced AI systems can perform real scientific...  ...task will be exercised against top AI models and iteratively refined until it matches...  ...solution validators in Python. Run and evaluate problems against state-of-the-art AI models... 
    Hourly pay
    Remote work

    SaidGig

    United States
    6 days ago
  • A data technology company is seeking a Statistician to enhance AI models by evaluating their logic and progress. The role demands expert-level mathematical reasoning and offers a flexible working schedule, allowing for full-time or part-time remote work. Responsibilities... 
    Hourly pay
    Full time
    Part time
    Remote work
    Flexible hours

    DataAnnotation

    Raleigh, NC
    5 days ago
  • Kake is seeking Mathematics professionals with Python proficiency to contribute to project-based AI development initiatives focused on enhancing frontier AI models. You will be responsible for identifying suitable mathematical software, developing problem solutions in... 
    Remote work

    Kake

    Austin, TX
    3 days ago
  • $40 per hour

    A leading data annotation company seeks a Statistician to enhance AI models through evaluating chatbot logic and solving complex mathematical problems. This remote position allows flexibility in projects and hours, with compensation starting at $40+ per hour. Ideal candidates... 
    Hourly pay
    Contract work
    Remote work

    DataAnnotation

    Bismarck, ND
    5 days ago
  • $80 - $135 per hour

     ...09.26574v3), a frontier research-level physics benchmark. The role produces fully human-verified reference data used to evaluate large language model performance on frontier physics reasoning. Work includes solving CritPt research-level problems end-to-end, auditing other... 
    Hourly pay
    Remote work
    10 hours per week

    SaidGig

    United States
    22 days ago
  •  ...Role Overview Lead the creation and evaluation of AI systems that perform advanced economic reasoning. This role focuses on building rigorous...  ..., benchmarks, and agentic workflows to measure and improve model performance on real world economic research and policy... 
    Full time
    Remote work

    SaidGig

    United States
    15 days ago
  • $40 per hour

     ...A technology-driven company is seeking a Statistician to support AI model development. In this flexible role, you will evaluate AI logic and solve complex mathematical problems, ensuring quality outputs. Candidates should possess strong mathematical reasoning and proficiency... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Boston, MA
    1 day ago
  • $75 per hour

     ...Role Overview Math PhDs and doctoral-level mathematicians apply advanced mathematical training to evaluate AI-generated mathematical content, design rigorous domain-relevant questions, and provide detailed feedback to improve mathematical reasoning, proof construction... 
    Hourly pay
    Full time
    Contract work
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    20 days ago
  • A leading AI development firm is seeking a Mathematician (PhD) to join their remote AI training project. In this role, you will evaluate AI-generated mathematical responses, ensuring accuracy and clarity. To qualify, you must have a PhD in Mathematics/Statistics, significant... 
    Weekly pay
    Remote work
    Flexible hours

    SME Careers

    Cambridge, MA
    1 day ago
  • $208k - $300k

     ...Machine Learning Engineer - Model Evaluations, Public Sector The Public Sector ML team at Scale deploys advanced AI systems—including LLMs, agentic models, and multimodal pipelines—into mission-critical government environments. We build evaluation frameworks that ensure... 
    Full time

    Scale Ai

    Washington DC
    10 hours ago
  • $80 per hour

     ...Role Overview Evaluate and improve frontier AI coding agents by applying professional data engineering judgment to realistic data infrastructure and pipeline scenarios. You will use and assess model-generated implementations for ETL, data warehouses, analytics platforms... 
    Hourly pay
    Remote work

    SaidGig

    United States
    19 days ago
  • $40 per hour

    A biotechnology company is seeking a Biotechnology R&D Scientist to train AI models and evaluate their outputs. This role involves measuring the progress of AI chatbots with complex biology questions and ensuring their performance and correctness. Candidates should have... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Vermont
    5 days ago
  • $224k - $356.5k

     ...people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts...  ...computing. As a Senior / Principal Deep Learning Engineer — Model Evaluation & AI Systems, you will play a meaningful role in crafting the... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant... 
    Remote work
    Flexible hours

    Prolific

    Charlotte, NC
    5 days ago
  • $20 per hour

     ...DataAnnotation is committed to creating quality AI. Join our team to help train AI chatbots while...  ...chatbots. You will develop complex prompts to test AI models, write high-quality responses to demonstrate excellence, and evaluate different model outputs based on accuracy and... 
    Hourly pay
    Full time
    Contract work
    Part time
    For contractors
    Self employment
    Freelance
    Remote work

    DataAnnotation

    Wyoming, OH
    5 days ago
  • $40 per hour

    A forward-thinking AI development firm seeks experienced quantitative professionals to evaluate AI-generated work, applying their skills in statistical analysis, predictive modeling, and technical writing. This fully remote opportunity offers a flexible schedule and projects... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Sioux Falls, SD
    5 days ago
  • $40 per hour

    A leading AI development company is seeking experienced quantitative professionals to evaluate AI-generated analyses and design quantitative problems for AI training. This fully remote role offers flexibility in project selection and scheduling, with competitive pay starting... 
    Hourly pay
    Remote work

    DataAnnotation

    Juneau, AK
    5 days ago
  • $40 per hour

     ...A leading AI development firm is looking for experienced quantitative professionals to evaluate AI-generated work and design problems for AI training. This fully remote position allows for a flexible schedule, offering competitive pay starting at $40+ per hour. Ideal candidates... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Bismarck, ND
    5 days ago
  •  ...role on a GenAI research team by applying hands-on legal practice experience to improve how frontier AI models perform real legal work. In this position you will evaluate model outputs, create high-quality instruction specifications and authoritative solutions, and... 
    Hourly pay
    Full time
    Freelance
    Internship
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    4 days ago
  • $40 per hour

    A leading AI training company is seeking a Biotechnology R&D Scientist to evaluate and improve AI models by posing complex biological questions. This remote position allows candidates to choose projects and work on their own schedule, with rates starting at $40+ per hour... 
    Hourly pay
    Remote work

    DataAnnotation

    Sioux Falls, SD
    3 days ago
  • $150 per hour

     ...Overview Aerospace engineering professionals apply their domain expertise to evaluate AI-generated outputs, assess technical content, and provide clear, structured feedback that improves models'' understanding of aerospace tasks, terminology, and practices. You will... 
    Hourly pay
    Temporary work
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $85 per hour

     ...Role Overview Environmental Analysts apply environmental assessment, GIS, and renewable energy siting expertise to evaluate AI-generated outputs and to create expert-level training data that improves AI understanding of environmental workflows and geospatial data practices... 
    Hourly pay
    Full time
    Contract work
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $40 per hour

    A leading AI development firm is seeking experienced quantitative professionals to evaluate AI-generated quantitative work and provide critical feedback. This role offers the flexibility of remote work, allowing you to set your own schedule while focusing on impactful projects... 
    Hourly pay
    Remote work

    DataAnnotation

    Wisconsin
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Mathematician for AI Model Evaluation. Be the first to apply!