Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Mathematician for AI Model Evaluation

SaidGig

Shape and evaluate advanced AI systems through rigorous mathematical reasoning, problem-solving, computational work, and clear written explanations. This remote contract role focuses on creating and assessing challenging mathematics tasks across undergraduate through Ph.D.-level subject areas. Key Responsibilities

  • Design original, challenging mathematics problems that test large language model reasoning in multi-step, abstract, and proof-based settings.
  • Solve problems independently and produce detailed, logically structured solutions with clear justifications.
  • Review model-generated solutions, identify mathematical errors and missing arguments, and provide precise feedback, annotations, and corrections.
  • Help define mathematics evaluation benchmarks spanning early undergraduate through Ph.D.-level curricula.
  • Design precise, closed-ended computational prompts, write reliable Python solutions, validate numerical answers, and provide clear rationales using approved scientific libraries.
  • Complete theorem-prover work in Lean, including translating mathematical problems and proofs into formal language and verifying that formal proofs compile correctly.
  • Break down complex mathematical concepts into clear explanations using simple language, visuals, and examples.
Qualifications
  • Strong mathematical foundation at engineering entrance-exam and graduate-program levels.
  • Research, analytical, creative, and lateral-thinking skills.
  • Ability to solve complex mathematics problems with a structured, logical approach.
  • Ability to provide constructive feedback and detailed annotations.
  • Excellent structured communication and collaboration skills for a remote environment.
  • Self-motivated, able to work independently, and able to work efficiently.
  • Desktop or laptop with a reliable internet connection.
Eligibility
  • Candidates pursuing a Master’s, Ph.D., or postdoctoral degree in Mathematics, Applied Mathematics, Statistics, or a related field are eligible and encouraged to apply.
Work Terms
  • Fully remote contract, independent-contractor assignment.
  • Choose a commitment of 20, 30, or 40 hours per week.
  • Commit at least 4 hours per day, with a minimum of 20 hours per week.
  • Maintain 4 hours of overlap with Pacific Time.
  • This contractor engagement does not include medical coverage or paid leave.
Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Mathematician for AI Model Evaluation in United States vacancy
  • $60 - $80 per hour

     ...Role Overview Join a Mathematician Expert Network that connects mathematicians with AI labs and companies to shape and evaluate cutting-edge AI in mathematics. Experts contribute domain expertise to model training and evaluation, create real-world tasks and deliverables... 
    Suggested
    Hourly pay
    Contract work
    Immediate start
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $60 - $80 per hour

     ...Role Overview Provide high-level mathematical expertise to support AI research and product development. Mathematicians in this expert network train and evaluate mathematical models, design realistic problem tasks and deliverables, and give domain-specific feedback that... 
    Suggested
    Hourly pay
    Contract work
    Immediate start
    Remote work

    SaidGig

    United States
    1 day ago
  • $50 per hour

     ...Role Overview This role focuses on improving and evaluating large language models through advanced mathematical reasoning, clear written solutions,...  ...efforts and the application of those advances into reliable AI systems for enterprise use. Key Responsibilities... 
    Suggested
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    5 days ago
  • Kake is seeking Mathematics professionals with Python proficiency to contribute to project-based AI development initiatives focused on enhancing frontier AI models. You will be responsible for identifying suitable mathematical software, developing problem solutions in... 
    Suggested
    Remote work

    Kake

    Austin, TX
    3 days ago
  • $60 - $90 per hour

     ...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors...  ...Jack Dorsey . Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract Compensation... 
    Suggested
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    Remote
    16 hours ago
  • $100 - $150 per hour

     ...data scientists who will be considered for future projects evaluating how well AI systems perform real-world data science tasks. Members of this...  ...criteria, assess AI or human-produced analyses and models, document decisions in writing, and iterate on evaluations with... 
    Hourly pay
    Immediate start
    Remote work

    SaidGig

    United States
    a month ago
  • A leading AI development firm is seeking a Mathematician (PhD) to join their remote AI training project. In this role, you will evaluate AI-generated mathematical responses, ensuring accuracy and clarity. To qualify, you must have a PhD in Mathematics/Statistics, significant... 
    Remote job
    Weekly pay
    Flexible hours

    SME Careers

    Cambridge, MA
    2 days ago
  • Cincinnatus LLC is recruiting for an Insurance Subject‑Matter Expert to join a leading AI lab's GenAI team in New York. This W-2 role involves evaluating insurance tasks and guiding model development for high‑quality underwriting judgments, with placement at the client... 
    Weekday work

    Mercor

    New York, NY
    1 day ago
  • Cincinnatus LLC is placing finance SMEs at a leading AI lab to critically evaluate AI model outputs for financial tasks. The role focuses on rigorous, rubric-based assessment of model performance and constructing finance-focused evaluation frameworks. Candidates should... 
    Weekday work

    Mercor

    San Francisco, CA
    1 day ago
  •  ...Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant... 
    Remote work
    Flexible hours

    Prolific

    Charlotte, NC
    16 hours ago
  •  ...is seeking Biology Experts and Life Science Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities... 
    Hourly pay
    Remote work
    Flexible hours

    Prolific

    Jacksonville, FL
    4 days ago
  • $136.44k - $265.11k

    We are rebuilding biotech for the AI era.When a breakthrough is delayed, the world waits...  ...structured data, and run AI agents and models directly in their workflows. Over 200,000...  ...our work here.You’ll build the datasets, evaluations, and systems that help close that gap.... 
    Work at office
    Local area
    Monday to Friday
    Shift work

    Benchling

    San Francisco, CA
    3 days ago
  • $100 per hour

     ...expertise to improve the performance of large language models on finance tasks. You will work with AI researchers to identify model weaknesses in areas...  ...focused on advanced AI systems. Key Responsibilities Evaluate LLM performance in finance areas where models... 
    Hourly pay
    Contract work
    For contractors
    Freelance
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $70 - $90 per hour

     ...Role Overview Help evaluate Neuron Kernel Interface development tasks that support the training and evaluation of advanced AI models. You will assess kernel quality, numerical correctness, CUDA-to-NKI migration fidelity, and whether implementations are well suited to... 
    Hourly pay
    Remote work

    SaidGig

    Remote
    11 days ago
  • $28 - $60 per hour

     ...Evaluate AI-generated music and lyrics across a wide range of genres, applying your knowledge of the Dutch music scene and strong editorial judgment to detailed quality standards. Key Responsibilities Assess AI-generated music and rate it against detailed quality criteria... 
    Hourly pay
    For contractors
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    13 days ago
  • $15 per hour

     ...Evaluate AI-generated music and lyrics in Malayalam and English, helping assess outputs across a broad range of genres against detailed quality standards. Key Responsibilities Compare AI-generated lyrics with published songs to identify similarities. Rate lyrics... 
    Hourly pay
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    13 days ago
  • $70 - $90 per hour

     ...Review GPU and accelerator kernel development tasks that support the training and evaluation of advanced AI models. This role focuses on assessing task quality, numerical correctness, completeness, fair performance benchmarking, appropriate scope, and whether kernels... 
    Hourly pay
    Remote work

    SaidGig

    Remote
    11 days ago
  • $70 - $80 per hour

     ...Role Overview Apply advanced drug safety expertise to help improve next-generation AI systems through high-quality evaluations, safety-report analysis, and structured feedback. This remote contractor opportunity is designed for professionals with experience authoring... 
    Hourly pay
    For contractors
    Remote work

    SaidGig

    United States
    a month ago
  •  ...Role Overview Apply research-grade expertise to help evaluate and improve AI reasoning across technical and humanities disciplines. This remote contractor role supports AI-model training through rigorous analysis, high-quality feedback, and clearly articulated academic... 
    Hourly pay
    For contractors
    Remote work

    SaidGig

    United States
    10 days ago
  • $60 - $80 per hour

     ...Help shape the training and evaluation of foundational large language models by applying real-world expertise in brand strategy, growth marketing, and campaign...  .... This role brings rigorous marketing judgment to AI tasks, model assessments, and training data for a leading... 
    Hourly pay
    Weekday work

    SaidGig

    United States
    a month ago
  • $11 - $19 per hour

     ...Evaluate AI-generated music and lyrics across a broad range of genres, applying your Bengali music expertise to help assess quality, originality, and natural expression. Key Responsibilities Compare AI-generated lyrics with published songs to identify similarities.... 
    Hourly pay
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    24 days ago
  • $100 - $150 per hour

     ...Shape how advanced AI models handle real-world legal work by applying senior-level legal judgment to the design, review, and evaluation of legal knowledge tasks. You will work closely with research and program management teams to turn practical legal expertise into rigorous... 
    Hourly pay
    Full time
    Internship
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    24 days ago
  • $400 per month

    About the Role Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments. The work focuses on realistic data engineering workflows... 

    Mercor Inc

    New York, NY
    16 hours ago
  • $20 - $60 per hour

     ...Role Overview Help train and evaluate next-generation AI systems by creating rigorous, real-world assessments that test how advanced models learn, reason, and perform. This remote contract opportunity is open to recent graduates and other researchers and writers with... 
    Hourly pay
    Contract work
    For contractors
    Remote work

    SaidGig

    United States
    13 days ago
  • We are seeking an expert to evaluate and improve our AI models through comprehensive testing and analysis. You will be responsible for designing evaluation frameworks, conducting model assessments, and providing actionable insights for model improvement. Key Responsibilities... 

    BAM VENTURES LLC

    New York, NY
    3 days ago
  • Xperteez Technology seeks a data science and ML-focused specialist to evaluate AI model outputs across statistics, ML, and quantitative problems. You will assess quality, identify errors, and provide actionable feedback to improve model capability. Responsibilities include... 

    Xperteez Technology

    New York, NY
    16 hours ago
  • Verita AI is seeking an Applied AI Researcher to work with clients on model evaluation and data strategy. You will assess model performance, identify failure modes, and design data-driven solutions, collaborating with operations and engineering to implement scalable data... 

    Verita AI

    San Francisco, CA
    4 days ago
  • Job Description - Member of Technical Staff (Language Model Evaluations) Location: San Francisco (preferred), Sydney, Melbourne, Brisbane About...  ...Analysis Artificial Analysis is the leading independent AI benchmarking company. We support labs, engineers and enterprises... 

    Artificial Analysis, Inc.

    San Francisco, CA
    1 day ago
  • Google DeepMind seeks a Senior Product Manager embedded in Gemini research and model training. You will read evaluations, analyze model outputs, and make judgment calls on quality alongside researchers, translating user needs into product priorities and feedback loops for... 

    Google DeepMind

    Mountain View, CA
    4 days ago
  • $50 - $75 per hour

    A leading tech company based in Australia is seeking an AI Model Evaluator on a contract basis. The role involves evaluating AI-generated responses, writing prompts, and providing justifications based on specific criteria. Ideal candidates will hold a Master's degree in... 
    Hourly pay
    Contract work

    Mercor

    San Francisco, CA
    16 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Mathematician for AI Model Evaluation. Be the first to apply!