Mathematician for AI Model Evaluation
$50 per hourSaidGig
Role Overview
Apply advanced mathematical reasoning to improve and evaluate large language models. You will create rigorous problems, assess model outputs, develop computational solutions, and explain complex concepts with clarity across topics ranging from early undergraduate mathematics to Ph.D.-level material.
Key Responsibilities
- Design original, challenging mathematics problems that test multi-step reasoning, abstraction, and proof-based thinking.
- Solve problems independently and produce detailed, logically structured solutions with clear justifications.
- Review model-generated solutions, identify errors or missing reasoning, and provide precise feedback, annotations, and corrections.
- Help define mathematics evaluation benchmarks spanning early undergraduate through Ph.D.-level curricula.
- Design precise, closed-ended computational prompts; write reliable Python code; verify numerical answers; and provide clear rationales.
- Develop and validate Python-based solutions using approved scientific libraries.
- Work on Lean theorem-prover tasks, including translating mathematical problems and proofs into formal language and confirming that formal proofs compile correctly.
Qualifications
- Strong mathematical foundation, including material at engineering entrance-exam and graduate or Ph.D. program levels.
- Research, analytical, creative, and lateral-thinking skills.
- Ability to solve complex mathematical problems using a structured, logical approach.
- Ability to explain mathematics clearly in simple language, using visuals and examples when helpful.
- Ability to provide constructive feedback and detailed annotations.
- Excellent structured communication and remote collaboration skills.
- Self-motivated, efficient, and able to work independently in a remote environment.
- Access to a desktop or laptop and a reliable internet connection.
Eligibility
- Candidates pursuing a Master''s, Ph.D., or postdoctoral degree in Mathematics, Applied Mathematics, Statistics, or a related field are eligible and encouraged to apply.
Work Terms
- Fully remote independent contractor engagement.
- Choose a commitment of 20, 30, or 40 hours per week.
- Minimum commitment is 4 hours per day and 20 hours per week, including 4 hours of overlap with Pacific Time.
- This engagement does not include medical benefits or paid leave.
$60 - $80 per hour
...Role Overview Join a Mathematician Expert Network that connects mathematicians with AI labs and companies to shape and evaluate cutting-edge AI in mathematics. Experts contribute domain expertise to model training and evaluation, create real-world tasks and deliverables...SuggestedHourly payContract workImmediate startRemote work- Kake is seeking Mathematics professionals with Python proficiency to contribute to project-based AI development initiatives focused on enhancing frontier AI models. You will be responsible for identifying suitable mathematical software, developing problem solutions in...SuggestedRemote work
$80 - $135 per hour
...09.26574v3), a frontier research-level physics benchmark. The role produces fully human-verified reference data used to evaluate large language model performance on frontier physics reasoning. Work includes solving CritPt research-level problems end-to-end, auditing other...SuggestedHourly payRemote work10 hours per week- A leading AI development firm is seeking a Mathematician (PhD) to join their remote AI training project. In this role, you will evaluate AI-generated mathematical responses, ensuring accuracy and clarity. To qualify, you must have a PhD in Mathematics/Statistics, significant...SuggestedWeekly payRemote workFlexible hours
$60 - $90 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...Jack Dorsey . Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract Compensation...SuggestedFull timeContract workSummer workRemote work$208k - $300k
...Machine Learning Engineer - Model Evaluations, Public Sector The Public Sector ML team at Scale deploys advanced AI systems—including LLMs, agentic models, and multimodal pipelines—into mission-critical government environments. We build evaluation frameworks that ensure...Full time$100 per hour
...Role Overview Apply deep domain expertise to train and evaluate next-generation AI systems by producing, refining, and validating high-quality,... ...data. This part-time contractor role focuses on improving model outputs through careful content review, prompt refinement,...Hourly payPart timeFor contractorsRemote work$60 - $90 per hour
...Role Overview Help advance frontier AI research by creating rigorous, real-world data analysis evaluations for generative AI models. You will design and complete complex analytical tasks that mirror practical research work, then use your reference analyses to assess where...Hourly payFull timeFreelanceRemote work$70 per hour
...original, executable scientific computing challenges that help evaluate the limits of advanced AI systems. This role focuses on creating rigorous... ...scientific computing that are difficult for current frontier AI models to solve. Source problem material from published...Hourly payPart timeFor contractorsImmediate startRemote work$224k - $356.5k
...people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts... ...computing. As a Senior / Principal Deep Learning Engineer — Model Evaluation & AI Systems, you will play a meaningful role in crafting the...Full time- Cincinnatus LLC is placing finance SMEs at a leading AI lab to critically evaluate AI model outputs for financial tasks. The role focuses on rigorous, rubric-based assessment of model performance and constructing finance-focused evaluation frameworks. Candidates should...Weekday work
- Cincinnatus LLC is recruiting for an Insurance Subject‑Matter Expert to join a leading AI lab's GenAI team in New York. This W-2 role involves evaluating insurance tasks and guiding model development for high‑quality underwriting judgments, with placement at the client...Weekday work
$60 per hour
...Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing...Hourly payRemote workWork from homeFlexible hours$65 - $90 per hour
...Apply deep financial judgment to help improve foundational AI models. In this role, you will create and evaluate finance-focused work that strengthens AI systems'' reasoning, analysis, and decision-making capabilities. Role Overview This is a W-2 employment opportunity...Hourly payWeekday work$13 - $54 per hour
...Evaluate AI-generated music and lyrics across a wide range of genres, using your knowledge of the Spanish (Mexico) music scene to assess quality against detailed standards. This remote role involves working in both Spanish (Mexico) and English. Key Responsibilities...Hourly payImmediate startRemote workFlexible hours- ...reasoning and computational problem solving to improve and evaluate large language models. You will design rigorous math problems, produce clear, logically... ...How this work supports customers Accelerate frontier AI research by contributing high quality data and evaluation...Contract workFor contractorsFreelanceRemote work
$70 - $110 per hour
...Help advance frontier AI systems by bringing rigorous materials science and engineering judgment to the evaluation, design, and improvement of technical knowledge work. You will... ...reasoning looks like in practice and ensure model outputs can withstand technical scrutiny....Hourly payFull timeLive inRelocationRelocation package$60 - $80 per hour
...Help shape the training and evaluation of foundational large language models by applying real-world expertise in brand strategy, growth marketing, and campaign... .... This role brings rigorous marketing judgment to AI tasks, model assessments, and training data for a leading...Hourly payWeekday work$35 - $62 per hour
...Apply your Japanese music expertise to evaluate AI-generated music and lyrics across a wide range of genres. You will assess outputs against detailed quality standards in both Japanese and English. Key Responsibilities Compare AI-generated lyrics with published songs...Hourly payFor contractorsImmediate startRemote workFlexible hours$17 - $42 per hour
...Evaluate AI-generated music and lyrics in Hebrew and English, applying your knowledge of the Hebrew music scene and detailed quality standards across a wide range of genres. Key Responsibilities Compare AI-generated lyrics with published songs to identify similarities...Hourly payFor contractorsImmediate startRemote workFlexible hours- ...Overview Apply advanced physics knowledge to help improve and evaluate large language models. You will design rigorous problems, produce clear reasoning... ...into accessible explanations while contributing to AI research projects. Key Responsibilities Design and solve...For contractorsFreelanceRemote work
$18 - $42 per hour
...Role Overview Evaluate generative music AI outputs across a wide range of genres, applying your knowledge of Portuguese-language music and lyrics to detailed quality standards. This role combines critical listening with lyric analysis in Portuguese and English. Key...Hourly payImmediate startRemote workFlexible hours$18 per hour
...Role Overview Evaluate AI-generated music and lyrics across a range of genres, applying detailed quality standards in both Thai and English... ...of Thai music, language, and lyrical expression to help assess model outputs. Key Responsibilities Compare AI-generated lyrics...Hourly payFor contractorsImmediate startRemote workFlexible hours$70 - $110 per hour
...Role Overview Help a leading AI research team improve how advanced AI models reason about real clinical work. In this hybrid, full-time role, you will... ...define high-quality clinical tasks, model answers, and evaluation standards alongside research and program management...Hourly payFull timeFreelanceLive inRelocationRelocation package$400 per month
About the Role Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments. The work focuses on realistic data engineering workflows...$20 - $60 per hour
...Help train next-generation AI systems by creating rigorous, real-world evaluations that test how well advanced models learn, reason, and perform. This remote contract opportunity is open to recent graduates, advanced-degree holders, and professionals from any background...Hourly payContract workFor contractorsRemote work- ...Evaluate generative music AI across a broad range of genres, applying your knowledge of Indonesian music and lyrics to detailed quality standards. You will work in both Indonesian and English to help assess lyric quality and authenticity. Key Responsibilities Compare...Hourly payImmediate startRemote workFlexible hours
- Cincinnatus LLC is seeking a Marketing SME to join a leading GenAI team focused on evaluating AI outputs against rubrics and strengthening brand strategy in AI training data. We require 8+ years of marketing experience with top-tier brands, plus hands-on evaluation of LLM...Weekday work
$70 - $110 per hour
...Role Overview Apply your legal practice experience to evaluate and improve AI-generated legal content and workflows. You will assess work grounded in litigation, legal drafting, and client advisory practice. Key Responsibilities Evaluate AI-generated legal memoranda...Hourly payImmediate startRemote work$60 per hour
...in Arizona, is seeking Biology Experts and Life Science Professionals to join their Expert Network. This role involves evaluating and training AI models with real scientific expertise. Successful candidates will review AI-generated content for accuracy and assist in fact...Hourly pay
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Mathematician for AI Model Evaluation. Be the first to apply!



