Remote Math Expert for AI Benchmarking & Evaluation
$100 per hourTuring
Turing is seeking PhD-level mathematicians to design challenging AI evaluation problems and craft rigorous, step-by-step solutions. You will assess AI reasoning, ensure clarity of feedback, and collaborate with researchers to build robust benchmarks across math topics from undergraduate to PhD level.
This fully remote contract role offers flexible hours and a $100/hour pay rate. Candidates should have a PhD in mathematics or related fields, strong reasoning, and communication skills.
#J-18808-Ljbffr$50 per hour
...focuses on improving and evaluating large language models... ...advances into reliable AI systems for enterprise... ...Contribute to new evaluation benchmarks spanning curricula from... ...effectively in a remote setting. Self-motivated... ...and solve complex math problems using a structured...Remote workContract workFor contractorsFreelance- ...Remote Mathematics Expert (AI/LLM) - 34877 Remote Mathematics Expert (AI/LLM) - 3... ...Design and solve challenging math problems to probe the... ...align problem types with evaluation goals, particularly in areas... ...to defining new evaluation benchmarks based on Mathematics curricula...Remote workHourly payContract workPart timeFor contractorsFreelanceInternshipWork from homeWorldwideAfternoon shift
$60 - $80 per hour
...technical talent with leading AI research labs.... ...our investors include Benchmark , General Catalyst ,... ...0/hour Location: Remote Role Responsibilities... ...real retail practice. Evaluate AI model outputs... ...with other subject matter experts to ensure consistency and...Remote workContract workSummer workWeekday work- ...to help train next-generation AI systems. Your work will shape... ...prompts. # Participate in remote collaboration, contributing via... ...the rigorous training and evaluation of advanced AI models.... ...producing data-annotation tasks, benchmark questions, technical reports,...Remote workTemporary work
$60 - $75 per hour
Role Description Join an advanced AI research initiative focused on improving how... ...professionals to design high-quality benchmark tasks that evaluate AI performance across software... ...structured outputs. This is a fully remote, independent contractor opportunity with...Remote workWeekly payContract workPart timeFor contractorsFlexible hours$50 per hour
A leading AI research accelerator is looking for remote PhD candidates in Chemistry, Chemical Engineering, or related fields... ...design advanced chemistry problems to evaluate AI performance and collaborate with researchers on benchmarks. This role offers flexible hours and a...Remote workHourly payFlexible hours- AuraOne is seeking a Ukrainian Law Expert for the Ukrainian Law Expert — Multiple-Choice AI Benchmark Review (Remote). This remote review track evaluates AI outputs in legal review workflows, with tasks like verifying citations and flagging policy-adherence gaps. As a...Remote jobFor contractors10 hours per weekFlexible hours
$60 - $70 per hour
...and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel... ...Compensation: $60–$70/hour Location: Remote Role Responsibilities Evaluate AI-generated responses for safety...Remote workContract workSummer work$90 per hour
...technical talent with leading AI research labs. Headquartered in... ..., our investors include Benchmark , General Catalyst , Peter... ...Compensation: $90/hour Location: Remote Role Responsibilities... ...that merely look correct. Evaluate responsive behavior and semantic...Remote workContract workSummer workLocal area- ...Role Overview Provide expert human judgment on commercial drug launches by creating and critiquing evaluation rubrics, building or reviewing launch curves, and assessing the... ...total over a 1 to 2 week pilot period. Remote work, candidates must be US-based or have deep...Remote workHourly pay
$80 per hour
...inventory planning, and supply chain operations to create expert training data and evaluate AI-generated responses for accuracy and relevance. This is... ...type, hourly contract work, part time. Fully remote, flexible and asynchronous schedule, no minimum weekly hour...Remote workHourly payContract workPart timeFlexible hours$150 - $180 per hour
...Virginia Beach is seeking Mental Health Professionals to train and evaluate AI models. The role involves reviewing AI responses to... ...attention to detail, and a reliable internet connection. Join our Expert Network to influence future AI innovations and work flexibly from...Remote jobWork from home- ...About Us Join an innovative AI initiative dedicated to improving... ...their real-world expertise by evaluating AI-generated content,... ...This is a flexible, fully remote, part-time opportunity requiring... ...multidisciplinary team of subject matter experts to improve AI performance...Remote workPart time10 hours per weekFlexible hours
$150 - $180 per hour
Prolific in Sacramento is seeking Mental Health Professionals to help train and evaluate advanced AI models. You will review AI-generated responses and engage in various training tasks, earning competitive pay rates up to $150-180/hr. Ideal candidates must hold a verified...Remote jobWork from home$8 - $65 per hour
Prolific is hiring Mental Health Professionals in New York to train and evaluate AI models. As a Domain Expert, you will be responsible for reviewing AI-generated responses, completing tasks related to psychology, and improving AI models based on your expertise. Pay rates...Remote jobHourly payWork from homeFlexible hours- Prolific in New York, NY, is seeking Chemistry Experts and Chemical Engineers to join its Expert Network. In this role, you will help train and evaluate AI models using your chemical expertise. Duties include evaluating AI-generated responses for accuracy and validating...Remote jobWork from homeFlexible hours
- ...Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant...Remote workFlexible hours
- ...A leading AI research firm is seeking Expert Prompt Curators to design challenging prompts for evaluating advanced AI models. The role requires advanced knowledge in diverse fields and offers flexible hours, remote work, and a competitive hourly wage. Ideal candidates...Remote workHourly payTemporary workFlexible hours
$60 per hour
Prolific in Seattle, WA is searching for Chemistry Experts and Chemical Engineers to join our Expert Network to train AI models with your expertise. The role involves evaluating AI-generated chemistry, fact-checking chemical reactions, and auditing technical documentation...Remote jobHourly payWork from homeFlexible hours$50 - $70 per hour
...Role Overview Help improve frontier AI models by evaluating the quality of real-world professional materials and AI-generated work. You will... ...Guidelines and context will be provided. Work Terms Remote, hourly engagement. Applicants must be based in the United...Remote workHourly pay$65 per hour
...security expertise to design domain-specific prompts and evaluate large language model outputs for AI research projects, improving model behavior, safety,... ..., or Hashcat. Ability to work independently in a remote, asynchronous setting and to document findings clearly...Remote workPart timeFlexible hours- A leading AI research accelerator is seeking an Associate to leverage expertise in internal... .... This role involves designing and evaluating clinical scenarios to enhance AI diagnostics... ...skills. The position is fully remote, requiring a commitment of 20-40 hours per...Remote job
- Prolific is seeking Mental Health Professionals in Indianapolis, Indiana, to assist in training and evaluating AI models. Ideal candidates will have a verified status as a Mental Health Professional, an understanding of psychological theory, and the ability to focus on...Remote jobHourly payFlexible hours
$8 - $65 per hour
...Mental Health Professionals in Houston, Texas, to train and evaluate cutting-edge AI models. The role offers flexible hours and competitive pay... ...This is an opportunity to influence AI development and work remotely. Join Prolific in just 15 minutes after passing the...Remote jobHourly payFlexible hours$8 - $65 per unit
Prolific is seeking Mental Health Professionals to train and evaluate AI models. In this role, you will review AI responses, analyze psychology... ...pay rates range from $8 to $65 per task, with flexible hours and remote work options available. #J-18808-Ljbffr ProlificRemote jobFlexible hours$60 per hour
Prolific seeks Chemistry Experts and Chemical Engineers to join our Expert Network and evaluate AI models using your chemical expertise. Successful candidates will be invited to assess and review AI-generated chemistry tasks and ensure their accuracy. Compensation can reach...Remote jobHourly payFlexible hours$60 per hour
Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing...Remote jobHourly payWork from homeFlexible hours$8 - $65 per hour
...for Mental Health Professionals to help train and evaluate AI models in Phoenix, Arizona. As a Domain Expert participant, you'll review AI-generated psychological... ...completed task, with flexible hours that allow for remote work. Applicants need verified professional status...Remote jobFlexible hours$8 - $65 per hour
Prolific is seeking Mental Health Professionals to train and evaluate AI models from home. Responsibilities include reviewing AI responses and enhancing model performance using psychological expertise. Ideal candidates will have verified professional status, a solid understanding...Remote jobFlexible hours- Prolific is seeking Mental Health Professionals to train and evaluate AI models. Candidates will review psychological scenarios, analyze... ...health professionals and must complete a quick skills assessment to join Prolific’s Domain Expert Network. #J-18808-Ljbffr ProlificRemote jobFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Remote Math Expert for AI Benchmarking & Evaluation. Be the first to apply!




