Researcher for AI Model Evaluation
$40 - $90 per hourSaidGig
Role Overview
Apply advanced research expertise to help train and evaluate next-generation AI systems by creating rigorous, real-world assessment material. This remote contract role focuses on developing high-quality evaluations that shape how AI models learn, reason, and perform. Prior AI experience is not required.
Key Responsibilities
- Create original, high-difficulty question-and-answer pairs in your discipline to assess AI model capabilities.
- Research, verify, and document answers using primary sources and authoritative references, with clear citations and reasoning.
- Design questions that require advanced reasoning, methodological nuance, or synthesis across multiple sources and cannot be solved through shortcuts.
- Test questions against AI systems, identify insufficiently challenging items, and increase complexity while preserving accuracy.
- Write precise, unambiguous questions and defensible answers.
- Incorporate reviewer feedback and follow all project guidelines and quality standards.
Qualifications
- Completed PhD, active PhD candidacy, or equivalent research experience as a specialist, researcher, or professor in your field.
- Demonstrated scholarly research experience, deep subject-matter expertise, and familiarity with primary literature.
- Strong analytical thinking, research and source-triangulation skills, attention to detail, written precision, and English communication skills.
- Ability to work independently and reliably in a remote setting.
- Experience developing original, challenging, methodologically sound questions.
- AI training or evaluation experience is a plus, but not required.
Work Terms
- Remote, independent contractor engagement.
- Compensation is output-based and paid per task that meets project specifications; completion time varies by experience and workflow.
- Minimum submission requirements apply, including a minimum number of tasks submitted each week.
- Selected candidates are expected to begin their first tasks within 24 to 48 hours after completing onboarding.
Compensation
$40 to $90 per hour.
Application Process
Apply using email or a Google account, then complete the required onboarding steps and agree to the applicable terms and privacy policies.
$70 - $90 per hour
...Role Overview Evaluate vulnerability reproduction and remediation tasks used to train and assess frontier AI models. You will determine whether CVE reproductions faithfully reflect... ...penetration testing, or vulnerability research. ~ Strong knowledge of CVE taxonomy and...SuggestedRemote jobHourly pay$70 - $110 per hour
...Help advance frontier AI systems by bringing rigorous materials... ...and engineering judgment to the evaluation, design, and improvement of... ...will work closely with an AI research team to define what high quality... ...looks like in practice and ensure model outputs can withstand...SuggestedHourly payFull timeLive inRelocationRelocation package- ...Role Overview Apply research-grade expertise to help evaluate and improve AI reasoning across technical and humanities disciplines, including STEM, English, literature, and journalism. This remote contract role focuses on producing rigorous training data, evaluations...SuggestedHourly payContract workFor contractorsRemote work
$60 - $90 per hour
...connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include... ...Jack Dorsey . Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract Compensation...SuggestedFull timeContract workSummer workRemote work$60 - $80 per hour
...that connects mathematicians with AI labs and companies to shape and evaluate cutting-edge AI in mathematics. Experts... ...contribute domain expertise to model training and evaluation, create real... ...detailed feedback that helps advance research. This listing is an open...SuggestedHourly payContract workImmediate startRemote work- ...reasoning and computational problem solving to improve and evaluate large language models. You will design rigorous math problems, produce clear,... ...How this work supports customers Accelerate frontier AI research by contributing high quality data and evaluation...Contract workFor contractorsFreelanceRemote work
- SME Careers is seeking biologists to contribute to an AI training project that involves reviewing AI-generated responses and providing... ...hold a MS or PhD in a relevant field and have experience in evaluating complex biology content. Strong communication skills and proficient...Immediate start
- ...Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities include fact-checking AI technical...Hourly payRemote workFlexible hours
$136.44k - $265.11k
...are rebuilding biotech for the AI era.When a breakthrough is... ...data, and run AI agents and models directly in their workflows.... ...here.You’ll build the datasets, evaluations, and systems that help close... ...engineers, scientists, and external research partners.Desire to work in a...Work at officeLocal areaMonday to FridayShift work- ...Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant...Remote workFlexible hours
- Cincinnatus LLC is recruiting for an Insurance Subject‑Matter Expert to join a leading AI lab's GenAI team in New York. This W-2 role involves evaluating insurance tasks and guiding model development for high‑quality underwriting judgments, with placement at the client...Weekday work
- Cincinnatus LLC is placing finance SMEs at a leading AI lab to critically evaluate AI model outputs for financial tasks. The role focuses on rigorous, rubric-based assessment of model performance and constructing finance-focused evaluation frameworks. Candidates should...Weekday work
$65 - $105 per hour
...Help advance frontier AI models by bringing rigorous life sciences research judgment into task design, evaluation, and model improvement. You will work closely with an AI lab''s research and program management teams to ensure models can reason credibly about real scientific...Hourly payFull timeFreelanceLive inRelocationRelocation package$60 - $90 per hour
...Role Overview Design rigorous, real-world materials science and engineering tasks that evaluate an AI model’s expert reasoning. You will create prompts, supporting data rooms, and grading criteria, then test and refine each task until it clearly distinguishes strong...Hourly payWork at officeRemote work- Thinking Machines in San Francisco seeks a researcher to advance internal evaluations and signals for post-training models. You will collaborate with researchers and engineers across the research organization, shaping evaluation creation, usability, and auditing. Your work...
$70 - $80 per hour
...Role Overview Apply advanced drug safety expertise to help improve next-generation AI systems through high-quality evaluations, safety-report analysis, and structured feedback. This remote contractor opportunity is designed for professionals with experience authoring...Hourly payFor contractorsRemote work- ...Role Overview Use your investment and finance expertise to evaluate and improve AI model performance on financial reasoning, valuation, markets, and real-world investment scenarios. Key Responsibilities Assess AI model outputs on valuation, financial modeling, markets...For contractorsRemote work
$70 - $90 per hour
...Role Overview Help evaluate Neuron Kernel Interface development tasks that support the training and evaluation of advanced AI models. You will assess kernel quality, numerical correctness, CUDA-to-NKI migration fidelity, and whether implementations are well suited to...Hourly payRemote work$100 per hour
...expertise to improve the performance of large language models on finance tasks. You will work with AI researchers to identify model weaknesses in areas such as... ...on advanced AI systems. Key Responsibilities Evaluate LLM performance in finance areas where models...Hourly payContract workFor contractorsFreelanceRemote work10 hours per weekFlexible hours$100 - $150 per hour
...Overview Apply senior legal judgment to help a leading AI research organization improve how advanced AI models reason about real-world legal work. You will work... ...define high-quality legal tasks, standards, and evaluations. This role is designed for a practicing legal...Hourly payFull timeFreelanceInternshipLive inRelocationRelocation package$60 - $80 per hour
...Help shape the training and evaluation of foundational large language models by applying real-world expertise in brand... ...rigorous marketing judgment to AI tasks, model assessments, and training... .... Key Responsibilities Advise research and engineering teams on knowledge...Hourly payWeekday work$70 - $90 per hour
...Review GPU and accelerator kernel development tasks that support the training and evaluation of advanced AI models. This role focuses on assessing task quality, numerical correctness, completeness, fair performance benchmarking, appropriate scope, and whether kernels...Hourly payRemote work$50 - $75 per hour
A leading tech company based in Australia is seeking an AI Model Evaluator on a contract basis. The role involves evaluating AI-generated responses, writing prompts, and providing justifications based on specific criteria. Ideal candidates will hold a Master's degree in...Hourly payContract work- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code). You will author original, executable research problems that today's frontier models cannot solve. The role emphasizes material sourcing, prompt design, and robust grading criteria against...Part timeImmediate start
- Google DeepMind seeks a Senior Product Manager embedded in Gemini research and model training. You will read evaluations, analyze model outputs, and make judgment calls on quality alongside researchers, translating user needs into product priorities and feedback loops...
- YO AI Labs seeks an Investment & Finance Expert (Contractor) to evaluate and improve AI models' financial reasoning from investment banking, private equity, VC, hedge funds, or growth equity backgrounds. Remote work allowed; strong English communication required. Responsibilities...Remote jobFor contractors
- ...Description - Member Of Technical Staff (Language Model Evaluations) Location: San Francisco (preferred),... ...Analysis is the leading independent AI benchmarking company. We support labs,... ...models, working directly with their research teams; our commercial team owns client relationships...
$300k - $320k
...role: We are seeking a Technical Program Manager to lead our AI model evaluation initiatives across multiple workstreams. This role will be... ...potential risks of our AI models. Working closely with our Research, Trust & Safety, Frontier Redteaming, and Policy teams, you...Work at officeHome officeVisa sponsorshipRelocation package- Sign in to set job alerts for “Regional Clinical Research Associate” roles. 1,000+ Regional Clinical Research Associate Jobs in United States Clinical Research Associate (CRA) - Cardiovascular (Remote) Clinical Research Associate (CRA) - Cardiovascular (Remote) Senior...Remote jobRelocation package
- ...Labs in Seattle is seeking a Member of Technical Staff for RL Research, aimed at recent PhD graduates in AI or ML. In this impactful role, you will own the RL and post-training for large-scale omni models and contribute to developing advanced AI systems. You will work...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Researcher for AI Model Evaluation. Be the first to apply!
- survey researcher United States
- content researcher United States
- legal researcher United States
- remote researcher United States
- independent researcher United States
- junior researcher United States
- music researcher United States
- academic researcher United States
- junior security researcher United States
- vulnerability researcher United States




