Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Psychologist for AI Model Evaluation

$55 per hour

SaidGig

Role Overview

Psychology experts apply clinical and research knowledge to evaluate and improve large language models (LLMs) in psychological domains, designing prompts, assessing responses for accuracy and nuance, and contributing research-informed feedback that shapes future AI behavior.

Key Responsibilities
  • Create and refine domain-specific prompts and test tasks to probe LLM performance.
  • Evaluate LLM outputs for factual accuracy, conceptual clarity, ethical considerations, and appropriate clinical framing.
  • Conduct independent literature or topic research, using AI tools as an aid, to support task design and evaluation.
  • Document findings and provide structured feedback to research teams to help improve model behavior.
  • Participate asynchronously in project work, collaborating with research teams and a small cohort of domain experts when assigned to a project.
Qualifications
  • Master''s degree from a US institution, or current master’s student in Psychology at a US institution.
  • Subject-matter expertise in psychology strong enough to critically evaluate and often outperform current AI systems in explaining key concepts.
  • Comfort working primarily in asynchronous formats and coordinating with research teams remotely.
  • Selection for individual projects is competitive, with cohorts chosen for disciplinary and research expertise.
Work Terms
  • Employment type: Part-time.
  • Work is remote and asynchronous, allowing independent scheduling from any location.
  • Flexible hours with no minimum required weekly commitment.
  • Placement into specific projects depends on available projects in the relevant domain.
Compensation
  • Pay rate: Up to $55/hr (depending on the project).
Eligibility
  • F-1 students may be eligible if they have CPT or OPT authorization; consult your Designated School Official to confirm eligibility.
  • If your school requires CPT to be tied to a course, this program may not meet that requirement.
  • STEM OPT is not supported for this program.
Project Availability

The program operates year-round, however project openings vary by domain and time. Assignment to projects is contingent on project availability and the fit between your expertise and project needs.

Application Process
  • Create a candidate account and submit your profile materials.
  • Complete identity verification as instructed.
  • Apply to or join relevant project openings and complete any project-specific onboarding when selected.
  • Begin assigned work and receive compensation according to the project terms.
Vacancy posted more than 2 months ago
Similar jobs that could be interesting for youBased on the Psychologist for AI Model Evaluation in United States vacancy
  • $75 per hour

     ...Recreational therapists can apply their clinical expertise to AI evaluation projects by assessing field-specific content and providing structured feedback that improves how AI models understand therapeutic practice, workplace tasks, and professional language. Role Overview... 
    Suggested
    For contractors
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $208k - $300k

     ...Machine Learning Engineer - Model Evaluations, Public Sector The Public Sector ML team at Scale deploys advanced AI systems—including LLMs, agentic models, and multimodal pipelines—into mission-critical government environments. We build evaluation frameworks that ensure... 
    Suggested
    Full time

    Scale Ai

    Washington DC
    21 hours ago
  •  ...is seeking Biology Experts and Life Science Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities... 
    Suggested
    Hourly pay
    Remote work
    Flexible hours

    Prolific

    Jacksonville, FL
    1 day ago
  • $224k - $356.5k

     ...people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts...  ...computing. As a Senior / Principal Deep Learning Engineer — Model Evaluation & AI Systems, you will play a meaningful role in crafting the... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    21 hours ago
  •  ...Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant... 
    Suggested
    Remote work
    Flexible hours

    Prolific

    Charlotte, NC
    2 days ago
  • $60 per hour

     ...Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing... 
    Hourly pay
    Remote work
    Work from home
    Flexible hours

    Prolific

    San Jose, CA
    2 days ago
  • Cincinnatus LLC is placing finance SMEs at a leading AI lab to critically evaluate AI model outputs for financial tasks. The role focuses on rigorous, rubric-based assessment of model performance and constructing finance-focused evaluation frameworks. Candidates should... 
    Weekday work

    Mercor Inc

    San Francisco, CA
    3 days ago
  • $40 per hour

    A technology company in Massachusetts is seeking an R&D Biologist to join their team to train AI models by evaluating chatbot outputs against complex biology questions. Ideal candidates will hold advanced qualifications in biology or biochemistry. This position allows full... 
    Hourly pay
    Full time
    Part time
    Remote work

    DataAnnotation

    Oklahoma City, OK
    2 days ago
  • $150 per hour

     ...Overview Aerospace engineering professionals apply their domain expertise to evaluate AI-generated outputs, assess technical content, and provide clear, structured feedback that improves models'' understanding of aerospace tasks, terminology, and practices. You will... 
    Hourly pay
    Temporary work
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $35 - $62 per hour

     ...Evaluate AI-generated music and lyrics across a wide range of genres, applying Korean music expertise and detailed quality standards in both Korean and English. Key Responsibilities Compare AI-generated lyrics with published songs to identify similarities. Rate... 
    Hourly pay
    For contractors
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    6 days ago
  • $65 - $90 per hour

     ...Role Overview Apply your real-world architecture expertise to evaluate and improve how AI systems understand and reason about architecture. In this flexible, part-time, remote role you will review content for technical accuracy, answer domain-specific questions, and provide... 
    Hourly pay
    Part time
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    Remote
    17 days ago
  • $60 - $80 per hour

     ...GenAI team building foundational large language models by applying deep insurance domain expertise to create, evaluate, and refine training data. You will translate real...  ...real underwriting and claims practice. Evaluate AI model outputs against structured rubrics,... 
    Hourly pay
    Weekday work

    SaidGig

    United States
    a month ago
  • $14 - $42 per hour

     ...Evaluate AI-generated music and lyrics across a wide range of genres, applying your knowledge of the Hindi music scene and detailed quality standards. Key Responsibilities Assess AI-generated music and rate it against established quality criteria. Compare AI-generated... 
    Hourly pay
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    6 days ago
  • $17 - $54 per hour

     ...Role Overview Evaluate AI-generated music and lyrics across a wide range of genres, applying detailed quality standards in both French and English. Key Responsibilities Compare AI-generated lyrics with published songs to identify similarities. Rate lyrics for... 
    Hourly pay
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    6 days ago
  • $85 per hour

     ...Role Overview Environmental Analysts apply environmental assessment, GIS, and renewable energy siting expertise to evaluate AI-generated outputs and to create expert-level training data that improves AI understanding of environmental workflows and geospatial data practices... 
    Hourly pay
    Full time
    Contract work
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  •  ...Role Overview Work with a leading AI lab to evaluate outputs from generative music models in German and English. This role focuses on listening, scoring, and annotating AI-generated music and lyrics across genres, using music production and audio engineering vocabulary... 
    Hourly pay
    Part time
    Immediate start
    Remote work
    10 hours per week

    SaidGig

    United States
    a month ago
  •  ...reasoning and computational problem solving to improve and evaluate large language models. You will design rigorous math problems, produce clear, logically...  ...How this work supports customers Accelerate frontier AI research by contributing high quality data and evaluation... 
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    22 days ago
  •  ...Role Overview Provide expert chemical engineering knowledge to evaluate and improve AI systems, ensuring domain accuracy and practical usefulness. You will review AI-generated content, answer technical questions, and share real-world practices, tools, and standards used... 
    Hourly pay
    Part time
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    Remote
    17 days ago
  • $85 - $120 per hour

     ...Apply senior-level legal judgment to help a leading AI research team improve how frontier models handle real-world legal work. You will turn specialized...  ...tasks, model instructions, reference solutions, and evaluation benchmarks. Key Responsibilities Review legal knowledge... 
    Hourly pay
    Full time
    Internship
    Live in
    Relocation
    Relocation package

    SaidGig

    Bay County, FL
    6 days ago
  • $75 per hour

     ...apply geospatial imaging, survey, and photogrammetric expertise to evaluate AI-generated maps and geospatial content, verify spatial accuracy, and provide clear, structured feedback that improves model outputs. No prior AI experience is required. Key Responsibilities... 
    Hourly pay
    Temporary work
    Part time
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $75 per hour

     ...Records Managers apply professional records, archives, and library expertise to evaluate AI-generated outputs, create field-accurate prompts, and provide structured feedback that improves model performance on records-management tasks. Candidates can include Archivists,... 
    Hourly pay
    Temporary work
    Part time
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • Apple Inc. is seeking a Research Scientist/Engineer to design evaluation systems for foundation models powering Apple products. You will work hands‑on across evaluation design, experimentation, and cross‑team collaboration to drive model improvement and product quality... 

    Apple Inc.

    Cupertino, CA
    2 days ago
  • $60 per hour

     ...in Arizona, is seeking Biology Experts and Life Science Professionals to join their Expert Network. This role involves evaluating and training AI models with real scientific expertise. Successful candidates will review AI-generated content for accuracy and assist in fact... 
    Hourly pay

    Prolific

    Arizona City, AZ
    5 days ago
  • Dorado is seeking an experienced Investment Banking SME to support the development, evaluation, and improvement of advanced AI models in finance. You will assess AI-generated analyses for accuracy, reasoned judgments, and data integrity, guiding model refinements and prompts... 
    Remote job

    Dorado

    New York, NY
    4 days ago
  • $60 per hour

     ...and Life Science Professionals to join their Expert Network to evaluate AI-generated science. This role allows you to work from home with...  ...offers a competitive pay rate of up to $60 per hour for reviewing model responses, validating technical claims, and critiquing... 
    Remote job
    Hourly pay
    Work from home
    Flexible hours

    Prolific

    Dallas, TX
    5 days ago
  • $17 - $42 per hour

     ...Role Overview Assess outputs from generative music AI models in Greek and English, applying professional music production and audio...  ...multiple genres and styles. Key Responsibilities Listen to and evaluate AI-generated music samples, including lyrics and synthesized... 
    Hourly pay
    Part time
    Immediate start
    Remote work
    10 hours per week

    SaidGig

    United States
    a month ago
  •  ...Role Overview Evaluate AI-generated music in Telugu and English for a research partnership with a leading AI lab. You will listen critically...  ...labels and feedback that help improve generative music models. Key Responsibilities Listen to and compare pairs or sets... 
    Hourly pay
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    a month ago
  •  ...Role Overview Evaluate and annotate AI-generated music in Arabic and English, providing detailed technical and creative feedback to improve generative musical models. This role focuses on listening, comparing, and scoring samples for musicality, vocal performance, lyrics... 
    Hourly pay
    Part time
    Immediate start
    Remote work
    10 hours per week

    SaidGig

    United States
    a month ago
  • $65 - $90 per hour

     ...Role Overview Bring real-world civil engineering expertise to a small, growing network of industry specialists who evaluate and improve how AI systems understand and reason about civil engineering topics. You will review technical content for accuracy, answer domain questions... 
    Hourly pay
    Part time
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    Remote
    17 days ago
  • $17 - $42 per hour

     ...Role Overview Evaluate AI-generated music across many genres for a leading AI research partner. You will listen critically to generated...  ...accuracy in Hebrew and English to help improve generative music models. Key Responsibilities Compare AI-generated song pairs... 
    Hourly pay
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    6 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Psychologist for AI Model Evaluation. Be the first to apply!