Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Evaluation Specialist for AI Model Evaluation

$20 - $60 per hour

SaidGig

Role Overview

Help train next-generation AI systems by creating rigorous, real-world evaluations that test how advanced models learn, reason, and perform. This remote contractor opportunity is open to recent graduates, advanced-degree holders, and professionals with strong research and writing capabilities. Prior AI experience is not required.

Key Responsibilities

  • Develop original, challenging question-and-answer pairs across diverse subjects.
  • Conduct in-depth research and triangulate multiple sources to produce accurate, comprehensive, well-documented answers.
  • Create multi-step questions that require synthesis and analytical reasoning rather than simple single-source lookups.
  • Test questions with AI models, analyze the results, and adjust difficulty as needed.
  • Document research methods with clear citations and logical explanations.
  • Revise content in response to reviewer feedback while following project guidelines and quality standards.

Qualifications

  • Demonstrated ability to independently research and critically assess information from multiple sources.
  • Excellent attention to detail, analytical thinking, and precise English writing. English fluency is required, though it need not be your first language.
  • Ability to write nuanced, well-structured questions that assess deep understanding.
  • Strong self-direction, reliability, and accountability in remote independent work.
  • Interest in advanced AI and evaluating current model capabilities.
  • Experience with AI training, evaluation, or content creation is helpful but not required.

Work Terms

  • Remote contract engagement.
  • Work is paid by qualifying task output, and completion time may vary based on experience and workflow.
  • Minimum submission requirements apply, including a minimum number of tasks submitted each week.
  • Roles are commonly filled within 48 hours. Selected candidates should be prepared to begin their first tasks within 24 to 48 hours after completing onboarding.

Compensation

Compensation is listed at $20 to $60 per hour, with payment determined on a per-task basis for work that meets project specifications.

Application Process

Submit an application using an email address or Google account. Candidates who advance complete onboarding before beginning project tasks. Questions can be reviewed through the available FAQs.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Evaluation Specialist for AI Model Evaluation in United States vacancy
  • $20 - $60 per hour

     ...document expertise to a project that trains next-generation AI systems. You will design realistic Fortune 500 style scenarios and interact iteratively with an advanced language model to create, edit, and evaluate Office Open XML files, with a focus on .pptx deliverables.... 
    Suggested
    Hourly pay
    Contract work
    For contractors
    Work at office
    Remote work

    SaidGig

    United States
    1 day ago
  • $75 - $115 per hour

     ...Role Overview Help advance frontier AI models by bringing rigorous pharmaceutical research and development judgment to the evaluation, design, and improvement of domain-specific...  ...Partner with researchers and adjacent-domain specialists to build pharmaceutical skills and tools... 
    Suggested
    Hourly pay
    Full time
    Contract work
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    2 days ago
  • $70 - $80 per hour

     ...Role Overview Apply advanced drug safety expertise to help improve AI systems through rigorous, real-world evaluation of pharmacovigilance documentation and data. This remote contract role focuses on the quality, accuracy, and regulatory alignment of complex safety reports... 
    Suggested
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    a month ago
  • $70 - $90 per hour

     ...Review GPU and accelerator kernel development tasks that support the training and evaluation of advanced AI models. This role focuses on assessing task quality, numerical correctness, completeness, fair performance benchmarking, appropriate scope, and whether kernels... 
    Suggested
    Hourly pay
    Remote work

    SaidGig

    Remote
    26 days ago
  • $100 - $150 per hour

     ...senior legal judgment to help a leading AI research organization improve how advanced AI models reason about real-world legal work....  ...quality legal tasks, standards, and evaluations. This role is designed for a practicing legal specialist with deep subject-matter expertise,... 
    Suggested
    Hourly pay
    Full time
    Freelance
    Internship
    Live in
    Relocation
    Relocation package

    SaidGig

    Bay County, FL
    a month ago
  •  ...Role Overview Use your investment and finance expertise to evaluate and improve AI model performance on financial reasoning, valuation, markets, and real-world investment scenarios. Key Responsibilities Assess AI model outputs on valuation, financial modeling, markets... 
    For contractors
    Remote work

    SaidGig

    United States
    24 days ago
  • $70 - $90 per hour

     ...Role Overview Help evaluate Neuron Kernel Interface development tasks that support the training and evaluation of advanced AI models. You will assess kernel quality, numerical correctness, CUDA-to-NKI migration fidelity, and whether implementations are well suited to... 
    Hourly pay
    Remote work

    SaidGig

    Remote
    26 days ago
  • $60 - $80 per hour

     ...expertise to help develop advanced large language models. In this role, you will bring practical brand, growth, and campaign judgment to AI training data, partnering with research...  ...reasoning quality. Develop and improve evaluation guidelines and scoring rubrics for... 
    Hourly pay
    Weekday work

    SaidGig

    United States
    a month ago
  •  ...Title: AI Evaluation & Model Risk Lead Location: Bellevue WA Engineer, AI - AI Evaluation & Model Risk Lead Are you ready to join the Un-carrier movement? This role leads how T-Mobile decides which AI models it can trust - the behavioral and model-risk... 
    Work experience placement

    Amaze Systems Inc.

    Washington DC
    2 days ago
  •  ...Role Title: Research Engineer - Code Generation & Model Evaluation Role Type: Contractor Location: Remote micro1 is engaging Research...  ...role, you'll apply your expertise to help train next-generation AI systems. Your work will shape how models learn, reason, and... 
    Temporary work
    For contractors
    Remote work

    micro1

    Remote
    2 days ago
  • $50 - $100 per hour

     ...Apply your software engineering expertise to help train and evaluate next-generation AI systems through real-world coding tasks, technical feedback...  ...contract role focuses on code generation workflows and model evaluation; prior AI experience is not required. Key Responsibilities... 
    Hourly pay
    Contract work
    For contractors
    Remote work

    SaidGig

    United States
    6 days ago
  • $136.44k - $265.11k

    We are rebuilding biotech for the AI era.When a breakthrough is delayed, the world waits...  ...structured data, and run AI agents and models directly in their workflows. Over 200,000...  ...our work here.You’ll build the datasets, evaluations, and systems that help close that gap.... 
    Work at office
    Local area
    Monday to Friday
    Shift work

    Benchling

    San Francisco, CA
    7 days ago
  •  ...reasoning and computational problem solving to improve and evaluate large language models. You will design rigorous math problems, produce clear, logically...  ...How this work supports customers Accelerate frontier AI research by contributing high quality data and evaluation... 
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    a month ago
  • $60 - $90 per hour

     ...engineering judgment to improve how advanced AI models reason through real-world engineering...  ...correct solutions, and create rigorous evaluations grounded in industry practice. Key...  ...with researchers and adjacent-domain specialists to calibrate consistent standards and convert... 
    Hourly pay
    Full time
    Remote work

    SaidGig

    United States
    3 days ago
  • $60 - $80 per hour

     ...mathematics experts considered for future contract opportunities with AI labs and companies. This is an open application, not a posting...  ...by applying mathematics knowledge to real-world tasks, model evaluation, and domain-specific feedback. Key Responsibilities... 
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $36 - $72 per hour

     ...seekers, 1 million+ employers, and 1,600 educational institutions. Handshake AI works directly with frontier AI lab researchers to create evaluations, publish benchmarks, and improve AI models through human expertise. Role Details Location: Onsite in Seattle, WA,... 
    Hourly pay
    Full time
    Monday to Friday
    Flexible hours

    Handshake

    Washington DC
    14 days ago
  • $60 - $80 per hour

     ...Overview Apply deep insurance expertise to help develop advanced large language models by bringing real-world underwriting, claims, and risk-assessment judgment to AI training and evaluation work. Key Responsibilities Partner with research and engineering teams to... 
    Hourly pay
    Weekday work

    SaidGig

    United States
    more than 2 months ago
  • $100 - $150 per hour

     ...data scientists who will be considered for future projects evaluating how well AI systems perform real-world data science tasks. Members of this...  ...criteria, assess AI or human-produced analyses and models, document decisions in writing, and iterate on evaluations with... 
    Hourly pay
    Immediate start
    Remote work

    SaidGig

    United States
    a month ago
  • $110 per hour

     ...Apply to join a physician talent network supporting AI labs and companies with medical expertise. This is an open application...  ...projects by bringing real-world clinical expertise to model development and evaluation. Key Responsibilities Train and evaluate AI models... 
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    a month ago
  • Zebra Technologies is seeking an AI Quality Analyst to ensure performance, safety, and reliability of cutting-edge AI/ML models. Design evaluation strategies, identify edge cases, bias sources, and provide actionable insights to drive model improvements across the development... 

    RXinsider LTD.

    Lincolnshire, IL
    4 days ago
  • $60 - $90 per hour

     ...Role Overview Design rigorous, real-world materials science and engineering tasks that evaluate an AI model’s expert reasoning. You will create prompts, supporting data rooms, and grading criteria, then test and refine each task until it clearly distinguishes strong... 
    Hourly pay
    Work at office
    Remote work

    SaidGig

    United States
    20 days ago
  • $65 - $105 per hour

     ...engineering judgment to help frontier AI models reason more accurately about real-world...  ...define high-quality engineering work, evaluate model performance, and turn expert practice...  .... Collaborate with researchers and specialists in adjacent disciplines to calibrate... 
    Hourly pay
    Full time
    Freelance
    Internship
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    a month ago
  •  ...Hallucination Failure Model Evaluation Specialist is a remote evaluation track for reviewing hallucination failure model evaluation evaluation prompts...  ...team can use to retrain. Why this role matters AI data reviewers help turn hallucination failure model evaluation... 
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    20 days ago
  • $70 - $110 per hour

     ...Overview Shape how advanced AI systems reason about real clinical...  ...with an AI research team to evaluate medical knowledge tasks,...  ...benchmarks that measure meaningful model improvement. Key Responsibilities...  ...and adjacent-domain specialists to calibrate standards and translate... 
    Hourly pay
    Full time
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    a month ago
  • $100 - $150 per hour

     ...finance expertise to improve how frontier AI systems reason through real-world...  ...team to define high-quality finance tasks, evaluate model performance, and translate professional...  ...Calibrate standards with researchers and specialists in adjacent fields, converting tacit financial... 
    Hourly pay
    Full time
    Live in
    Relocation
    Relocation package

    SaidGig

    Bay County, FL
    a month ago
  •  ...Apply your dermatology expertise to image-based work that helps develop and evaluate advanced AI models for clinical image interpretation. This is a non-clinical, remote opportunity with no direct patient care, focused on ensuring AI-generated medical outputs reflect real... 
    Hourly pay
    Remote work

    SaidGig

    United States
    11 days ago
  • $150 per hour

     ...Role Overview Apply sell-side equity research judgment to create forecasting and research data that helps evaluate and improve AI models on real financial analysis. You will assess, using only information available at a defined point in time, when a named analyst is... 
    Hourly pay
    For contractors

    SaidGig

    United States
    6 days ago
  • $60 - $90 per hour

     ...Role Overview Help shape how an advanced performance-transfer model evaluates character animation, preserving an actor''s timing, emotion,...  ...evaluation methods, and help build a reliable human-review process for AI-generated performance results. Key Responsibilities... 
    Hourly pay
    Part time

    SaidGig

    Playa Vista, CA
    a month ago
  • $60 - $80 per hour

     ...Role Overview Apply deep retail domain expertise to help build and evaluate foundational generative AI models. You will design realistic retail tasks, produce authoritative solutions grounded in real-world practice, and judge model outputs to improve model correctness... 
    Hourly pay
    Weekday work

    SaidGig

    United States
    2 days ago
  • $150 per hour

     ...authored political forecasting and research data that helps evaluate and improve AI models performing real-world political and financial analysis....  ...campaign strategist, polling or statistical-methodology specialist, or political-risk analyst. A demonstrated record of making... 
    Hourly pay
    Odd job

    SaidGig

    United States
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Evaluation Specialist for AI Model Evaluation. Be the first to apply!