Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Evaluation Specialist [Remote]

$30 - $90 per hour

SaidGig

Remote
  • Remote job

Role Overview

Help improve enterprise AI assistants by evaluating their outputs, identifying weaknesses, and delivering feedback that strengthens how models learn, reason, and perform. Your subject matter expertise and careful judgment are central to this remote contract role. Prior AI experience is not required.

Key Responsibilities

  • Assess AI-generated responses using detailed rubrics and quality standards for accuracy, relevance, and guideline adherence.
  • Apply impartial, consistent judgment across a high volume of examples.
  • Identify reasoning gaps, logic errors, and tool-use failures, then provide actionable feedback for improvement.
  • Write concise feedback that highlights both strengths and opportunities to improve.
  • Participate in discussions about rubric interpretation, evolving quality standards, process improvements, and best practices.
  • Maintain thorough evaluation records and recommendations to support transparent, traceable assessment workflows.

Qualifications

  • Experience in grading, quality assurance, editorial review, assessment, annotation, or a comparable field requiring detailed analysis and feedback.
  • Advanced daily use of AI assistants, such as ChatGPT or Claude, as a work or productivity tool.
  • Ability to synthesize complex information and communicate findings clearly in writing.
  • Experience with process improvement, rubric development, or operational quality assessment in an enterprise or educational setting.
  • Strong critical thinking, consistency, integrity, fairness, and attention to detail.
  • Comfort working independently through large volumes of similar examples.
  • A collaborative approach to sharing insights, resolving ambiguous cases, and refining evaluation criteria.

Work Terms

  • Independent contractor engagement.
  • Remote work is available to candidates located in the United States, Canada, United Kingdom, Ireland, Australia, or New Zealand.

Compensation

$30 to $90 per hour.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the AI Evaluation Specialist [Remote] in Remote vacancy
  •  ...AI Evaluation Specialist Role Type: Contractor Location: Remote (US, CA, UK, IE, AU, NZ) Micro1 is engaging AI Evaluation Specialists to assess and elevate the quality of AI assistant outputs for an enterprise AI training initiative. In this role, you'll apply... 
    Suggested
    For contractors
    Remote work

    micro1

    United States
    1 day ago
  • $120 per hour

     ...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors...  ...hour Location: Remote Role Responsibilities Evaluate complex technical tasks using deep language expertise in Scala... 
    Suggested
    Hourly pay
    Weekly pay
    Full time
    Contract work
    For contractors
    Summer work
    Remote work

    Mercor

    Remote
    20 hours ago
  •  ...Join a fast-paced AI evaluation initiative supporting one of the world's leading AI research organizations. We are seeking detail-oriented professionals to evaluate AI-generated outputs by applying structured grading rubrics with precision and consistency. This is... 
    Suggested
    Temporary work
    Immediate start

    Weekday

    Remote
    a month ago
  •  ...Join a pioneering AI initiative focused on building next-generation evaluation benchmarks for frontier AI models. We are seeking analytical and technically skilled professionals to identify where advanced AI systems fail in subtle, real-world scenarios. Working in a red... 
    Suggested
    Full time
    Contract work
    For contractors
    Remote work
    Flexible hours

    Weekday

    Remote
    a month ago
  • $25 - $30 per hour

     ...Bilingual Simplified Chinese AI Evaluation Specialist is a remote Chinese specialist track for evaluating chinese evaluation outputs against native-speaker standards. Reviewers spot fluency, register, and cultural-context errors that automated checks miss, and write structured... 
    Suggested
    For contractors
    Remote work
    10 hours per week

    AuraOne Human Data

    Remote
    a month ago
  • $60 per hour

     ...seeking contributors for a part-time QA project focused on autonomous AI agents. This flexible remote opportunity requires strong...  ...familiarity with structured data formats. Candidates will review evaluation tasks, identify inconsistencies, and help define expected AI behaviors... 
    Part time
    Remote work
    Flexible hours

    Mind Rift

    Kansas City, MO
    3 days ago
  • $150k - $250k

     ...About Distyl AI Distyl is an applied AI technology company partnering with the world’s most ambitious institutions to rearchitect...  ...What We Are Looking For At Distyl, we build AI systems using Evaluation-Driven Development —an approach where evaluation is not an afterthought... 
    Full time
    Work at office
    Flexible hours
    3 days per week

    Distyl Ai

    Remote
    20 hours ago
  •  ...experienced frontend and full-stack software developers across eligible global regions to support leading AI labs in training frontier models on frontend code evaluation. This role focuses on leveraging your web development expertise to evaluate and grade AI-generated... 
    Temporary work
    For contractors
    Remote work

    Mercor

    Remote
    2 days ago
  • $20 - $80 per hour

     ...Role Overview Help improve next-generation AI systems by supplying precise, real-world evaluation, annotation, and feedback. This remote contractor role focuses on how AI models learn, reason, and perform across diverse subject areas. Key Responsibilities Evaluate... 
    Hourly pay
    Contract work
    For contractors
    Remote work

    SaidGig

    United States
    a month ago
  •  ...Summary This is a fully remote, hourly contractor role supporting AI data and language projects on a project-based, flexible hour...  ..., and other content to support AI training datasets. LLM evaluation: reviewing AI-generated responses for accuracy, reasoning quality... 
    Hourly pay
    For contractors
    Remote work
    Flexible hours

    CNTXT AI

    Brooklyn, NY
    8 days ago
  •  ...Employment Type: Project-based | Contract  We are looking for detail-oriented Image Quality Evaluator for a multilingual AI data Annotation and Transcription Specialists with strong proficiency in English. In this role, you will support AI/ML projects by annotating,... 
    Contract work
    Remote work
    Work from home
    Monday to Friday
    Day shift

    iMerit Technologies

    San Jose, CA
    a month ago
  • $35 - $120 per hour

     ...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors...  ...Improve task and rubric quality through structured review. Evaluate the accuracy and depth of AI-generated content to strengthen reasoning... 
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    Remote
    20 hours ago
  •  ...We are looking for an AI Evaluation Scientist to design and execute evaluation processes that ensure our predictive and generative AI systems are accurate, reliable, safe, and aligned with mission requirements. This role is essential for establishing trust in AI solutions... 
    Remote work

    Convergenz

    United States
    1 day ago
  • $116.04k - $168.29k

     ...your future. Responsibilities AI/ML Engineers at AI Validation & Monitoring...  ..., human-factors and user-experience specialists, data scientists, IT, architecture, patient...  ...As the AI/ML Engineer - Validation & Evaluation, with the functional assignment of Validation... 
    Full time
    Remote work
    Flexible hours
    Weekend work

    Mayo Clinic

    Rochester, MN
    4 days ago
  • $141.02k - $204.53k

     ...expenses. Retirement: Competitive retirement package to secure your future. Responsibilities As the Senior AI/ML Engineer - Validation & Evaluation within AI Validation & Monitoring (AVM), you will provide practice leadership for validation pathways, applied... 
    Full time
    Work at office
    Remote work
    Flexible hours
    Weekend work

    Mayo Clinic

    Rochester, MN
    4 days ago
  • $30 per hour

     ...Entry-Level AI Evaluation Specialist is a remote review track for evaluating AI outputs across ai training workflows. Reviewers grade workflow correctness, policy adherence, and stakeholder fit; flag operational risk; and document the right next step so the modeling team... 
    Remote job
    For contractors
    Work experience placement
    10 hours per week

    AuraOne Human Data

    Remote
    2 days ago
  • $49 - $98 per hour

     ...Bilingual Japanese AI Evaluation Specialist is a remote evaluation track for reviewing japanese generalist evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback... 
    Remote job
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    a month ago
  •  ...Join a pioneering AI initiative focused on building next-generation evaluation benchmarks for frontier AI models. We are seeking analytical and technically skilled professionals to identify where advanced AI systems fail in subtle, real-world scenarios. Working in a red... 
    Full time
    Contract work
    For contractors
    Remote work
    Flexible hours

    Weekday

    Remote
    a month ago
  • $163.28k - $236.75k

     ...Retirement: Competitive retirement package to secure your future. Responsibilities As the Principal AI/ML Engineer - Validation & Evaluation Governance within AI Validation & Monitoring (AVM), you will serve as the enterprise subject-matter authority for... 
    Full time
    Interim role
    Work at office
    Remote work
    Flexible hours
    Weekend work

    Mayo Clinic

    Rochester, MN
    5 days ago
  •  ...Innodata is expanding its team of technical experts in LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you will build and optimize the technical foundations that power model improvement for foundation model... 
    Full time

    Innodata

    Remote
    a month ago
  • $135 per hour

     ...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors...  ...Measure accuracy against a held-out set and assess data quality. Evaluate the realistic performance ceiling of the computer vision system... 
    Hourly pay
    Full time
    Contract work
    For contractors
    Summer work
    Remote work

    Mercor

    Remote
    20 hours ago
  • $20 - $80 per hour

     ...role, you''ll apply your expertise to help train next-generation AI systems. Your work will shape how models learn, reason, and...  ...through high-quality, real-world input. Key Responsibilities: Evaluate and score AI-generated responses using well-defined rubrics and... 
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    20 hours ago
  • $100k - $150k

     ...Generative AI Specialist- Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud...  ...with the discipline of building reusable design patterns, evaluation frameworks, and developer tooling that scale across many teams... 
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    United States
    3 days ago
  • $140k - $180k

     ...AI/ML Integration Specialist Location: San Diego, CA Work Type: Hybrid, 3 Days In-Person (Client Site) Clearance Level: Active DoD Secret...  ...for Subject Matter Experts (SMEs), helping them identify, evaluate, and integrate accessible AI/ML capabilities into existing... 
    Contract work
    Work at office
    Remote work
    3 days per week

    The Marlin Alliance

    San Diego, CA
    1 day ago
  •  ...Financial services is one of the highest-stakes domains for generative AI. Numerical accuracy, regulatory compliance, model risk...  ...financial workflows. As an Applied Data Scientist, Financial AI Evaluation & Datasets , you own the design, measurement quality, and... 
    Full time
    Shift work

    Innodata

    Remote
    a month ago
  •  ...engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the...  ...responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems... 
    Full time
    Shift work

    Innodata

    Remote
    a month ago
  •  ...creative development firm that addresses clients' most pressing needs and challenges. We are currently looking for a Data Scientist - AI Evaluation Analytics Location: San Diego, CA (Remote but resource should be in PST time zone) Position: Data Scientist - AI... 
    Remote work

    Momento USA

    San Diego, CA
    2 days ago
  • $20 per hour

     ...A technology firm is seeking DataAnnotators to create diverse conversations and evaluate AI models. The role offers remote work and the flexibility to choose projects, allowing successful candidates to work between 5-40 hours per week. Applicants should be fluent in English... 
    Hourly pay
    Remote work

    SupportFinity

    United States
    4 days ago
  • $150k - $210k

     ...Enterprise Knowledge (EK) is hiring a full-time Semantic Data and AI Engineer to join our growing Semantic Engineering and AI...  ...retrieval pipeline architecture, embedding strategies, and response evaluation Contribute to agentic AI solution design and implementation,... 
    Full time
    For contractors
    H1b
    Work at office
    Local area
    Remote work

    Enterprise Knowledge, Llc

    United States
    20 hours ago
  • $60 - $90 per hour

     ...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors...  ...expert-level prompts across specialized cybersecurity topics. Evaluate and annotate model responses for technical accuracy, helpfulness... 
    Full time
    Contract work
    Summer work
    Immediate start
    Remote work

    Mercor

    Remote
    20 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Evaluation Specialist [Remote]. Be the first to apply!