Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Remote AI Agent Evaluation Specialist

$60 per hour

Mind Rift

An innovative tech company in Missouri is seeking contributors for a part-time QA project focused on autonomous AI agents. This flexible remote opportunity requires strong analytical and critical thinking skills, alongside attention to detail and familiarity with structured data formats. Candidates will review evaluation tasks, identify inconsistencies, and help define expected AI behaviors. The role offers competitive pay up to $60/hour and is ideal for those in academia or consulting with an interest in AI. #J-18808-Ljbffr

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Remote AI Agent Evaluation Specialist in Kansas City, MO vacancy
  • $125 per hour

     ...QGIS specialists leverage their expertise in geographic information systems...  ...data analysis to support AI research through flexible,...  ...based work. This role involves evaluating AI-generated content and providing...  ...Work Terms This position is remote and allows for asynchronous... 
    Remote work
    Flexible hours

    SaidGig

    United States
    4 days ago
  • $20 - $30 per hour

     ...As an Image Evaluation Generalist, you will play a pivotal role in supporting an image assessment...  ...to the training of next-generation AI systems. Your expertise will be essential...  ...ambiguities or gaps. Operate independently in a remote environment, delivering high-quality,... 
    Remote work
    Hourly pay
    For contractors

    SaidGig

    Indiana
    a month ago
  •  ...technology company is looking for a detail-oriented individual to design structured evaluation scenarios for AI agents. This entry-level part-time role allows you to contribute remotely, creating test cases for LLM-based agents. The ideal candidate will have a background... 
    Remote work
    Part time

    Mind Rift

    Brooklyn, NY
    1 day ago
  • $80 per hour

     ...-time opportunity focused on quality assurance for autonomous AI agents. You will analyze complex systems, review tasks for logic, and...  ...analytical and detail-oriented skills, with experience in policy evaluation or logic puzzles preferred. Compensation can reach up to $80/... 
    Remote work
    Part time
    Flexible hours

    Mind Rift

    Raleigh, NC
    5 days ago
  • $80 per hour

    A leading AI consultancy is seeking a detail-oriented individual to design evaluation scenarios for LLM-based agents. This part-time, remote role allows you to create structured test cases while working flexibly around your commitments. The ideal candidate holds a relevant... 
    Remote work
    Part time

    Mindrift

    Houston, TX
    3 days ago
  • $80 per hour

    A technology firm is seeking QAs for autonomous AI agents to validate and improve task structures within a new project. Candidates will...  ...thinkers with strong attention to detail and experience in evaluating scenarios. This flexible, project-based role offers competitive... 
    Remote job
    Flexible hours

    Mindrift

    Providence, RI
    3 days ago
  • $60 per hour

     ...A leading AI firm in Austin is looking for QA experts to validate...  ...and improve AI systems. This remote, freelance role requires...  ...detail. Candidates will review AI evaluation tasks, identify inconsistencies...  ...define expected behaviors for agents. Ideal applicants have experience... 
    Remote work
    Freelance

    Mind Rift

    Austin, TX
    5 days ago
  •  ...SWE Agent Evaluation Specialist is a remote engineering review track for evaluating production code, debugging traces, and developer-facing AI outputs against real-world correctness standards. Reviewers reproduce failures, write the unit test the model should have written... 
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    6 hours ago
  • $80 per hour

     ...Do The Mindrift platform connects specialists with AI projects from major tech...  ...design realistic and structured evaluation scenarios for LLM‑based agents. You'll create test cases that simulate...  ...needs Take part in a flexible, remote, freelance project that fits around... 
    Remote work
    Freelance
    Flexible hours

    Mindrift

    Jackson, MS
    4 days ago
  • $80 per hour

     ...intelligence to ethically shape the future of AI. What We Do The Mindrift platform,...  ...servers and internal tools for running and evaluating agent behavior. You’ll implement base methods...  ...project needs Take part in a flexible, remote, freelance project that fits around your... 
    Remote work
    Part time
    Freelance
    Flexible hours

    Mind Rift

    New York, NY
    5 days ago
  • $80 - $120 per hour

     ...elite creative and technical talent with leading AI research labs. Headquartered in San Francisco...  .... Position: Process improvement / SOPs Evaluator Type: Contract Compensation: $80–$120/hour Location: Remote Role Responsibilities Evaluate AI-... 
    Remote work
    Contract work
    Summer work
    Work at office

    Mercor

    Houston, TX
    5 days ago
  •  ...IT Helpdesk Agent Evaluator is a remote evaluation track for reviewing it helpdesk agent evaluation prompts...  ...to retrain. Why this role matters AI data reviewers help turn it helpdesk...  ...— US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR... 
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    6 hours ago
  • $60 - $90 per hour

     ...Join a leading AI lab''s cutting-edge GenAI team to be at the core...  ...the next generation of agentic evaluation benchmarks for frontier models and needs specialists who can find where those models...  ...benchmark tasks. This role is fully remote within the United States, at... 
    Remote work
    Hourly pay
    Full time
    Contract work
    Part time
    Freelance

    SaidGig

    United States
    3 days ago
  • $15 - $25 per hour

     ...Role Overview As a Health Care Specialist, you will leverage your...  ...training of next-generation AI systems. Your insights will be...  ...exceptional attention to detail when evaluating medical information....  ...multitask and work efficiently in a remote environment. Strong... 
    Remote work
    Hourly pay
    Part time
    For contractors
    Flexible hours

    SaidGig

    United States
    4 days ago
  • $1,750 - $2,150 per month

     ...connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our...  ...,750–$2,150 per completed task Location: Remote Role Responsibilities Review and evaluate AI-generated outputs related to threat analysis,... 
    Remote work
    Hourly pay
    Full time
    Contract work
    Summer work

    Mercor

    Remote
    7 hours ago
  • $15 - $25 per hour

     ...As a Legal Specialist, you will leverage your legal expertise to contribute...  ...training of next-generation AI systems. Your insights will...  ...to support structured legal evaluations. Prepare concise written...  ...priorities and deadlines in a remote environment. Strong... 
    Remote work
    Hourly pay
    Contract work
    Part time
    For contractors

    SaidGig

    United States
    3 days ago
  • $25 - $30 per hour

     ...DataAnnotation is seeking a Certified Coding Specialist (CCS) to join their team and help train AI models. This role requires expertise in healthcare to evaluate AI performance and improve model quality. Applicants should have fluency in English and a medical or healthcare... 
    Remote work
    Hourly pay
    For contractors

    DataAnnotation

    Boston, MA
    4 days ago
  •  ...To support program evaluation and advanced analytics, the full-time Program Evaluation Specialist will lead the design and implementation of evaluations, utilizing quantitative and qualitative methodologies to generate actionable insights for health and civilian programs... 
    Remote work
    Full time

    Virtual Vocations Inc

    United States
    3 days ago
  • $152k - $240k

     ...control spend effortlessly. Brex’s AI-native automation and world-...  ...’ll do We're building AI agents to automate and augment...  ...four weeks per year of fully remote work! Responsibilities:...  ..., and data sources. Define evaluation frameworks, success metrics, and... 
    Remote work
    Full time
    Work at office
    Work from home

    Brex Inc.

    New York, NY
    7 hours ago
  •  ...drive revenue performance. About the Role AI agents are becoming central to how CaptivateIQ...  ...SDK, LLM orchestration layer, and the evaluation and observability infrastructure that underpins...  ...in‑office 3 days per week) – Austin, TX Remote – Raleigh, NC Remote – Nashville, TN... 
    Remote work
    Work at office
    Flexible hours
    Shift work
    3 days per week

    CAPTIVATEIQ INC

    Austin, TX
    5 days ago
  • $200k - $320k

     ...technical depth and a passion for AI-driven product development....  .... This is a full-time remote opportunity, with preference for...  ...production-ready LLM pipelines and AI agent systems Develop AI-driven...  ...new product initiatives Evaluate and recommend AI architectures... 
    Remote work
    Full time
    Visa sponsorship

    Virco Talent

    Remote
    7 hours ago
  • ManpowerGroup is seeking an AI Enablement Lead to accelerate AI adoption across non‑engineering teams. This role involves identifying...  ...and strong communication skills. Responsibilities include evaluating processes for AI opportunities, aligning initiatives with governance... 
    Remote job

    ManpowerGroup

    Jacksonville, FL
    1 day ago
  • $100k - $115k

     ...is seeking a detail-oriented Program Evaluation Specialist to support program evaluation, implementation...  ...-grade platforms and mission-ready AI to federal agencies at commercial speed...  ...security clearances, due to the nature of the work. Job Locations US-Remote
    Remote work
    Contract work

    LMI

    United States
    8 days ago
  • $1,150 - $1,450 per month

     ...technical talent with leading AI research labs. Headquartered in...  ...task Location: Remote Role Responsibilities Build...  ..., grant applications, program evaluations, health-assessment reports, policy...  ...stakeholders to challenge frontier AI agents. Collaborate with other... 
    Remote work
    Hourly pay
    Full time
    Contract work
    Summer work

    Mercor

    Philadelphia, PA
    8 days ago
  • $202.5k - $247.5k

     ...sharing localhost or running AI workloads in production. We...  ...worth your time.   About the Agent Team Our Agent team...  ...AWS. Engineers develop by using remote development tools and/or ssh to...  ...and actual compensation will be evaluated based on factors including,... 
    Remote work
    Permanent employment
    Full time
    Work at office
    Local area
    Immediate start
    Home office
    Flexible hours

    Ngrok

    Remote
    7 hours ago
  • $80 - $120 per hour

     ...Mercor is seeking a Nonprofit / philanthropy / community programs Evaluator to work remotely. The role involves evaluating AI-generated artifacts, providing structured feedback, and ensuring quality in documents, spreadsheets, and presentations. The ideal candidate should... 
    Remote work
    Hourly pay
    Contract work
    Work at office

    Mercor Inc

    New York, NY
    1 day ago
  • $60 - $90 per hour

     ...creative and technical talent with leading AI research labs. Headquartered in San...  ...Compensation: $60–$90/hour Location: Remote Commitment: 15–40 hours/week...  ...specialized cybersecurity topics. Evaluate and annotate model responses for technical... 
    Remote work
    Full time
    Contract work
    Summer work
    Immediate start

    Mercor

    Remote
    7 hours ago
  •  ...Prolific is seeking Product Designers and UX Specialists to join our Expert Network, contributing to the training and evaluation of cutting-edge AI models. This role requires expertise in usability, design systems, and user research, providing essential feedback on design... 
    Remote work
    Work from home
    Flexible hours

    Prolific

    Tucson, AZ
    4 days ago
  • $80 - $150 per hour

     ...Prolific is seeking Medical Doctors to train and evaluate AI models. You'll review AI-generated clinical responses and evaluate their accuracy and appropriateness, earning competitive pay of $80-$150 per hour depending on your skills and experience. Ideal candidates must... 
    Remote work
    Hourly pay
    Work from home
    Flexible hours

    Prolific

    Portland, OR
    4 days ago
  •  ...Kraken app. As a fully remote company, we have...  ...architect and builder of the AI-native finance...  ...Finance operations – Evaluate how financial work currently...  ...engines, and MCP or similar agent coordination layers....  ...without relying on a single specialist. Train finance... 
    Remote work
    Local area

    Kraken

    Poland, NY
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Remote AI Agent Evaluation Specialist. Be the first to apply!