Remote AI Agent Evaluation Specialist
$60 per hourMind Rift
An innovative tech company in Missouri is seeking contributors for a part-time QA project focused on autonomous AI agents. This flexible remote opportunity requires strong analytical and critical thinking skills, alongside attention to detail and familiarity with structured data formats. Candidates will review evaluation tasks, identify inconsistencies, and help define expected AI behaviors. The role offers competitive pay up to $60/hour and is ideal for those in academia or consulting with an interest in AI. #J-18808-Ljbffr
$125 per hour
...QGIS specialists leverage their expertise in geographic information systems... ...data analysis to support AI research through flexible,... ...based work. This role involves evaluating AI-generated content and providing... ...Work Terms This position is remote and allows for asynchronous...Remote workFlexible hours$20 - $30 per hour
...As an Image Evaluation Generalist, you will play a pivotal role in supporting an image assessment... ...to the training of next-generation AI systems. Your expertise will be essential... ...ambiguities or gaps. Operate independently in a remote environment, delivering high-quality,...Remote workHourly payFor contractors- ...technology company is looking for a detail-oriented individual to design structured evaluation scenarios for AI agents. This entry-level part-time role allows you to contribute remotely, creating test cases for LLM-based agents. The ideal candidate will have a background...Remote workPart time
$80 per hour
...-time opportunity focused on quality assurance for autonomous AI agents. You will analyze complex systems, review tasks for logic, and... ...analytical and detail-oriented skills, with experience in policy evaluation or logic puzzles preferred. Compensation can reach up to $80/...Remote workPart timeFlexible hours$80 per hour
A leading AI consultancy is seeking a detail-oriented individual to design evaluation scenarios for LLM-based agents. This part-time, remote role allows you to create structured test cases while working flexibly around your commitments. The ideal candidate holds a relevant...Remote workPart time$80 per hour
A technology firm is seeking QAs for autonomous AI agents to validate and improve task structures within a new project. Candidates will... ...thinkers with strong attention to detail and experience in evaluating scenarios. This flexible, project-based role offers competitive...Remote jobFlexible hours$60 per hour
...A leading AI firm in Austin is looking for QA experts to validate... ...and improve AI systems. This remote, freelance role requires... ...detail. Candidates will review AI evaluation tasks, identify inconsistencies... ...define expected behaviors for agents. Ideal applicants have experience...Remote workFreelance- ...SWE Agent Evaluation Specialist is a remote engineering review track for evaluating production code, debugging traces, and developer-facing AI outputs against real-world correctness standards. Reviewers reproduce failures, write the unit test the model should have written...Remote jobHourly payFor contractors10 hours per week
$80 per hour
...Do The Mindrift platform connects specialists with AI projects from major tech... ...design realistic and structured evaluation scenarios for LLM‑based agents. You'll create test cases that simulate... ...needs Take part in a flexible, remote, freelance project that fits around...Remote workFreelanceFlexible hours$80 per hour
...intelligence to ethically shape the future of AI. What We Do The Mindrift platform,... ...servers and internal tools for running and evaluating agent behavior. You’ll implement base methods... ...project needs Take part in a flexible, remote, freelance project that fits around your...Remote workPart timeFreelanceFlexible hours$80 - $120 per hour
...elite creative and technical talent with leading AI research labs. Headquartered in San Francisco... .... Position: Process improvement / SOPs Evaluator Type: Contract Compensation: $80–$120/hour Location: Remote Role Responsibilities Evaluate AI-...Remote workContract workSummer workWork at office- ...IT Helpdesk Agent Evaluator is a remote evaluation track for reviewing it helpdesk agent evaluation prompts... ...to retrain. Why this role matters AI data reviewers help turn it helpdesk... ...— US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR...Remote jobHourly payFor contractors10 hours per week
$60 - $90 per hour
...Join a leading AI lab''s cutting-edge GenAI team to be at the core... ...the next generation of agentic evaluation benchmarks for frontier models and needs specialists who can find where those models... ...benchmark tasks. This role is fully remote within the United States, at...Remote workHourly payFull timeContract workPart timeFreelance$15 - $25 per hour
...Role Overview As a Health Care Specialist, you will leverage your... ...training of next-generation AI systems. Your insights will be... ...exceptional attention to detail when evaluating medical information.... ...multitask and work efficiently in a remote environment. Strong...Remote workHourly payPart timeFor contractorsFlexible hours$1,750 - $2,150 per month
...connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our... ...,750–$2,150 per completed task Location: Remote Role Responsibilities Review and evaluate AI-generated outputs related to threat analysis,...Remote workHourly payFull timeContract workSummer work$15 - $25 per hour
...As a Legal Specialist, you will leverage your legal expertise to contribute... ...training of next-generation AI systems. Your insights will... ...to support structured legal evaluations. Prepare concise written... ...priorities and deadlines in a remote environment. Strong...Remote workHourly payContract workPart timeFor contractors$25 - $30 per hour
...DataAnnotation is seeking a Certified Coding Specialist (CCS) to join their team and help train AI models. This role requires expertise in healthcare to evaluate AI performance and improve model quality. Applicants should have fluency in English and a medical or healthcare...Remote workHourly payFor contractors- ...To support program evaluation and advanced analytics, the full-time Program Evaluation Specialist will lead the design and implementation of evaluations, utilizing quantitative and qualitative methodologies to generate actionable insights for health and civilian programs...Remote workFull time
$152k - $240k
...control spend effortlessly. Brex’s AI-native automation and world-... ...’ll do We're building AI agents to automate and augment... ...four weeks per year of fully remote work! Responsibilities:... ..., and data sources. Define evaluation frameworks, success metrics, and...Remote workFull timeWork at officeWork from home- ...drive revenue performance. About the Role AI agents are becoming central to how CaptivateIQ... ...SDK, LLM orchestration layer, and the evaluation and observability infrastructure that underpins... ...in‑office 3 days per week) – Austin, TX Remote – Raleigh, NC Remote – Nashville, TN...Remote workWork at officeFlexible hoursShift work3 days per week
$200k - $320k
...technical depth and a passion for AI-driven product development.... .... This is a full-time remote opportunity, with preference for... ...production-ready LLM pipelines and AI agent systems Develop AI-driven... ...new product initiatives Evaluate and recommend AI architectures...Remote workFull timeVisa sponsorship- ManpowerGroup is seeking an AI Enablement Lead to accelerate AI adoption across non‑engineering teams. This role involves identifying... ...and strong communication skills. Responsibilities include evaluating processes for AI opportunities, aligning initiatives with governance...Remote job
$100k - $115k
...is seeking a detail-oriented Program Evaluation Specialist to support program evaluation, implementation... ...-grade platforms and mission-ready AI to federal agencies at commercial speed... ...security clearances, due to the nature of the work. Job Locations US-RemoteRemote workContract work$1,150 - $1,450 per month
...technical talent with leading AI research labs. Headquartered in... ...task Location: Remote Role Responsibilities Build... ..., grant applications, program evaluations, health-assessment reports, policy... ...stakeholders to challenge frontier AI agents. Collaborate with other...Remote workHourly payFull timeContract workSummer work$202.5k - $247.5k
...sharing localhost or running AI workloads in production. We... ...worth your time. About the Agent Team Our Agent team... ...AWS. Engineers develop by using remote development tools and/or ssh to... ...and actual compensation will be evaluated based on factors including,...Remote workPermanent employmentFull timeWork at officeLocal areaImmediate startHome officeFlexible hours$80 - $120 per hour
...Mercor is seeking a Nonprofit / philanthropy / community programs Evaluator to work remotely. The role involves evaluating AI-generated artifacts, providing structured feedback, and ensuring quality in documents, spreadsheets, and presentations. The ideal candidate should...Remote workHourly payContract workWork at office$60 - $90 per hour
...creative and technical talent with leading AI research labs. Headquartered in San... ...Compensation: $60–$90/hour Location: Remote Commitment: 15–40 hours/week... ...specialized cybersecurity topics. Evaluate and annotate model responses for technical...Remote workFull timeContract workSummer workImmediate start- ...Prolific is seeking Product Designers and UX Specialists to join our Expert Network, contributing to the training and evaluation of cutting-edge AI models. This role requires expertise in usability, design systems, and user research, providing essential feedback on design...Remote workWork from homeFlexible hours
$80 - $150 per hour
...Prolific is seeking Medical Doctors to train and evaluate AI models. You'll review AI-generated clinical responses and evaluate their accuracy and appropriateness, earning competitive pay of $80-$150 per hour depending on your skills and experience. Ideal candidates must...Remote workHourly payWork from homeFlexible hours- ...Kraken app. As a fully remote company, we have... ...architect and builder of the AI-native finance... ...Finance operations – Evaluate how financial work currently... ...engines, and MCP or similar agent coordination layers.... ...without relying on a single specialist. Train finance...Remote workLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Remote AI Agent Evaluation Specialist. Be the first to apply!
- operations agent Kansas City, MO
- cruise agent Kansas City, MO
- telemarketer - state farm agent team member Kansas City, MO
- import export agent Kansas City, MO
- state farm agent Kansas City, MO
- special agent Kansas City, MO
- remote chat agent Kansas City, MO
- executive protection agent Kansas City, MO
- commissioning agent Kansas City, MO
- agent Kansas City, MO




