AI Evaluation Specialist
Weekday
Join a fast-paced AI evaluation initiative supporting one of the world's leading AI research organizations.
We are seeking detail-oriented professionals to evaluate AI-generated outputs by applying structured grading rubrics with precision and consistency.
This is an immediate-start opportunity with onboarding beginning right away. You'll review JSON-formatted model outputs, interpret detailed evaluation guidelines, and provide accurate, rubric-based assessments without the use of AI tools.
If you have experience working with structured data, evaluation frameworks, or quality assurance processes, this role offers an opportunity to contribute directly to improving next-generation AI systems.
Requirements
- Strong ability to read and interpret JSON files and structured data.
- Experience applying detailed rubrics, quality standards, or evaluation frameworks with consistency.
- Excellent analytical thinking and exceptional attention to detail.
- Comfortable reviewing technical or structured documentation and following written guidelines precisely.
- Previous experience in data annotation, quality assurance, content evaluation, AI model assessment, or similar analytical work is highly preferred.
- Strong written communication skills and the ability to justify evaluation decisions clearly.
- Ability to work independently while maintaining accuracy under tight deadlines.
- Must be based in the United States.
Engagement Details
- Immediate onboarding with the opportunity to begin work right away.
- Expected commitment of 20–40 hours per week .
- Ability to complete assigned work within 12–24 hours of receiving tasks.
- Requires access to a desktop or laptop computer. Chromebook devices are not supported.
- Project duration may be extended based on business needs and individual performance.
- ...To support AI development, the part-time AI Evaluation Specialist will review and assess AI-generated outputs for quality and usability while collaborating with teams to refine evaluation standards in a remote contract role. Key responsibilities Review and critically...SuggestedContract workPart timeRemote work
- ...Join a pioneering AI initiative focused on building next-generation evaluation benchmarks for frontier AI models. We are seeking analytical and technically skilled professionals to identify where advanced AI systems fail in subtle, real-world scenarios. Working in a red...SuggestedFull timeContract workFor contractorsRemote workFlexible hours
$20 per hour
A leading AI training company is seeking analytical, detail-oriented individuals to remotely teach AI chatbots. Responsibilities include developing prompts, writing high-quality responses, and evaluating different AI models. The role is ideal for those with experience in...SuggestedHourly payFull timePart timeRemote workFlexible hours$25 - $30 per hour
...independent contractors in the Town of Vermont, Wisconsin to help train AI chatbots. Ideal candidates are proactive individuals who can... ...include creating complex prompts, writing responses, and evaluating AI models. Candidates are required to have strong writing and research...SuggestedHourly payFor contractorsRemote workFlexible hours$25 - $30 per hour
...DataAnnotation is seeking analytical and detail-oriented individuals to help train AI chatbots. In this role, you will develop and evaluate complex prompts while working on your schedule from home. Ideal candidates should have excellent writing and research skills. This...SuggestedHourly payFor contractorsRemote workWork from home$70 per hour
...clients Compensation: $70 per hour Join a cutting-edge AI research initiative focused on improving the quality, accuracy,... ...with exceptional critical thinking and communication skills to evaluate AI-generated responses across a variety of topics. In this role...Hourly payWeekly payContract workFor contractorsRemote workFlexible hours$20 per hour
A leading AI quality assurance company is seeking individuals to develop prompts and test AI chatbots. This role allows for remote... ...creating diverse conversations, providing high-quality responses, and evaluating AI model outputs. A bachelor's degree and experience in business...Hourly payFreelanceRemote workFlexible hours$49 - $98 per hour
...Bilingual Japanese AI Evaluation Specialist is a remote evaluation track for reviewing japanese generalist evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback...For contractorsRemote work10 hours per week$20 per hour
A leading AI training company is seeking remote workers to assist in training AI chatbots. Responsibilities include developing conversations, writing responses, and evaluating AI performance. The ideal candidates are analytical and detail-oriented with strong communication...Hourly payRemote workFlexible hours$20 per hour
A growing AI development company is seeking individuals for a remote position to help train AI chatbots. You will create diverse conversations, provide high-quality responses, and evaluate AI model performance. This role is perfect for those with experience in managing...Hourly payRemote workFlexible hours$20 per hour
A technology company focused on AI is seeking individuals for a remote position teaching AI chatbots. This role allows you to work flexibly, developing prompts and evaluating AI responses while managing your own schedule. Candidates should possess strong writing and research...Hourly payRemote work$20 per hour
A tech firm specializing in AI seeks analytical and detail-oriented individuals for remote work in training AI chatbots. The role involves... ...diverse conversation prompts, writing high-quality answers, and evaluating different AI models. Ideal candidates should have a bachelor's...For contractorsFreelanceRemote workFlexible hours$25 - $30 per hour
...Bilingual German AI Evaluation Specialist is a remote German specialist track for evaluating german evaluation outputs against native-speaker standards. Reviewers spot fluency, register, and cultural-context errors that automated checks miss, and write structured rationale...For contractorsRemote work10 hours per week$70 per hour
...AI Research Initiative Role Compensation: $70 per hour Join a cutting-edge AI research initiative focused on improving the quality... ...with exceptional critical thinking and communication skills to evaluate AI-generated responses across a variety of topics. In this...Hourly payContract workRemote workFlexible hours$45 - $185 per hour
...using Model Context Protocol tools, plugins, and connectors for complex personal workflows. This role supports an AI research initiative focused on evaluating how effectively AI assistants complete personalised, multi-step tasks using connected tools such as Google Drive...Hourly payWeekly payContract workPart timeFor contractorsRemote workTrial period$130 - $180 per hour
Role Description Join a leading AI research initiative focused on advancing healthcare-focused artificial intelligence. We are seeking... ...to contribute their clinical expertise toward training and evaluating next-generation AI models capable of sophisticated medical reasoning...Weekly payContract workPart timeFor contractorsRemote workFlexible hours$60 - $85 per hour
...About the job Remote | Licensed Chemical Engineer & AI Evaluation Specialist - $60-$85/hour We are sharing a specialised part-time consulting opportunity for licensed US chemical engineers with professional experience in process design, process safety, plant operations...Hourly payFull timeContract workPart timeRemote work10 hours per weekFlexible hours- AuraOne is seeking a remote Business Document Expert (Chinese Speaker) to review Chinese evaluation prompts and responses against a quality rubric. You will compare paired outputs, label edge cases, and provide structured feedback to help retrain models. This contractor...Remote jobFor contractors
- ...create role-play scenarios across domains such as travel, financial services, telecoms and technical support, contributing to diverse evaluation datasets. Responsibilities include evaluating performance with metrics on task completion, naturalness, audio comprehension, and...Remote jobContract work
- Dorado is seeking Speech AI Evaluation Specialists to support AI content improvement. This freelance, part-time role is based remotely from Malaysia, with 10+ hours per week and a starting date immediately. You will evaluate Vietnamese-language responses and provide structured...Remote jobPart timeFreelanceImmediate start10 hours per week
- AuraOne seeks a Mandarin Language Expert for a remote evaluation track. Review mandarin generalist prompts and responses against AuraOne's quality rubric, compare paired outputs, and write structured feedback to retrain models. This contractor role requires clear written...Remote jobFor contractors10 hours per week
$80 - $120 per hour
Mercor is seeking a Clinical / Biomedical / Pharma Evaluator to assess AI-generated artifacts and ensure quality. Candidates should have a minimum of 5 years of relevant experience and fluency in English. This role involves evaluating outputs, providing structured feedback...Remote jobHourly payWork at office- A tech company specializing in AI projects is seeking skilled LibreSprite users to assist in evaluating AI-generated visual content. As an independent contractor, you can work flexibly from anywhere, contributing around 5-20 hours per week depending on project needs. Ideal...Remote jobFor contractors
- A leading AI research accelerator is seeking Geospatial Experts to enhance AI systems through evaluations and real-world applications. This entry-level contractor position is fully remote with flexible hours, primarily focusing on geospatial reasoning tasks. Responsibilities...Remote jobFor contractorsFlexible hours
- AuraOne is seeking a Bilingual Korean AI Evaluation Specialist for a remote, contractor role. You will review korean generalist prompts and model outputs, annotate using a versioned rubric, and provide structured feedback to retrain the model. You will compare paired responses...Remote jobFor contractors
- AuraOne is seeking an AI Evaluation Specialist for remote review of AI outputs across operations, focusing on workflow adherence, policy alignment, and stakeholder fit. You will grade tone, escalation logic, and provide clear next steps to help the modeling team improve...Remote jobFor contractors
$40 per hour
# AI Image & Video Evaluation Specialist (PT, MT and CT)$40per hourContractor Remote Last verified 21 Jul 2026Listed on micro1 micro1.aiApplyAffiliate disclosureWork Expert (WE) is an independent publishing and referral website. We are not a recruiter, hiring manager,...Remote work$80 per hour
...Science Expert with Python experience for part-time, remote projects. This role focuses on tasks related to AI systems, including designing problems, evaluating solutions, and validating calculations. Ideal candidates will have a degree in Computer Science, proficiency...Remote jobHourly payPart time- AuraOne seeks a Music Audio Expert - French for a remote evaluation track. You will review french generalist prompts and responses, compare outputs, and provide structured feedback to retrain models. This independent contract role supports quality labeling, edge-case tagging...Remote jobContract work
$150k - $250k
...About Distyl AI Distyl is an applied AI technology company partnering with the world’s most ambitious institutions to rearchitect... ...What We Are Looking For At Distyl, we build AI systems using Evaluation-Driven Development —an approach where evaluation is not an afterthought...Full timeWork at officeFlexible hours3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Evaluation Specialist. Be the first to apply!





