AI Evaluation Specialist
micro1
AI Evaluation Specialist
Role Type: Contractor
Location: Remote (US, CA, UK, IE, AU, NZ)
Micro1 is engaging AI Evaluation Specialists to assess and elevate the quality of AI assistant outputs for an enterprise AI training initiative. In this role, you'll apply your expertise to help train next-generation AI systems. Your work will shape how models learn, reason, and perform through high-quality, real-world input. No prior experience in AI is required — your domain knowledge is what matters.
Scope of Work
- Evaluate AI-generated outputs against detailed rubrics and defined quality standards, focusing on accuracy, relevance, and adherence to guidelines.
- Apply consistent, impartial judgment across a high volume of examples, ensuring a fair and reliable assessment process.
- Identify reasoning gaps, tool-use failures, or logic errors in AI assistant responses, providing actionable feedback for iterative improvement.
- Produce clear, concise written feedback on both strengths and areas for improvement, directly influencing model refinement and AI adoption practices.
- Participate in discussions regarding rubric interpretation and evolving quality standards, contributing to process optimization and best practices.
- Maintain meticulous documentation of evaluations and recommendations, ensuring transparency and traceability in assessment workflows.
Preferred Qualifications
- Experience in grading, quality assurance, editorial review, assessment, annotation, or similar fields demanding careful analysis and detailed feedback.
- Advanced, daily use of AI assistants (such as ChatGPT, Claude, or similar) as an essential work and productivity tool.
- Demonstrated ability to synthesize complex information and communicate findings effectively in writing.
- Background in process improvement, rubric development, or operational quality assessment in an enterprise or educational context.
- Strong critical thinking skills with a focus on consistency, integrity, and fairness in evaluations.
- Comfort working independently on large volumes of similar examples while maintaining high attention to detail.
- Collaborative mindset for sharing insights, discussing ambiguous cases, and refining evaluation criteria as models evolve.
- ...A leading AI organization in Australia is seeking individuals with strong writing and analytical skills to evaluate and improve AI outputs. The ideal candidate must possess the ability to assess emotional nuances and detail while adhering to structured guidelines. Responsibilities...SuggestedImmediate start
$120 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...hour Location: Remote Role Responsibilities Evaluate complex technical tasks using deep language expertise in Scala...SuggestedHourly payWeekly payFull timeContract workFor contractorsSummer workRemote work- ...Join a fast-paced AI evaluation initiative supporting one of the world's leading AI research organizations. We are seeking detail-oriented professionals to evaluate AI-generated outputs by applying structured grading rubrics with precision and consistency. This is...SuggestedTemporary workImmediate start
- ...Join a pioneering AI initiative focused on building next-generation evaluation benchmarks for frontier AI models. We are seeking analytical and technically skilled professionals to identify where advanced AI systems fail in subtle, real-world scenarios. Working in a red...SuggestedFull timeContract workFor contractorsRemote workFlexible hours
- ...About the Opportunity We are looking for experienced professionals with strong expertise in the video game industry to support AI evaluation initiatives. The role involves creating gaming-related evaluation scenarios, reviewing AI-generated responses, and providing...Suggested
- ...About the Opportunity We are looking for experienced entertainment professionals and subject matter experts to support AI evaluation initiatives focused on Movies and TV content. The role involves creating challenging evaluation scenarios, reviewing AI-generated responses...
$25 - $30 per hour
...Bilingual Simplified Chinese AI Evaluation Specialist is a remote Chinese specialist track for evaluating chinese evaluation outputs against native-speaker standards. Reviewers spot fluency, register, and cultural-context errors that automated checks miss, and write structured...For contractorsRemote work10 hours per week$60 per hour
...seeking contributors for a part-time QA project focused on autonomous AI agents. This flexible remote opportunity requires strong... ...familiarity with structured data formats. Candidates will review evaluation tasks, identify inconsistencies, and help define expected AI behaviors...Part timeRemote workFlexible hours$150k - $250k
...About Distyl AI Distyl is an applied AI technology company partnering with the world’s most ambitious institutions to rearchitect... ...What We Are Looking For At Distyl, we build AI systems using Evaluation-Driven Development —an approach where evaluation is not an afterthought...Full timeWork at officeFlexible hours3 days per week- ...experienced frontend and full-stack software developers across eligible global regions to support leading AI labs in training frontier models on frontend code evaluation. This role focuses on leveraging your web development expertise to evaluate and grade AI-generated...Temporary workFor contractorsRemote work
$36 - $72 per hour
...power 25 million job seekers, 1 million+ employers, and 1,600 educational institutions. Handshake AI works directly with frontier AI lab researchers to create evaluations, publish benchmarks, and improve AI models through human expertise. Role Details Location:...Hourly payFull timeMonday to FridayFlexible hours- ...Employment Type: Project-based | Contract We are looking for detail-oriented Image Quality Evaluator for a multilingual AI data Annotation and Transcription Specialists with strong proficiency in English. In this role, you will support AI/ML projects by annotating,...Contract workRemote workWork from homeMonday to FridayDay shift
- ...Summary This is a fully remote, hourly contractor role supporting AI data and language projects on a project-based, flexible hour... ..., and other content to support AI training datasets. LLM evaluation: reviewing AI-generated responses for accuracy, reasoning quality...Hourly payFor contractorsRemote workFlexible hours
$20 - $80 per hour
...Role Overview Help improve next-generation AI systems by supplying precise, real-world evaluation, annotation, and feedback. This remote contractor role focuses on how AI models learn, reason, and perform across diverse subject areas. Key Responsibilities Evaluate...Hourly payContract workFor contractorsRemote work$35 - $120 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...Improve task and rubric quality through structured review. Evaluate the accuracy and depth of AI-generated content to strengthen reasoning...Full timeContract workSummer workRemote work- ...We are looking for an AI Evaluation Scientist to design and execute evaluation processes that ensure our predictive and generative AI systems are accurate, reliable, safe, and aligned with mission requirements. This role is essential for establishing trust in AI solutions...Remote work
$229.9k - $262.4k
Senior Lead AI Engineer (SDK's: Gen AI Evaluation and MCP) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized...Full timePart timeLocal area$130k - $220k
...Artificial Analysis** Artificial Analysis is the leading independent AI benchmarking and insights company. They help engineers,... ...What This Role Actually Is** This role is best described as an AI Evaluation Engineer / Technical Generalist. It is not a traditional software...Full timeWorldwide- ...Responsibilities Run human and automated evaluations across large question sets, including hundreds of thousands or millions of evaluations. Statistically measure factual grounding and before-and-after accuracy lift. Build metrics frameworks that quantify answer...Full time
$141.02k - $204.53k
...expenses. Retirement: Competitive retirement package to secure your future. Responsibilities As the Senior AI/ML Engineer - Validation & Evaluation within AI Validation & Monitoring (AVM), you will provide practice leadership for validation pathways, applied...Full timeWork at officeRemote workFlexible hoursWeekend work$116.04k - $168.29k
...your future. Responsibilities AI/ML Engineers at AI Validation & Monitoring... ..., human-factors and user-experience specialists, data scientists, IT, architecture, patient... ...As the AI/ML Engineer - Validation & Evaluation, with the functional assignment of Validation...Full timeRemote workFlexible hoursWeekend work$49 - $98 per hour
...Bilingual Japanese AI Evaluation Specialist is a remote evaluation track for reviewing japanese generalist evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback...Remote jobFor contractors10 hours per week$30 per hour
...Entry-Level AI Evaluation Specialist is a remote review track for evaluating AI outputs across ai training workflows. Reviewers grade workflow correctness, policy adherence, and stakeholder fit; flag operational risk; and document the right next step so the modeling team...Remote jobFor contractorsWork experience placement10 hours per week- ...Join a pioneering AI initiative focused on building next-generation evaluation benchmarks for frontier AI models. We are seeking analytical and technically skilled professionals to identify where advanced AI systems fail in subtle, real-world scenarios. Working in a red...Full timeContract workFor contractorsRemote workFlexible hours
$30 - $90 per hour
...Role Overview Help improve enterprise AI assistants by evaluating their outputs, identifying weaknesses, and delivering feedback that strengthens how models learn, reason, and perform. Your subject matter expertise and careful judgment are central to this remote contract...Remote jobHourly payContract workFor contractors$163.28k - $236.75k
...Retirement: Competitive retirement package to secure your future. Responsibilities As the Principal AI/ML Engineer - Validation & Evaluation Governance within AI Validation & Monitoring (AVM), you will serve as the enterprise subject-matter authority for...Full timeInterim roleWork at officeRemote workFlexible hoursWeekend work- ...Innodata is expanding its team of technical experts in LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you will build and optimize the technical foundations that power model improvement for foundation model...Full time
$229.9k - $262.4k
...Overview Senior Manager, AI Engineer (Gen AI Platform Services: Agentic AI, Guardrails, Evaluation) Overview : At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader...Full timePart timeLocal area$20 - $80 per hour
...role, you''ll apply your expertise to help train next-generation AI systems. Your work will shape how models learn, reason, and... ...through high-quality, real-world input. Key Responsibilities: Evaluate and score AI-generated responses using well-defined rubrics and...Hourly payContract workRemote work$135 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...Measure accuracy against a held-out set and assess data quality. Evaluate the realistic performance ceiling of the computer vision system...Hourly payFull timeContract workFor contractorsSummer workRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Evaluation Specialist. Be the first to apply!





