AI Evaluation Specialist
Weekday
Join a fast-paced AI evaluation initiative supporting one of the world's leading AI research organizations.
We are seeking detail-oriented professionals to evaluate AI-generated outputs by applying structured grading rubrics with precision and consistency.
This is an immediate-start opportunity with onboarding beginning right away. You'll review JSON-formatted model outputs, interpret detailed evaluation guidelines, and provide accurate, rubric-based assessments without the use of AI tools.
If you have experience working with structured data, evaluation frameworks, or quality assurance processes, this role offers an opportunity to contribute directly to improving next-generation AI systems.
Requirements
- Strong ability to read and interpret JSON files and structured data.
- Experience applying detailed rubrics, quality standards, or evaluation frameworks with consistency.
- Excellent analytical thinking and exceptional attention to detail.
- Comfortable reviewing technical or structured documentation and following written guidelines precisely.
- Previous experience in data annotation, quality assurance, content evaluation, AI model assessment, or similar analytical work is highly preferred.
- Strong written communication skills and the ability to justify evaluation decisions clearly.
- Ability to work independently while maintaining accuracy under tight deadlines.
- Must be based in the United States.
Engagement Details
- Immediate onboarding with the opportunity to begin work right away.
- Expected commitment of 20–40 hours per week .
- Ability to complete assigned work within 12–24 hours of receiving tasks.
- Requires access to a desktop or laptop computer. Chromebook devices are not supported.
- Project duration may be extended based on business needs and individual performance.
- ...AI Evaluation Specialist Role Type: Contractor Location: Remote (US, CA, UK, IE, AU, NZ) Micro1 is engaging AI Evaluation Specialists to assess and elevate the quality of AI assistant outputs for an enterprise AI training initiative. In this role, you'll apply...SuggestedFor contractorsRemote work
$120 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...hour Location: Remote Role Responsibilities Evaluate complex technical tasks using deep language expertise in Scala...SuggestedHourly payWeekly payFull timeContract workFor contractorsSummer workRemote work- ...Join a pioneering AI initiative focused on building next-generation evaluation benchmarks for frontier AI models. We are seeking analytical and technically skilled professionals to identify where advanced AI systems fail in subtle, real-world scenarios. Working in a red...SuggestedFull timeContract workFor contractorsRemote workFlexible hours
$25 - $30 per hour
...Bilingual Simplified Chinese AI Evaluation Specialist is a remote Chinese specialist track for evaluating chinese evaluation outputs against native-speaker standards. Reviewers spot fluency, register, and cultural-context errors that automated checks miss, and write structured...SuggestedFor contractorsRemote work10 hours per week$60 per hour
...seeking contributors for a part-time QA project focused on autonomous AI agents. This flexible remote opportunity requires strong... ...familiarity with structured data formats. Candidates will review evaluation tasks, identify inconsistencies, and help define expected AI behaviors...SuggestedPart timeRemote workFlexible hours$150k - $250k
...About Distyl AI Distyl is an applied AI technology company partnering with the world’s most ambitious institutions to rearchitect... ...What We Are Looking For At Distyl, we build AI systems using Evaluation-Driven Development —an approach where evaluation is not an afterthought...Full timeWork at officeFlexible hours3 days per week- ...experienced frontend and full-stack software developers across eligible global regions to support leading AI labs in training frontier models on frontend code evaluation. This role focuses on leveraging your web development expertise to evaluate and grade AI-generated...Temporary workFor contractorsRemote work
$20 - $80 per hour
...Role Overview Help improve next-generation AI systems by supplying precise, real-world evaluation, annotation, and feedback. This remote contractor role focuses on how AI models learn, reason, and perform across diverse subject areas. Key Responsibilities Evaluate...Hourly payContract workFor contractorsRemote work- ...Employment Type: Project-based | Contract We are looking for detail-oriented Image Quality Evaluator for a multilingual AI data Annotation and Transcription Specialists with strong proficiency in English. In this role, you will support AI/ML projects by annotating,...Contract workRemote workWork from homeMonday to FridayDay shift
- ...Summary This is a fully remote, hourly contractor role supporting AI data and language projects on a project-based, flexible hour... ..., and other content to support AI training datasets. LLM evaluation: reviewing AI-generated responses for accuracy, reasoning quality...Hourly payFor contractorsRemote workFlexible hours
$35 - $120 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...Improve task and rubric quality through structured review. Evaluate the accuracy and depth of AI-generated content to strengthen reasoning...Full timeContract workSummer workRemote work- ...We are looking for an AI Evaluation Scientist to design and execute evaluation processes that ensure our predictive and generative AI systems are accurate, reliable, safe, and aligned with mission requirements. This role is essential for establishing trust in AI solutions...Remote work
$116.04k - $168.29k
...your future. Responsibilities AI/ML Engineers at AI Validation & Monitoring... ..., human-factors and user-experience specialists, data scientists, IT, architecture, patient... ...As the AI/ML Engineer - Validation & Evaluation, with the functional assignment of Validation...Full timeRemote workFlexible hoursWeekend work$50 per hour
...Generalist AI Evaluation Specialist is a remote review track for evaluating AI outputs across generalists specialist operations workflows. Reviewers grade workflow correctness, policy adherence, and stakeholder fit; flag operational risk; and document the right next step...Remote jobFor contractorsWork experience placement10 hours per week$30 - $90 per hour
...Role Overview Help improve enterprise AI assistants by evaluating their outputs, identifying weaknesses, and delivering feedback that strengthens how models learn, reason, and perform. Your subject matter expertise and careful judgment are central to this remote contract...Remote jobHourly payContract workFor contractors$49 - $98 per hour
...Bilingual Japanese AI Evaluation Specialist is a remote evaluation track for reviewing japanese generalist evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback...Remote jobFor contractors10 hours per week$141.02k - $204.53k
...expenses. Retirement: Competitive retirement package to secure your future. Responsibilities As the Senior AI/ML Engineer - Validation & Evaluation within AI Validation & Monitoring (AVM), you will provide practice leadership for validation pathways, applied...Full timeWork at officeRemote workFlexible hoursWeekend work- ...Join a pioneering AI initiative focused on building next-generation evaluation benchmarks for frontier AI models. We are seeking analytical and technically skilled professionals to identify where advanced AI systems fail in subtle, real-world scenarios. Working in a red...Full timeContract workFor contractorsRemote workFlexible hours
$163.28k - $236.75k
...Retirement: Competitive retirement package to secure your future. Responsibilities As the Principal AI/ML Engineer - Validation & Evaluation Governance within AI Validation & Monitoring (AVM), you will serve as the enterprise subject-matter authority for...Full timeInterim roleWork at officeRemote workFlexible hoursWeekend work- ...Innodata is expanding its team of technical experts in LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you will build and optimize the technical foundations that power model improvement for foundation model...Full time
$20 - $80 per hour
...role, you''ll apply your expertise to help train next-generation AI systems. Your work will shape how models learn, reason, and... ...through high-quality, real-world input. Key Responsibilities: Evaluate and score AI-generated responses using well-defined rubrics and...Hourly payContract workRemote work$135 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...Measure accuracy against a held-out set and assess data quality. Evaluate the realistic performance ceiling of the computer vision system...Hourly payFull timeContract workFor contractorsSummer workRemote work$100k - $150k
...Generative AI Specialist- Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud... ...with the discipline of building reusable design patterns, evaluation frameworks, and developer tooling that scale across many teams...Full timeH1bLocal areaImmediate startRemote workVisa sponsorship$140k - $180k
...AI/ML Integration Specialist Location: San Diego, CA Work Type: Hybrid, 3 Days In-Person (Client Site) Clearance Level: Active DoD Secret... ...for Subject Matter Experts (SMEs), helping them identify, evaluate, and integrate accessible AI/ML capabilities into existing...Contract workWork at officeRemote work3 days per week- ...Financial services is one of the highest-stakes domains for generative AI. Numerical accuracy, regulatory compliance, model risk... ...financial workflows. As an Applied Data Scientist, Financial AI Evaluation & Datasets , you own the design, measurement quality, and...Full timeShift work
- ...engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the... ...responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems...Full timeShift work
- ...creative development firm that addresses clients' most pressing needs and challenges. We are currently looking for a Data Scientist - AI Evaluation Analytics Location: San Diego, CA (Remote but resource should be in PST time zone) Position: Data Scientist - AI...Remote work
$20 per hour
...A technology firm is seeking DataAnnotators to create diverse conversations and evaluate AI models. The role offers remote work and the flexibility to choose projects, allowing successful candidates to work between 5-40 hours per week. Applicants should be fluent in English...Hourly payRemote work$100 - $150 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...exfiltration , ransomware , worms , and exploits . Evaluate POC exploit development to determine boundaries between security...Hourly payWeekly payFull timeContract workFor contractorsSummer workRemote work- ...POSITION SUMMARY: Natera is seeking a Senior Agentic AI Engineer to design, build, and productionize intelligent agent workflows... ...approach and delivering measurable improvements through careful evaluation and phased rollout. The ideal candidate combines strong hands...Full timeImmediate startWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Evaluation Specialist. Be the first to apply!





