Humanities / arts / culture Evaluator
$80 - $120 per hourAI Trainer Jobs
Humanities / arts / culture Evaluator is a remote evaluation track for reviewing humanities / arts / culture evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain. Category: Frontier Model Evaluation · Pay: $80–$120 / hr · Location: Remote — US-eligible · Contractor Humanities / arts / culture Evaluator is a remote evaluation track for reviewing humanities / arts / culture evaluation prompts and responses against AuraOne's quality rubric. About the role Humanities / arts / culture Evaluator is a remote evaluation track for reviewing humanities / arts / culture evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain. AI data reviewers help turn humanities / arts / culture evaluation outputs into auditable labels, rationales, and regression cases for AuraOne Human Data. Review frontier model outputs. Judge benchmark failures and calibrate other evaluators. Responsibilities Evaluate humanities / arts / culture evaluation model outputs against a versioned rubric and assign severity tags for Humanities / arts / culture Evaluator assignments. Compare paired responses and pick the stronger answer with a written rationale. Label hallucinations, instruction-following failures, and unsafe content with structured tags. Capture ambiguous prompts and route them back to the program team for rubric updates. Role details Track Evaluation & annotation Work model Remote · Independent specialist contractor Compensation $80–$120 / hr Eligible from US What you should bring Prior evaluation, annotation, or human-rater experience on humanities / arts / culture evaluation or adjacent content for Humanities / arts / culture Evaluator work. Comfort applying multi-page rubrics consistently across long batches. Clear written reasoning that names the issue and the rubric clause being applied. Strong attention to detail and the ability to flag when a prompt itself is the problem. Reliable async availability for at least 10 hours per week. Role signals Example tasks Compare two humanities / arts / culture evaluation model responses to the same prompt and pick the stronger one with rationale. Tag an unsafe response with the correct policy category and severity. Audit a 50-row batch for rubric consistency and report drift to the program lead. Propose a rubric clarification after spotting a recurring failure mode. Useful experience Background in linguistics, content moderation, or trust & safety review. Experience with inter-rater agreement metrics and calibration cycles. Domain expertise that lets you spot subject-matter errors automated checks miss. Compensation and schedule $80–$120 / hr Expected arrangement: contractor , with program-defined task volume and review pacing. Placement depends on current program demand and reviewer confirmation. Skills used in matching Model output evaluation Rubric-based annotation Severity tagging Inter-rater calibration Humanities / arts / culture evaluation #J-18808-Ljbffr AI Trainer Jobs
- Obsidian is seeking expert Evaluators in Humanities / arts / culture to review AI-generated work products for accuracy, rigor, and domain quality. You will apply subject-matter expertise to grade outputs and provide structured feedback. This is a remote, hourly engagement...SuggestedRemote jobHourly payWork at office
- Mercor is hiring musicians to evaluate generative music AI models in partnership with a leading AI lab. You will assess AI-generated lyrics across genres, rating quality, creativity, prompt adherence, and originality in Malayalam and English. Start date is immediate with...SuggestedHourly payImmediate startShift work
- Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models in partnership with a leading AI lab. You will assess AI-generated music across genres and rate it against detailed quality standards, working in Telugu and English....SuggestedImmediate startRemote workFlexible hours
- Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Portuguese...SuggestedContract workImmediate startFlexible hours
- Mercor is hiring experienced Musicians to evaluate generative musical AI models in partnership with a leading AI lab. You will assess model outputs across genres in your bilingual language, focusing on lyrics, vocal quality, and musicality. Ideal candidates have 3+ years...SuggestedPart time10 hours per weekFlexible hours
- Mercor sucht erfahrene Musikproduzenten und Audioingenieure zur Bewertung generativer Musik‑KI‑Modelle in Zusammenarbeit mit einem führenden KI‑Lab. Sie prüfen AI‑generierte Musik in verschiedenen Genres und bewerten sie nach detaillierten Qualitätsstandards, dabei arbeiten...
$30 per hour
...in the AI space - we are building the biggest pool of quality human data in the world. About Prolific Prolific is not just another... ...Visual Designers to act as Domain Experts for a high-level AI evaluation project. AI models are evolving beyond simple image generation...Remote jobWork from homeFlexible hours- ...realistic UI/UX briefs and test frontier AI tools against professional standards, judging output with rigorous design criteria. You will evaluate interfaces, compare model versions, and document failures across typography, layout, and visual hierarchy while collaborating...Remote jobContract workFlexible hours
$14 per hour
...Compensation: $14/hour Location: Remote Duration: Up to 6 months Commitment: 10+ hours/week Role Responsibilities Evaluate AI model output for lyrics and voice generation across various music categories. Score the quality of musical training data...Contract workSummer workImmediate startRemote work- YO AI Labs is seeking Humanities Evaluation Specialists for a remote contract to support an AI training project. You will research, analyze, and craft challenging humanities questions with high-quality answers, while ensuring clear documentation of reasoning and sources...Remote jobContract work
$37.5 per hour
...generated ad copy for various advertisers looking for placements within Amazon originals. In this role you will be responsible for evaluating advertising copy generated by a large language model (LLM) to ensure that it meets their benchmark in quality, structure, brand...Contract workTemporary workRemote work$160k - $210k
BVAL (Bloomberg's Evaluated Pricing Service) Evaluator - US Agency Structured Products Location New York Business Area Product... ...what makes Bloomberg unique - watch our podcast series for an inside look at our culture, values, and the people behind our success.Price workTemporary workFor contractorsWork experience placement$80 - $120 per hour
...This role is for one of our clients Compensation: $80 - $120 per hour We are hiring expert Evaluators in Special education / IEP to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality....Hourly payContract workFor contractorsWork at officeRemote work- ...Obsidian is hiring expert Evaluators in real estate, hospitality, and events to review AI-generated work products for accuracy and quality. This is a remote position requiring strong expertise in the relevant domains. Ideal candidates should have over 5 years of professional...Work at officeRemote work
$50 - $70 per hour
...Expert (US/Canada) Type: Contract Compensation: $50–$70/hour Location: Remote Role Responsibilities Review and evaluate documents, slides, spreadsheets, and similar materials for quality and correctness. Provide clear, well-reasoned written...Contract workSummer workRemote workFlexible hours$75 per hour
...Role Responsibilities Develop and deliver sociology content through AI training initiatives for postsecondary education. Evaluate and review AI-generated coursework, assignments, and simulated academic papers. Create lectures, learning modules, and case studies...Hourly payContract workRemote work- ...About the role We are hiring expert Evaluators in Data analysis / quantitative readouts to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise...Hourly payWork at officeRemote work
$20 - $30 per hour
A remote-focused technology firm is looking for an individual proficient in evaluating large language model outputs. The role involves assessing AI systems, reviewing workflows, and providing actionable feedback to enhance product quality. Ideal candidates will have strong...Remote jobHourly pay- AI Trainer Jobs is seeking a Hebrew Bilingual Expert for remote evaluation work. You will review Hebrew prompts and responses, compare outputs, and provide structured feedback to support AuraOne’s quality rubric. Responsibilities include labeling edge cases, tracking severity...Remote jobHourly payContract workFor contractorsFlexible hours
$30 - $60 per hour
Web Browsing Evaluator is a remote evaluation track for reviewing web browsing evaluation prompts and responses against AuraOne's quality... ...auditable labels, rationales, and regression cases for AuraOne Human Data. Review browser automation and multi-step search agents....Hourly payFor contractorsRemote work10 hours per week- Obsidian is seeking expert Evaluators for a remote, hourly role focused on reviewing AI-generated work products for quality and accuracy. The successful candidates will leverage their subject matter expertise to assess various documents, spreadsheets, and slide decks....Remote jobHourly payWork at office
- RWS - TrainAI is seeking freelancers to evaluate generative audio in Arabic (Bahrain). This remote, part-time role involves testing AI-generated audio for video and text-to-audio tasks, scoring performance, and providing actionable feedback to improve models. Flexibility...Remote jobPart timeFreelance
- ...Health, part of CVS Health, is seeking a Part-Time Clinician (Nurse Practitioner or Physician Assistant) to provide in-home health evaluations. Visiting members in their own homes, you will review medical history, perform a physical exam and support gaps in care. Visits...Part time
- ...leading digital solutions provider is seeking a Personalized Ads Evaluator for a part-time position. This entry-level role involves... ...relevance to search terms, providing detailed feedback on content and cultural context. Candidates should have excellent communication skills...Remote jobPart time
- ...supported by significant investments and an entrepreneurial drive fueled by our partnership with EQT. Website: Linkedin Job Title: Evaluator - Political Science Location: Remote (USA) Job Type: Contract Key Responsibilities Review and Validate AI Responses: Evaluator...Contract workRemote workWorldwide
- About the role JavaScript and TypeScript AI Evaluator is a remote evaluation track for reviewing javascript and typescript ai evaluation... ...auditable labels, rationales, and regression cases for AuraOne Human Data. Review frontier model outputs. Judge benchmark failures and...Hourly payFor contractorsRemote work10 hours per week
- Cincinnatus LLC is recruiting a senior character animator to help shape evaluation standards for an AI-driven performance transfer model. The role focuses on facial animation, performance capture, and translating complex visuals into precise data captions, with part-time...Part time
- Archangel Health AI is seeking Clinical AI Evaluators to review AI-generated clinical outputs, benchmark diagnostic reasoning, and refine responses to real-world medical queries. You will perform clinical accuracy auditing, RLHF ranking, error and harm identification,...Remote workFlexible hoursShift work
- A leading research accelerator is seeking a contractor to evaluate North American teen humor. The role involves reviewing short-form content, rating based on cultural relevance, and explaining humor dynamics clearly. Ideal candidates are 18 to 19 years old, familiar with...Remote jobContract workFor contractorsFreelanceFlexible hours
- YO AI Labs is seeking Content Writers & Editors to contribute to a customer project focused on text quality and AI evaluation. You will apply your writing and editing expertise to help train and evaluate next-generation AI systems through high-quality input. No prior AI...Remote job
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Humanities / arts / culture Evaluator. Be the first to apply!
- ai evaluator New York, NY
- education evaluator New York, NY
- work from home web search evaluator New York, NY
- evaluator New York, NY
- program evaluator New York, NY
- work from home social media evaluator New York, NY
- clinical evaluator New York, NY
- social media evaluator New York, NY
- quality evaluator New York, NY
- humanities faculty New York, NY



