Music Generation Quality Evaluator [Remote]
AuraOne Human Data
- Remote job
Music Generation Quality Evaluator is a remote evaluation track for reviewing music generation quality evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain.
Why this role matters
AI data reviewers help turn music generation quality evaluation outputs into auditable labels, rationales, and regression cases for AuraOne Human Data.
Responsibilities
- Evaluate music generation quality evaluation model outputs against a versioned rubric and assign severity tags for Music Generation Quality Evaluator assignments.
- Compare paired responses and pick the stronger answer with a written rationale.
- Label hallucinations, instruction-following failures, and unsafe content with structured tags.
- Capture ambiguous prompts and route them back to the program team for rubric updates.
- Maintain reviewer-quality scores by calibrating against gold-standard examples each week.
- Document recurring failure modes so the modeling team can target them in the next training run.
Qualifications
- Prior evaluation, annotation, or human-rater experience on music generation quality evaluation or adjacent content for Music Generation Quality Evaluator work.
- Comfort applying multi-page rubrics consistently across long batches.
- Clear written reasoning that names the issue and the rubric clause being applied.
- Strong attention to detail and the ability to flag when a prompt itself is the problem.
- Reliable async availability for at least 10 hours per week.
Example tasks
- Compare two music generation quality evaluation model responses to the same prompt and pick the stronger one with rationale.
- Tag an unsafe response with the correct policy category and severity.
- Audit a 50-row batch for rubric consistency and report drift to the program lead.
- Propose a rubric clarification after spotting a recurring failure mode.
Nice to have
- Background in linguistics, content moderation, or trust & safety review.
- Experience with inter-rater agreement metrics and calibration cycles.
- Domain expertise that lets you spot subject-matter errors automated checks miss.
Skills
- Model output evaluation
- Rubric-based annotation
- Severity tagging
- Inter-rater calibration
- Music Generation Quality evaluation
- Creative review
- Production quality
- Design critique
- Music
- Generation
Work model
Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.
Compensation
Hourly rate confirmed after the interview process.
Application process
Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.
$20 - $36 per hour
...Role Overview Evaluate AI-generated music and lyrics in Slovak and English, applying music industry knowledge to judge quality, originality, and naturalness across a wide range of genres. This role supports development of generative music models by providing detailed...QualityHourly payImmediate startRemote workFlexible hours$17 per hour
...and Jack Dorsey . Position: Music Audio Expert - Thai Type: Contract... .../week Role Responsibilities Evaluate AI model output lyrics , voice generation, and other standards in various music categories. Score the quality of musical training data to ensure...QualityRemote jobContract workSummer workImmediate start$42 per hour
...Larry Summers , and Jack Dorsey . Position: Music & Lyrics Expert - Hebrew Type: Contract... ...Commitment: 20+ hours/week Role Responsibilities Evaluate AI-generated music across various genres for quality and creativity. Compare AI-generated lyrics...QualityRemote jobContract workSummer workImmediate start$14 - $42 per hour
...Summers , and Jack Dorsey . Position: Music & Lyrics Expert - Hindi Type: Contract... ...20+ hours/week Role Responsibilities Evaluate AI-generated music across various genres. Rate against detailed quality standards. Compare AI-generated lyrics with...QualityContract workSummer workImmediate startRemote work- Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models in partnership with a leading AI lab. You will assess AI... ...music across genres and rate it against detailed quality standards, working in Telugu and English. Start date...QualityImmediate startRemote workFlexible hours
$39 per hour
...Summers , and Jack Dorsey . Position: Music Audio Expert - English (US) Type:... ...hours/week Role Responsibilities Evaluate AI model output for lyrics, voice generation, and other musical standards. Score the quality of musical training data to enhance...QualityContract workSummer workImmediate startRemote work$20 - $30 per hour
...focused technology firm is looking for an individual proficient in evaluating large language model outputs. The role involves assessing AI... ..., and providing actionable feedback to enhance product quality. Ideal candidates will have strong analytical skills, attention...QualityRemote jobHourly pay$30 per hour
...$20-$30/hour Location: Remote Commitment: 10-40 hours/week Role Responsibilities Evaluate outputs from large language models and autonomous agent systems using defined rubrics and quality standards. Review multi-step agent workflows, including screenshots and reasoning...QualityRemote jobHourly payContract work$80 - $120 per hour
...legal and compliance expertise to assess AI-generated documents, spreadsheets, and slide decks... ...for accuracy, rigor, and overall domain quality. You will grade outputs using subject-... ...standards. Key Responsibilities Evaluate AI-generated work products against legal...QualityHourly payWork at officeRemote work$80 - $120 per hour
...Role Overview Assess AI-generated data analysis deliverables, including documents, spreadsheets... ..., methodological rigor, and domain quality. You will use deep subject-matter... ...Terms Remote, hourly engagement. You will evaluate AI-generated deliverables on an as-assigned...QualityHourly payWork at officeRemote work$80 - $120 per hour
...Role Overview Assess the accuracy, rigor, and domain quality of AI-generated legal work products related to intellectual property, trademark... ..., and slide decks, for legal accuracy and domain rigor. Evaluate outputs against domain-specific quality rubrics and grading...QualityHourly payWork at officeLocal areaRemote work$80 - $120 per hour
...Role Overview Assess AI-generated clinical, biomedical, and pharmaceutical work products, including documents, spreadsheets, and slide decks, for accuracy, scientific rigor, and domain quality. Use your subject-matter expertise to grade outputs and deliver clear, actionable...QualityHourly payWork at officeRemote work- Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models. You will assess AI-generated music across genres and rate it against detailed quality standards, working in Hindi and English. Responsibilities include head-to-head...QualityRemote work
$80 - $120 per hour
...Role Overview Assess AI-generated product launch and experiment readiness deliverables,... ...decks, for accuracy, rigor, and domain quality. Use deep subject-matter expertise to grade... ...completeness and domain correctness. Evaluate outputs against domain-specific quality...QualityHourly payContract workWork at officeRemote work- United States Digital Space LLC is seeking a part-time media search analyst to perform comprehensive assessments of music, video, and home pod evaluations across diverse media domains. You will analyze search results for App Store content and conduct online research to...QualityPart time
- ...Alignerr is seeking a Search Quality Evaluator to assess search engine results, AI-generated answers, and content recommendations. This fully remote, flexible contract role values strong critical thinking and quality instincts, with no technical background required beyond...QualityContract workRemote workFlexible hours
$80 - $120 per hour
...Special Education / IEP Evaluator $80-$120 per hour Hourly contract, remote About... ...Education / IEP to review and assess AI-generated work products (documents, spreadsheets,... ...decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise...QualityHourly payContract workFor contractorsWork at officeRemote work- ...Obsidian is hiring expert Evaluators in real estate, hospitality, and events to review AI-generated work products for accuracy and quality. This is a remote position requiring strong expertise in the relevant domains. Ideal candidates should have over 5 years of professional...QualityWork at officeRemote work
- ...role. About the role We are hiring expert Evaluators in Program management / implementation planning to review and assess AI-generated work products (documents, spreadsheets, and... ...decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise...QualityHourly payWork at officeRemote work
- ...Web Browsing Evaluator Contractor Remote Micro1 is engaging Web Browsing Evaluators to contribute high-quality web interaction data for a valued customer project. In this role... ...your expertise to help train next-generation AI systems. Your work will shape how models...QualityFor contractorsRemote work
- Obsidian is seeking expert Evaluators in Privacy / regulatory compliance to review AI-generated work products for accuracy and quality. The role is remote and requires 5+ years of relevant experience along with proficiency in Microsoft Office and Google Workspace. As an...QualityRemote jobHourly payWork at office
$100 per hour
...looking for a highly experienced software engineer (SR+) to help evaluate the quality of interactions with modern coding agents such as OpenAI... ...syntax correctness. What You’ll Be Doing: ~Evaluate AI-generated coding interactions end-to-end ~Judge whether outputs...QualityContract workImmediate start- Obsidian is seeking expert Evaluators in Biology/environmental science to review and assess AI-generated work products for accuracy and quality. In this remote, hourly role, you will leverage your expertise to provide feedback on documents and presentations, ensuring they...QualityRemote jobHourly pay
- Mercor is seeking a Contract Process Improvement / SOPs Evaluator to join our remote team. You will assess AI-generated artifacts against domain-specific quality rubrics and provide actionable feedback to improve accuracy and presentation. Collaborating with subject matter...QualityRemote jobContract work
$80 - $120 per hour
...is looking for a Legal contracts / diligence / redlines Evaluator to assess AI-generated artifacts. This role requires 5+ years of relevant experience... ...fluency in English. The ideal candidate will evaluate quality rubrics, provide structured feedback, and work independently...QualityRemote job- Obsidian is hiring expert Evaluators in Investment analysis / valuation / credit to review AI-generated work products for accuracy and quality. This remote, hourly position requires deep subject-matter expertise and professional fluency in English to provide structured...QualityHourly payWork at officeRemote work
- Obsidian is seeking expert Evaluators for a remote, hourly role focused on reviewing AI-generated work products for quality and accuracy. The successful candidates will leverage their subject matter expertise to assess various documents, spreadsheets, and slide decks....QualityRemote jobHourly payWork at office
- MERIT Beauty is seeking Turkish-speaking annotators for a contract role evaluating AI-generated content. You will review content for coherence, consistency, and alignment with real-world expectations in Turkish. The ideal candidates are native or fluent Turkish speakers...QualityRemote jobContract workTemporary workImmediate start
$45 - $85 per hour
Ixolabs is seeking an Animation Quality Evaluator to refine AI-generated animations. The ideal candidate will evaluate motion fluidity and naturalness, ensuring animations convey believable narratives. The role is part-time (15-25 hours/week) and offers competitive compensation...QualityRemote jobPart time$10 per hour
Dorado is seeking a Speech AI Evaluation Specialist to help improve AI-generated content in Korean. This freelance, remote role offers a flexible schedule... ..., and rate responses for relevance, accuracy, and clarity while meeting quality targets. #J-18808-Ljbffr DoradoQualityRemote jobHourly payPart timeFreelanceImmediate startWork from homeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Music Generation Quality Evaluator [Remote]. Be the first to apply!




