Expert Prompt Curators for Advanced AI Evaluation Dataset
CloudDevs
Expert Prompt Curators for Advanced AI Evaluation Dataset This description is a summary of our understanding of the job description. Click on ‘Apply’ button to find out more. Role Description Mercor is collaborating with a leading AI research lab to develop a next-generation evaluation dataset for frontier AI models. We are seeking experts with advanced domain knowledge across diverse fields to design extremely challenging prompts that cannot be solved by existing AI systems without internet search or browsing capabilities. The goal is to create a benchmark dataset that pushes the limits of current AI reasoning and retrieval. This is a short-term research engagement with significant impact on AI evaluation. Key Responsibilities Create original, expert-level prompts that require tool use (e.g., search, browse, or code execution). Ensure prompts are objective, self-contained, and yield clear, unambiguous answers. Test prompts against advanced AI models and document failures/successes. Provide reasoning steps and solutions for each prompt. Classify prompts into subject domains for dataset organization. Collaborate with reviewers for expert validation and prompt refinement. Qualifications Advanced academic or professional expertise in a specialized subject (STEM, law, finance, history, cultural studies, etc.). Strong ability to design precise, high-difficulty questions requiring deep knowledge and external references. Experience in academic research, benchmarking, or test question design preferred. Attention to detail and ability to provide concise reasoning explanations. Familiarity with AI models and their limitations is a plus. Requirements Remote and asynchronous — set your own hours. Expected commitment: ~10–20 hours/week. Project duration: ~2 months, with possible extensions based on dataset needs. Opportunity to contribute to high-impact AI safety and evaluation research. Compensation & Contract Terms Competitive hourly compensation based on expertise. Independent contractor engagement. Payments for services rendered processed weekly via Stripe Connect. Application Process Submit your resume or CV highlighting your subject matter expertise. Complete a brief questionnaire about your background and areas of specialization. Selected applicants may be asked to draft a short test prompt. You’ll receive follow-up within a few days regarding next steps. #J-18808-Ljbffr CloudDevs
- A leading AI research firm is seeking Expert Prompt Curators to design challenging prompts for evaluating advanced AI models. The role requires advanced knowledge in diverse fields and offers flexible hours, remote work, and a competitive hourly wage. Ideal candidates...SuggestedRemote jobHourly payTemporary workFlexible hours
$20 - $36 per hour
Italian Music & Lyrics Expert - AI Evaluation (Remote) is a remote evaluation track for reviewing italian generalist evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured...SuggestedRemote jobFor contractors10 hours per week- A leading educational organization is seeking a Linguistics Expert for a remote position to research and curate advanced linguistics materials into datasets for AI training. This role focuses on scientific accuracy and academic rigor, requiring a strong background in linguistics...SuggestedRemote jobHourly pay10 hours per week
- Turing Global India is seeking a Politics Domain Expert to evaluate and improve Large Language Models. You will design challenging prompts across political science, governance, elections, public policy and international relations, and assess factual accuracy and reasoning...SuggestedFor contractors
$73 per hour
A leading AI innovation firm is seeking individuals for a flexible, project-based role focusing on creating complex tasks for evaluating AI performance. Ideal candidates should possess a postgraduate... ...while contributing to advanced AI projects that shape the future...SuggestedRemote workFlexible hours- ...education and research sector is seeking a Social Sciences PhD Expert for a remote part-time role, requiring strong analytical... ...work. The position involves creating historically relevant prompts, evaluating AI outputs, and contributing to research initiatives. Ideal candidates...Remote jobPart timeFlexible hours
- Mercor is seeking experienced musicians to evaluate generative music AI models, collaborating with a leading AI lab. You will assess AI-generated... ...comparing lyrics to published songs, rating creativity and prompt adherence, and judging naturalness of language, slang, and...Part timeImmediate start
- Rise Data Labs is seeking advanced Mathematics and Statistics experts to support the training and evaluation of state-of-the-art AI systems. We need subject-matter experts who can apply deep quantitative knowledge to AI evaluation problems, assess AI-generated reasoning...Remote jobContract workImmediate startFlexible hours
- SME Careers is seeking a remote Data Scientist to contribute to AI training content and ensure model integrity. The role involves developing AI prompts, evaluating AI responses, and testing for model reliability. Ideal candidates will have a degree in a quantitative field...Remote jobHourly payFor contractors
- ...contribute deep scientific expertise to AI model evaluation. You will craft and assess challenging... ...develop and evaluate high‑difficulty physics prompts across classical mechanics to... ...prompting to surface errors, and provide expert critique of AI responses while working...Remote job
- ...Freelancer - Biology Expert for GenAI Prompts Review About the Role: We are seeking... ...involving Generative AI (GenAI) by creating biology-... ...dangerous CBRN-related outputs). Evaluate AI-generated responses to biology... ...evaluations. Requirements Advanced degree in Biology (Ph.D....FreelanceRemote work
$60 - $80 per hour
...creative and technical talent with leading AI research labs. Headquartered in San... ...solutions grounded in real retail practice. Evaluate AI model outputs against structured... ...Collaborate with other subject matter experts to ensure consistency and accuracy in training...Contract workSummer workRemote workWeekday work$30 per hour
Prolific is looking for Fluent Urdu Speakers to join our Expert Network to help train AI models using real legal expertise. The role involves completing... ...with pay rates of up to $30/hr. Candidates must possess advanced Urdu language skills and strong attention to detail....Remote jobWork from homeFlexible hours- ...context to support a range of AI training and evaluation projects. In this role, you will... ...types, including evaluating prompts and AI-generated outputs,... ...of linguists, subject matter experts, and language professionals who are advancing human knowledge together. Grow...Remote jobFor contractorsLocal area
- ...we believe the safest AI is the one that’s already... ...project - human data experts who probe AI models with... ...and agents: jailbreaks, prompt injections, misuse... ...reproducibly: produce reports, datasets, and attack cases... ...strengthen customer AI systems Evaluation coverage expands: more...Remote work
- Mercor is seeking a Music Audio Expert - German for a remote, project-based engagement. You will evaluate AI output lyrics and voice generation, score training data quality, and write music in the domain language listed in the title. Applicants should have 3+ years as...Remote job10 hours per week
- HireArt is seeking a Godot Software Expert for a short-term AI evaluation project. You will design complex Godot workflows and record end-to-end tasks for quality evaluation. This contract position runs ~1.5 weeks and is remote from the United States, with flexible hours...Remote jobContract workTemporary workPart time10 hours per weekFlexible hours
$30 per hour
Prolific is seeking fluent Hindi speakers to join their Expert Network, helping to train and evaluate AI models with real legal expertise. Responsibilities include analyzing and writing tasks in Hindi, judging AI’s performance, and aiding in the improvement of AI models...Remote jobHourly payFlexible hours- A tech company focusing on AI research is looking for experienced Krita users for a flexible, project-based contract opportunity. This role allows you to earn while evaluating AI-generated content related to digital painting and concept art. Candidates should have at least...Contract workRemote workFlexible hours
- A leading research services provider is seeking a Chemistry Expert (PhD) to evaluate complex chemistry problems and review AI-generated outputs for accuracy. This remote, hourly contract role requires deep subject-matter expertise and excellent communication skills. Candidates...Remote jobHourly payContract workFlexible hours
$100 - $120 per hour
...motivated Insurance Underwriting AI Expert to join our growing team.... ...domain expertise and advanced analytics, focusing on transforming... ..., improving accuracy in evaluating applicant profiles, exposures... ...historical data and external datasets to identify risk patterns, trends...Hourly payContract workPart timeRemote workWeekday work- ...re hiring PhD‑level biologists to help make advanced AI models safer. You'll apply your scientific expertise to evaluate and strengthen how these models handle... ...you on the workflow. Responsibilities Write expert‑level prompts across specialized life‑science topics. Evaluate...Part timeImmediate start
- 1. Overview Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models... .... We’re seeking talented Retail subject-matter experts (SMEs) with deep domain expertise and hands-on experience evaluating AI model outputs against rubrics to bring rigor...Contract workWeekday work
- Volga Partners is seeking experienced finance, audit, insurance, compliance, and legal professionals to join our on-call AI Evaluation Specialists network. This remote, project-based role involves evaluating AI outputs, applying critical thinking, and providing high-quality...Remote jobFlexible hours
- HireArt is seeking an experienced Audacity Software Expert for a short-term AI evaluation project. You will design complex Audacity tasks, develop evaluation rubrics, and record workflows from start to finish. This contract role focuses on high-quality, professional outputs...Remote jobContract workTemporary workFlexible hours
$73 per hour
...ethically shape the future of AI. What We Do The... ...systems are tested and evaluated? This is a flexible,... ...hand‑holding. Real expert complexity only. You're... ...From creating training prompts to refining model... ...commitments. Work on advanced AI projects and gain valuable...Permanent employmentPart timeFreelanceRemote workFlexible hours$55 per hour
...Freelance Biology Expert with Python - AI Trainer 5 days ago – Be among the... ...Responsibilities Generate prompts that challenge AI.... ...comprehensive scoring criteria to evaluate the accuracy of the AI's answers... ..., Quantitative Biology. Advanced English proficiency (C1 or...Part timeFreelanceRemote work$80 - $120 per hour
...creative and technical talent with leading AI research labs. Headquartered in San... ...Position: Software / AI / IT / data Evaluator Type: Contract Compensation: $... ...especially Slides . Preferred ~ Advanced degree ( Master's or higher ) from a reputable...Contract workSummer workWork at officeRemote work$80 - $120 per hour
...creative and technical talent with leading AI research labs. Headquartered in San... ...Position: Software / AI / IT / data Evaluator Type: Contract Compensation: $... ...especially Slides . Preferred ~ Advanced degree ( Master's or higher) from a reputable...Contract workSummer workWork at officeRemote work$8 - $65 per hour
...Remote Overview Are you a Guaraní language expert eager to shape the future of AI? Large-scale language models are... ...traces, and suggest improvements to our prompt engineering and evaluation metrics. You’ll challenge advanced language models on topics such as contextual...Hourly payContract workFor contractorsFreelanceRemote workWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Expert Prompt Curators for Advanced AI Evaluation Dataset. Be the first to apply!


