Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Expert Prompt Curators for Advanced AI Evaluation Dataset

CloudDevs

Expert Prompt Curators for Advanced AI Evaluation Dataset This description is a summary of our understanding of the job description. Click on ‘Apply’ button to find out more. Role Description Mercor is collaborating with a leading AI research lab to develop a next-generation evaluation dataset for frontier AI models. We are seeking experts with advanced domain knowledge across diverse fields to design extremely challenging prompts that cannot be solved by existing AI systems without internet search or browsing capabilities. The goal is to create a benchmark dataset that pushes the limits of current AI reasoning and retrieval. This is a short-term research engagement with significant impact on AI evaluation. Key Responsibilities Create original, expert-level prompts that require tool use (e.g., search, browse, or code execution). Ensure prompts are objective, self-contained, and yield clear, unambiguous answers. Test prompts against advanced AI models and document failures/successes. Provide reasoning steps and solutions for each prompt. Classify prompts into subject domains for dataset organization. Collaborate with reviewers for expert validation and prompt refinement. Qualifications Advanced academic or professional expertise in a specialized subject (STEM, law, finance, history, cultural studies, etc.). Strong ability to design precise, high-difficulty questions requiring deep knowledge and external references. Experience in academic research, benchmarking, or test question design preferred. Attention to detail and ability to provide concise reasoning explanations. Familiarity with AI models and their limitations is a plus. Requirements Remote and asynchronous — set your own hours. Expected commitment: ~10–20 hours/week. Project duration: ~2 months, with possible extensions based on dataset needs. Opportunity to contribute to high-impact AI safety and evaluation research. Compensation & Contract Terms Competitive hourly compensation based on expertise. Independent contractor engagement. Payments for services rendered processed weekly via Stripe Connect. Application Process Submit your resume or CV highlighting your subject matter expertise. Complete a brief questionnaire about your background and areas of specialization. Selected applicants may be asked to draft a short test prompt. You’ll receive follow-up within a few days regarding next steps. #J-18808-Ljbffr CloudDevs

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Expert Prompt Curators for Advanced AI Evaluation Dataset in New York, NY vacancy
  • A leading AI research firm is seeking Expert Prompt Curators to design challenging prompts for evaluating advanced AI models. The role requires advanced knowledge in diverse fields and offers flexible hours, remote work, and a competitive hourly wage. Ideal candidates... 
    Suggested
    Remote job
    Hourly pay
    Temporary work
    Flexible hours

    CloudDevs

    New York, NY
    2 days ago
  • $20 - $36 per hour

    Italian Music & Lyrics Expert - AI Evaluation (Remote) is a remote evaluation track for reviewing italian generalist evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured... 
    Suggested
    Remote job
    For contractors
    10 hours per week

    AuraOne

    New York, NY
    3 days ago
  • A leading educational organization is seeking a Linguistics Expert for a remote position to research and curate advanced linguistics materials into datasets for AI training. This role focuses on scientific accuracy and academic rigor, requiring a strong background in linguistics... 
    Suggested
    Remote job
    Hourly pay
    10 hours per week

    Crossing Hurdles

    New York, NY
    2 days ago
  • Turing Global India is seeking a Politics Domain Expert to evaluate and improve Large Language Models. You will design challenging prompts across political science, governance, elections, public policy and international relations, and assess factual accuracy and reasoning... 
    Suggested
    For contractors

    Turing Global India

    New York, NY
    3 days ago
  • $73 per hour

    A leading AI innovation firm is seeking individuals for a flexible, project-based role focusing on creating complex tasks for evaluating AI performance. Ideal candidates should possess a postgraduate...  ...while contributing to advanced AI projects that shape the future... 
    Suggested
    Remote work
    Flexible hours

    Mindrift

    New York, NY
    3 days ago
  •  ...education and research sector is seeking a Social Sciences PhD Expert for a remote part-time role, requiring strong analytical...  ...work. The position involves creating historically relevant prompts, evaluating AI outputs, and contributing to research initiatives. Ideal candidates... 
    Remote job
    Part time
    Flexible hours

    Crossing Hurdles

    New York, NY
    2 days ago
  • Mercor is seeking experienced musicians to evaluate generative music AI models, collaborating with a leading AI lab. You will assess AI-generated...  ...comparing lyrics to published songs, rating creativity and prompt adherence, and judging naturalness of language, slang, and... 
    Part time
    Immediate start

    Obsidian

    New York, NY
    3 days ago
  • Rise Data Labs is seeking advanced Mathematics and Statistics experts to support the training and evaluation of state-of-the-art AI systems. We need subject-matter experts who can apply deep quantitative knowledge to AI evaluation problems, assess AI-generated reasoning... 
    Remote job
    Contract work
    Immediate start
    Flexible hours

    BAM Ventures

    New York, NY
    3 days ago
  • SME Careers is seeking a remote Data Scientist to contribute to AI training content and ensure model integrity. The role involves developing AI prompts, evaluating AI responses, and testing for model reliability. Ideal candidates will have a degree in a quantitative field... 
    Remote job
    Hourly pay
    For contractors

    SME Careers

    New York, NY
    4 days ago
  •  ...contribute deep scientific expertise to AI model evaluation. You will craft and assess challenging...  ...develop and evaluate high‑difficulty physics prompts across classical mechanics to...  ...prompting to surface errors, and provide expert critique of AI responses while working... 
    Remote job

    Dorado

    New York, NY
    2 days ago
  •  ...Freelancer - Biology Expert for GenAI Prompts Review About the Role: We are seeking...  ...involving Generative AI (GenAI) by creating biology-...  ...dangerous CBRN-related outputs). Evaluate AI-generated responses to biology...  ...evaluations. Requirements Advanced degree in Biology (Ph.D.... 
    Freelance
    Remote work

    ActiveFence

    New York, NY
    4 days ago
  • $60 - $80 per hour

     ...creative and technical talent with leading AI research labs. Headquartered in San...  ...solutions grounded in real retail practice. Evaluate AI model outputs against structured...  ...Collaborate with other subject matter experts to ensure consistency and accuracy in training... 
    Contract work
    Summer work
    Remote work
    Weekday work

    Mercor

    New York, NY
    3 days ago
  • $30 per hour

    Prolific is looking for Fluent Urdu Speakers to join our Expert Network to help train AI models using real legal expertise. The role involves completing...  ...with pay rates of up to $30/hr. Candidates must possess advanced Urdu language skills and strong attention to detail.... 
    Remote job
    Work from home
    Flexible hours

    Prolific

    New York, NY
    3 days ago
  •  ...context to support a range of AI training and evaluation projects. In this role, you will...  ...types, including evaluating prompts and AI-generated outputs,...  ...of linguists, subject matter experts, and language professionals who are advancing human knowledge together. Grow... 
    Remote job
    For contractors
    Local area

    Lilt

    New York, NY
    4 days ago
  •  ...we believe the safest AI is the one that’s already...  ...project - human data experts who probe AI models with...  ...and agents: jailbreaks, prompt injections, misuse...  ...reproducibly: produce reports, datasets, and attack cases...  ...strengthen customer AI systems Evaluation coverage expands: more... 
    Remote work

    Mercor Inc

    New York, NY
    4 days ago
  • Mercor is seeking a Music Audio Expert - German for a remote, project-based engagement. You will evaluate AI output lyrics and voice generation, score training data quality, and write music in the domain language listed in the title. Applicants should have 3+ years as... 
    Remote job
    10 hours per week

    United States Digital Space LLC

    New York, NY
    4 days ago
  • HireArt is seeking a Godot Software Expert for a short-term AI evaluation project. You will design complex Godot workflows and record end-to-end tasks for quality evaluation. This contract position runs ~1.5 weeks and is remote from the United States, with flexible hours... 
    Remote job
    Contract work
    Temporary work
    Part time
    10 hours per week
    Flexible hours

    HireArt

    New York, NY
    3 days ago
  • $30 per hour

    Prolific is seeking fluent Hindi speakers to join their Expert Network, helping to train and evaluate AI models with real legal expertise. Responsibilities include analyzing and writing tasks in Hindi, judging AI’s performance, and aiding in the improvement of AI models... 
    Remote job
    Hourly pay
    Flexible hours

    Prolific

    New York, NY
    3 days ago
  • A tech company focusing on AI research is looking for experienced Krita users for a flexible, project-based contract opportunity. This role allows you to earn while evaluating AI-generated content related to digital painting and concept art. Candidates should have at least... 
    Contract work
    Remote work
    Flexible hours

    Handshake

    New York, NY
    1 day ago
  • A leading research services provider is seeking a Chemistry Expert (PhD) to evaluate complex chemistry problems and review AI-generated outputs for accuracy. This remote, hourly contract role requires deep subject-matter expertise and excellent communication skills. Candidates... 
    Remote job
    Hourly pay
    Contract work
    Flexible hours

    Crossing Hurdles

    New York, NY
    2 days ago
  • $100 - $120 per hour

     ...motivated Insurance Underwriting AI Expert to join our growing team....  ...domain expertise and advanced analytics, focusing on transforming...  ..., improving accuracy in evaluating applicant profiles, exposures...  ...historical data and external datasets to identify risk patterns, trends... 
    Hourly pay
    Contract work
    Part time
    Remote work
    Weekday work

    Weekday AI (YC W21)

    New York, NY
    2 days ago
  •  ...re hiring PhD‑level biologists to help make advanced AI models safer. You'll apply your scientific expertise to evaluate and strengthen how these models handle...  ...you on the workflow. Responsibilities Write expert‑level prompts across specialized life‑science topics. Evaluate... 
    Part time
    Immediate start

    Obsidian

    New York, NY
    5 days ago
  • 1. Overview Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models...  .... We’re seeking talented Retail subject-matter experts (SMEs) with deep domain expertise and hands-on experience evaluating AI model outputs against rubrics to bring rigor... 
    Contract work
    Weekday work

    Obsidian

    New York, NY
    3 days ago
  • Volga Partners is seeking experienced finance, audit, insurance, compliance, and legal professionals to join our on-call AI Evaluation Specialists network. This remote, project-based role involves evaluating AI outputs, applying critical thinking, and providing high-quality... 
    Remote job
    Flexible hours

    Dorado

    New York, NY
    1 day ago
  • HireArt is seeking an experienced Audacity Software Expert for a short-term AI evaluation project. You will design complex Audacity tasks, develop evaluation rubrics, and record workflows from start to finish. This contract role focuses on high-quality, professional outputs... 
    Remote job
    Contract work
    Temporary work
    Flexible hours

    HireArt

    New York, NY
    3 days ago
  • $73 per hour

     ...ethically shape the future of AI. What We Do The...  ...systems are tested and evaluated? This is a flexible,...  ...hand‑holding. Real expert complexity only. You're...  ...From creating training prompts to refining model...  ...commitments. Work on advanced AI projects and gain valuable... 
    Permanent employment
    Part time
    Freelance
    Remote work
    Flexible hours

    Mind Rift

    New York, NY
    4 days ago
  • $55 per hour

     ...Freelance Biology Expert with Python - AI Trainer 5 days ago – Be among the...  ...Responsibilities Generate prompts that challenge AI....  ...comprehensive scoring criteria to evaluate the accuracy of the AI's answers...  ..., Quantitative Biology. Advanced English proficiency (C1 or... 
    Part time
    Freelance
    Remote work

    Mind Rift

    New York, NY
    4 days ago
  • $80 - $120 per hour

     ...creative and technical talent with leading AI research labs. Headquartered in San...  ...Position: Software / AI / IT / data Evaluator Type: Contract Compensation: $...  ...especially Slides . Preferred ~ Advanced degree ( Master's or higher ) from a reputable... 
    Contract work
    Summer work
    Work at office
    Remote work

    Mercor

    New York, NY
    20 days ago
  • $80 - $120 per hour

     ...creative and technical talent with leading AI research labs. Headquartered in San...  ...Position: Software / AI / IT / data Evaluator Type: Contract Compensation: $...  ...especially Slides . Preferred ~ Advanced degree ( Master's or higher) from a reputable... 
    Contract work
    Summer work
    Work at office
    Remote work

    Mercor

    New York, NY
    11 days ago
  • $8 - $65 per hour

     ...Remote Overview Are you a Guaraní language expert eager to shape the future of AI? Large-scale language models are...  ...traces, and suggest improvements to our prompt engineering and evaluation metrics. You’ll challenge advanced language models on topics such as contextual... 
    Hourly pay
    Contract work
    For contractors
    Freelance
    Remote work
    Worldwide

    Meridial

    New York, NY
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Expert Prompt Curators for Advanced AI Evaluation Dataset. Be the first to apply!