Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Evaluation Analyst [Remote]

$20 - $30 per hour

SaidGig

Role Overview

Contribute domain expertise to shape how next-generation language models learn and behave by authoring and evaluating multi-turn conversations, rubrics, and other evaluation assets. This contractor role focuses on producing high-quality, specification-driven evaluation and training data that influence model reasoning and performance. No prior AI work is required, domain knowledge and strong written analysis are the primary assets.

Key Responsibilities
  • Author detailed, task-based multi-turn conversations and corresponding evaluation rubrics that follow project specifications.
  • Test conversation drafts against frontier large language models, refine examples to meet required quality and difficulty, and iterate based on model behavior.
  • Deliver comprehensive evaluation assets, including transcripts, identified target behaviors, binary rubrics, and supporting evidence for judgments.
  • Maintain strict fidelity to evolving project specifications while sustaining high throughput and attention to detail.
  • Validate and calibrate outputs with team leads and quality control as guidelines change.
  • Work independently to meet expected output rates and complete deliverables on time.
Qualifications
  • Working knowledge of frontier LLM behavior and common model failure patterns.
  • Experience with data annotation, or demonstrated ability to follow detailed annotation specifications at scale.
  • Strong written English clarity, structure, and attention to detail, preferred at native level.
  • Self-direction to interpret and apply highly detailed specifications without supervision.
  • Critical thinking and analytical skills, especially in writing-heavy or analysis-heavy domains.
  • Preferred experience areas include RLHF, SFT, evaluation, prompt engineering, authoring evaluation items or rubrics, research, editorial work, technical writing, or quality assurance.
Work Terms
  • Role type, location: Contractor, Remote.
  • Start timeline: roles are typically filled within 48 hours. If selected, you should be ready to begin your first tasks within 24 to 48 hours after completing onboarding.
  • Work independently to meet expected output rates for deliverable completion.
  • Minimum submission requirements apply. Experts must submit a minimum of tasks per week.
Compensation
  • Pay range: $20 to $30 per hour.
  • Compensation structure: output-based, experts are paid per task that meets the project specifications. The time required to complete work will vary by expert and workflow.
Eligibility
  • No prior AI employment is required, applicants with domain expertise in finance, healthcare, STEM engineering, or other fields are encouraged.
  • Ability to follow detailed written specifications and to perform sustained, writing-heavy evaluation work is required.
Vacancy posted 1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Evaluation Analyst [Remote]. Be the first to apply!