AI Evaluation Analyst [Remote]
$20 - $30 per hourSaidGig
Frontier County, NE
- Remote job
Role Overview
Contribute domain expertise to shape how next-generation language models learn and behave by authoring and evaluating multi-turn conversations, rubrics, and other evaluation assets. This contractor role focuses on producing high-quality, specification-driven evaluation and training data that influence model reasoning and performance. No prior AI work is required, domain knowledge and strong written analysis are the primary assets.
Key Responsibilities- Author detailed, task-based multi-turn conversations and corresponding evaluation rubrics that follow project specifications.
- Test conversation drafts against frontier large language models, refine examples to meet required quality and difficulty, and iterate based on model behavior.
- Deliver comprehensive evaluation assets, including transcripts, identified target behaviors, binary rubrics, and supporting evidence for judgments.
- Maintain strict fidelity to evolving project specifications while sustaining high throughput and attention to detail.
- Validate and calibrate outputs with team leads and quality control as guidelines change.
- Work independently to meet expected output rates and complete deliverables on time.
- Working knowledge of frontier LLM behavior and common model failure patterns.
- Experience with data annotation, or demonstrated ability to follow detailed annotation specifications at scale.
- Strong written English clarity, structure, and attention to detail, preferred at native level.
- Self-direction to interpret and apply highly detailed specifications without supervision.
- Critical thinking and analytical skills, especially in writing-heavy or analysis-heavy domains.
- Preferred experience areas include RLHF, SFT, evaluation, prompt engineering, authoring evaluation items or rubrics, research, editorial work, technical writing, or quality assurance.
- Role type, location: Contractor, Remote.
- Start timeline: roles are typically filled within 48 hours. If selected, you should be ready to begin your first tasks within 24 to 48 hours after completing onboarding.
- Work independently to meet expected output rates for deliverable completion.
- Minimum submission requirements apply. Experts must submit a minimum of tasks per week.
- Pay range: $20 to $30 per hour.
- Compensation structure: output-based, experts are paid per task that meets the project specifications. The time required to complete work will vary by expert and workflow.
- No prior AI employment is required, applicants with domain expertise in finance, healthcare, STEM engineering, or other fields are encouraged.
- Ability to follow detailed written specifications and to perform sustained, writing-heavy evaluation work is required.
Vacancy posted 1 day ago
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Evaluation Analyst [Remote]. Be the first to apply!
