Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

LLM Red-Teamer for AI Model Evaluation [Remote]

$40 - $65 per hour

SaidGig

Join a high-impact project focused on the evaluation and enhancement of frontier language models as an LLM Red-Teamer. In this role, you will leverage your expertise to train next-generation AI systems, shaping how models learn, reason, and perform through high-quality, real-world input. Key Responsibilities

  • Develop complex, adversarial multi-turn conversations and task-based scenarios aligned with detailed project specifications.
  • Author clear, precise evaluation rubrics to rigorously assess model responses against defined behavioral targets.
  • Iteratively test conversations and tasks against frontier LLMs, escalating difficulty and nuance until the desired quality threshold is achieved.
  • Deliver comprehensive task packages, including transcripts, target behaviors, binary rubrics, and supporting rationale or evidence.
  • Validate LLM outputs, documenting model strengths and failure modes relative to the project specification.
  • Maintain calibration with team leads and quality control contacts as project requirements evolve.
  • Contribute independently, producing high-quality deliverables at a steady and consistent pace.
Qualifications
  • Exceptional written English skills, with clarity, precision, and strong structural organization.
  • Prior experience in AI human data environments (RLHF, SFT, evaluations, annotation, or prompt engineering) is preferred.
  • Deep familiarity with large language models, including the ability to anticipate and identify common failure patterns.
  • Demonstrated ability to work autonomously, interpreting and executing complex specifications with minimal oversight.
  • Proven critical thinking and meticulous attention to detail.
  • Experience designing evaluation items or rubrics is advantageous.
  • Background in writing-intensive or analysis-centric fields such as research, editorial, technical writing, or quality assurance is a plus.
Work Terms

This is a contractor position with remote work flexibility. Experts are expected to submit a minimum number of tasks per week.

Compensation

Compensation is output-based, ranging from $40 to $65 per hour, depending on the expert''s experience and workflow.

Eligibility

We typically fill roles within 48 hours and are looking for experts ready to start immediately. Selected candidates are expected to begin their first tasks within 24, 48 hours of completing onboarding.

Vacancy posted 4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to LLM Red-Teamer for AI Model Evaluation [Remote]. Be the first to apply!