Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

LLM Red-Teamer for AI Model Evaluation [Remote]

$40 - $65 per hour

SaidGig

Remote
  • Remote job

Join a high-impact project focused on the evaluation and enhancement of frontier language models as an LLM Red-Teamer. In this role, you will leverage your expertise to train next-generation AI systems, shaping how models learn, reason, and perform through high-quality, real-world input. Key Responsibilities

  • Develop complex, adversarial multi-turn conversations and task-based scenarios aligned with detailed project specifications.
  • Author clear, precise evaluation rubrics to rigorously assess model responses against defined behavioral targets.
  • Iteratively test conversations and tasks against frontier LLMs, escalating difficulty and nuance until the desired quality threshold is achieved.
  • Deliver comprehensive task packages, including transcripts, target behaviors, binary rubrics, and supporting rationale or evidence.
  • Validate LLM outputs, documenting model strengths and failure modes relative to the project specification.
  • Maintain calibration with team leads and quality control contacts as project requirements evolve.
  • Contribute independently, producing high-quality deliverables at a steady and consistent pace.
Qualifications
  • Exceptional written English skills, with clarity, precision, and strong structural organization.
  • Prior experience in AI human data environments (RLHF, SFT, evaluations, annotation, or prompt engineering) is preferred.
  • Deep familiarity with large language models, including the ability to anticipate and identify common failure patterns.
  • Demonstrated ability to work autonomously, interpreting and executing complex specifications with minimal oversight.
  • Proven critical thinking and meticulous attention to detail.
  • Experience designing evaluation items or rubrics is advantageous.
  • Background in writing-intensive or analysis-centric fields such as research, editorial, technical writing, or quality assurance is a plus.
Work Terms

This is a contractor position with remote work flexibility. Experts are expected to submit a minimum number of tasks per week.

Compensation

Compensation is output-based, ranging from $40 to $65 per hour, depending on the expert''s experience and workflow.

Eligibility

We typically fill roles within 48 hours and are looking for experts ready to start immediately. Selected candidates are expected to begin their first tasks within 24, 48 hours of completing onboarding.

Vacancy posted 15 days ago
Similar jobs that could be interesting for youBased on the LLM Red-Teamer for AI Model Evaluation [Remote] in Remote vacancy
  • $40 - $65 per hour

     ...Join a high-impact project focused on the evaluation and enhancement of frontier language models as an LLM Red-Teamer. In this role, you will leverage your expertise to train next-generation AI systems, shaping how models learn, reason, and perform through high-quality... 
    Suggested
    Remote job
    Hourly pay
    For contractors
    Immediate start

    SaidGig

    Frontier County, NE
    15 days ago
  • $60 - $90 per hour

     ...where frontier language models appear competent but quietly...  ...findings into robust evaluation benchmarks. Key...  ...engineering, security, or AI evaluation. Proven ability...  ...or ML systems, through red teaming, adversarial...  ...Strong familiarity with LLM capabilities, limitations... 
    Suggested
    Hourly pay
    Full time
    Freelance
    Remote work

    SaidGig

    United States
    2 days ago
  • $350k

     ...role in shaping the future of AI-powered legal reasoning. This...  ...intersection of large language models, agentic systems, and legal workflows...  ...the development of rigorous evaluation frameworks to measure and...  ...Advanced degree in Law (JD, LLM, SJD, PhD in Law, or equivalent... 
    Suggested
    Remote job
    Full time

    SaidGig

    United States
    4 days ago
  • $105 per hour

     ...leverage their expertise to contribute to AI research projects focused on high-...  ...community of experts to refine and evaluate the capabilities of Large Language Models (LLMs) in creating impactful...  ...domain-specific prompts and evaluate LLM responses for business contexts.... 
    Suggested
    Work experience placement
    Remote work
    Flexible hours

    SaidGig

    United States
    4 days ago
  • $60 - $150 per hour

     ...Role Overview Provide legal subject-matter expertise to improve and evaluate AI systems, by designing realistic legal tasks, reviewing model outputs, and giving domain-specific feedback that advances frontier AI research. This is an open application to join a Law Expert... 
    Suggested
    Hourly pay
    Contract work
    Immediate start
    Remote work

    SaidGig

    United States
    2 days ago
  • $80 - $110 per hour

     ...Overview Work on the forefront of generative AI by designing and executing real-world...  ...and reasoning gaps in advanced models. You will author tasks, produce reference...  ...and executable tests where applicable, run evaluations against a target model, and analyze failures... 
    Hourly pay
    Part time
    Freelance
    Remote work

    SaidGig

    United States
    11 days ago
  • $85 per hour

     ...Role Overview GIS Analysts apply hands-on expertise in environmental assessment, GIS analysis, and renewable energy siting to evaluate AI-generated geospatial outputs and develop expert training data that improves AI understanding of environmental workflows and mapping... 
    Hourly pay
    Contract work
    Part time
    Work at office
    Remote work
    Flexible hours

    SaidGig

    United States
    4 days ago
  • $85 per hour

     ...Role Overview REC Traders evaluate AI-generated content using their renewable energy and REC trading expertise, creating expert training...  ...clear, detailed feedback on AI-generated responses to help refine model behavior and domain correctness. Work independently and... 
    Hourly pay
    Full time
    Contract work
    Part time
    For contractors
    Remote work
    Flexible hours

    SaidGig

    United States
    4 days ago
  • $75 per hour

     ...Managers apply archival, library, and collections expertise to evaluate and improve AI-generated content related to records, archives, and...  ...will create prompts that reflect real workplace tasks, review model outputs for accuracy and relevance, and provide clear, structured... 
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    4 days ago
  •  ...Physicians apply clinical judgment and frontline medical experience to evaluate AI-generated medical content, ensuring clinical accuracy, sound...  ...planning. Assess clarity, relevance, and safety of model outputs in realistic care scenarios. Provide detailed, constructive... 
    Full time
    For contractors
    Private practice
    Remote work
    Flexible hours

    SaidGig

    United States
    4 days ago
  • $80 - $150 per hour

     ...in a high-impact project that shapes the future of AI systems by applying your physics expertise. As a...  ...Junior Professor), you will play a critical role in evaluating and enhancing the training of next-generation AI models, ensuring they learn and reason effectively... 
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Remote
    15 days ago
  • $75 per hour

     ...Role Overview Physics experts apply advanced physics training to evaluate AI-generated scientific content and provide detailed feedback that improves AI physical reasoning, mathematical modeling, theoretical analysis, and quantitative problem solving. This is a project... 
    Hourly pay
    Full time
    Contract work
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    1 day ago
  • $220k

     ...This role focuses on advancing the evaluation and development of cutting-edge coding agents. You will operate at the intersection of AI research, software engineering, and model evaluation, designing the benchmarks, methodologies, and data systems that shape how next-... 
    Full time
    Remote work

    SaidGig

    United States
    4 days ago
  • $90 - $130 per hour

     ...seasoned funds-focused legal expertise to improve how advanced AI systems read, evaluate, and negotiate investment-related contracts. In this part-...  ...and give precise legal feedback that trains and refines AI models for fund formation and fund management workflows. Key... 
    Hourly pay
    Contract work
    Part time
    For contractors
    Work at office
    Remote work

    SaidGig

    United States
    14 days ago
  • $80 - $105 per hour

     ...Role Overview Help define how advanced AI understands and evaluates fund-related contracts by applying hands-on funds law experience to contract redlining, simulated negotiations, and model evaluation. This part-time, contractor role supports the development of AI systems... 
    Hourly pay
    Contract work
    Part time
    For contractors
    Remote work

    SaidGig

    United States
    8 days ago
  • $80 - $135 per hour

     ...end reference solutions for the CritPt benchmark (arXiv:2509.26574v3). This role produces definitive solutions used to evaluate large language models on frontier physics reasoning, by solving research-level problems, auditing peer submissions, or adjudicating between competing... 
    Hourly pay
    Remote work
    10 hours per week

    SaidGig

    United States
    3 days ago
  • $60 per hour

     ...and contribute to developing cutting-edge AI systems, while enjoying the flexibility...  ...professionals to help advance AI development. AI models are increasingly capable of performing...  ...state-of-the-art AI models on tasks like evaluating AI-generated quantitative analysis,... 
    Hourly pay
    Full time
    Remote work
    Flexible hours

    DataAnnotation

    Little Rock, AR
    2 days ago
  • $40 per hour

     ...A leading AI development firm is looking for experienced quantitative professionals to evaluate AI-generated work and design problems for AI training. This fully remote position allows for a flexible schedule, offering competitive pay starting at $40+ per hour. Ideal candidates... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Bismarck, ND
    1 day ago
  • $40 per hour

     ...A leading AI development firm is seeking experienced quantitative professionals to evaluate AI-generated quantitative analysis and provide impactful feedback. This fully remote role allows for flexible scheduling and competitive pay starting at $40 per hour. Candidates... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Honolulu, HI
    3 days ago
  • $40 per hour

    A leading AI development firm in Michigan is seeking experienced quantitative professionals to evaluate AI-generated analyses and contribute to the evolution of AI models. Candidates should have a background in data science, statistics, or similar fields, with at least... 
    Hourly pay
    Remote work

    DataAnnotation

    Lansing, MI
    1 day ago
  • $40 per hour

     ...A leading AI development firm is seeking experienced quantitative professionals to join their team remotely. The role involves evaluating AI-generated quantitative work, providing insights, and shaping the future of AI systems. Candidates should have over two years of... 
    Hourly pay
    Remote work

    DataAnnotation

    Indiana, PA
    1 day ago
  • $40 per hour

     ...data annotation company is seeking professionals in quantitative fields to enhance AI development. This fully remote role allows individuals to set flexible schedules while evaluating AI-generated analyses and solving complex quantitative problems. Candidates should have... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Oregon, WI
    1 day ago
  • $40 per hour

     ...A forward-thinking AI team is seeking quantitative professionals to evaluate and improve cutting-edge AI systems. The role involves working on AI-generated analyses and providing critical feedback for model enhancement. Candidates should have a strong quantitative background... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Helena, MT
    3 days ago
  • $40 per hour

     ...A forward-thinking AI solutions company is seeking experienced quantitative professionals to evaluate AI-generated analyses and contribute to the development of cutting-edge...  ...skills. Join us to directly impact the future of AI analytics and model reasoning. #J-18808-Ljbffr... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Lincoln, NE
    1 day ago
  • $40 per hour

    A leading AI development firm is seeking experienced quantitative professionals to evaluate AI-generated quantitative work and provide critical feedback. This role offers the flexibility of remote work, allowing you to set your own schedule while focusing on impactful projects... 
    Hourly pay
    Remote work

    DataAnnotation

    Wisconsin
    1 day ago
  • $40 per hour

     ...A leading AI company in the United States is seeking experienced quantitative professionals to evaluate and validate AI-generated analytical work. This fully remote position allows you to set your own schedule, with competitive hourly pay starting at $40 USD. Responsibilities... 
    Hourly pay
    Remote work

    DataAnnotation

    Jackson, MS
    1 day ago
  • $40 per hour

    A data science team is seeking experienced quantitative professionals to evaluate AI-generated work and contribute to the development of cutting-edge AI systems. This fully remote position offers flexible scheduling and competitive hourly pay starting at $40+. Ideal candidates... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Madison, WI
    1 day ago
  • $40 per hour

     ...An innovative AI development company is seeking experienced quantitative professionals to contribute to AI advancements. This fully remote role involves evaluating AI-generated analyses and ensuring they are technically accurate and valid in real-world scenarios. Candidates... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Santa Fe, NM
    3 days ago
  • $40 per hour

    A leading AI development company is seeking experienced quantitative professionals to evaluate AI-generated analyses and design quantitative problems for AI training. This fully remote role offers flexibility in project selection and scheduling, with competitive pay starting... 
    Hourly pay
    Remote work

    DataAnnotation

    Juneau, AK
    2 days ago
  • $40 per hour

    A leading AI development company is seeking experienced quantitative professionals for a remote role. Candidates will evaluate AI-generated quantitative work, solve complex problems, and provide valuable feedback. The ideal candidate has 2+ years of experience in a quantitative... 
    Hourly pay
    Full time
    Remote work
    Flexible hours

    DataAnnotation

    Topeka, KS
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to LLM Red-Teamer for AI Model Evaluation [Remote]. Be the first to apply!