Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

LLM Red Team Specialist for AI Model Evaluation

$60 - $90 per hour

SaidGig

Join a leading AI lab''s cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. Overview

A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models and needs specialists who can find where those models break. Working in a red-teaming setup, you will design and probe complex, multi-step tasks to expose vulnerabilities, edge cases, and failure modes in frontier AI systems, the places where a model looks competent but is quietly wrong.

Each task represents one to two days of continuous, focused effort and spans multiple technical skills: coding, experimentation, and careful analysis. You will work in a tight feedback loop with the lab''s researchers, turning the failure modes you find into stronger benchmark tasks. This role is fully remote within the United States, at approximately 35 hours per week.

Key Responsibilities
  • Probe models: Explore how frontier AI models behave on coding, ML, and analysis tasks, and find the spots where they quietly get things wrong.
  • Design challenges: Turn the weaknesses you find into well-crafted tasks that are hard for models but fair to grade.
  • Document findings: Write up what you discover clearly, with evidence and steps others can reproduce.
  • Strengthen tasks: Team up with task authors to close loopholes, shortcuts, and grading gaps.
  • Work as a team: Share insights with researchers and fellow experts so the benchmark keeps getting better.
Core Qualifications
  • MSc or PhD in a STEM field, or equivalent practical experience in a research-heavy domain requiring data analysis and coding.
  • 1+ years of experience in a research, research-engineering, security, or AI-evaluation role.
  • Demonstrated ability to identify vulnerabilities, edge cases, or failure modes in LLMs or ML systems, through red teaming, adversarial testing, security research, or rigorous model evaluation.
  • Working proficiency in Python and Git, with the ability to script your own probes and analyses.
  • Strong familiarity with LLM capabilities, limitations, and evaluation techniques.
  • Past experience in AI training, model evaluation, or benchmark/task authoring is preferred.
  • A perfectionist mindset: high attention to detail, creativity in finding what others missed, strong written communication, and the ability to work independently through ambiguous, open-ended problems.
  • Ability to engage reliably for approximately 35 hours per week.
Work Terms

This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce.

Compensation

Hourly compensation ranges from $60 to $90.

Eligibility

This role is fully remote within the United States.

Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives.

Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client''s internal teams, and integration into standard enterprise workflows.

Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, disability, or any other characteristic protected by applicable law.]]><

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the LLM Red Team Specialist for AI Model Evaluation in United States vacancy
  • $40 - $65 per hour

     ...high-impact project focused on the evaluation and enhancement of frontier language models as an LLM Red-Teamer. In this role, you will...  ...to train next-generation AI systems, shaping how models learn...  ...specification. Maintain calibration with team leads and quality control... 
    Suggested
    Remote job
    Hourly pay
    For contractors
    Immediate start

    SaidGig

    Frontier County, NE
    11 days ago
  • $208k - $300k

     ...Machine Learning Engineer - Model Evaluations, Public Sector The Public Sector ML team at Scale deploys advanced AI systems—including LLMs, agentic...  ...safety metrics, including LLM-judge–based evaluations....  ...and execute stress tests and red-teaming workflows to uncover... 
    Suggested
    Full time

    Scale Ai

    Washington DC
    1 day ago
  • $350k

     ...Join a dynamic research team as a Member of Technical...  ...in shaping the future of AI-powered legal reasoning....  ...of large language models, agentic systems, and legal...  ...development of rigorous evaluation frameworks to measure and...  ...Advanced degree in Law (JD, LLM, SJD, PhD in Law, or... 
    Suggested
    Remote job
    Full time

    SaidGig

    United States
    1 day ago
  • $15 - $25 per hour

     ...As a Legal Specialist, you will leverage your legal expertise...  ...of next-generation AI systems. Your insights...  ...role in shaping how these models learn, reason, and...  ...support structured legal evaluations. Prepare concise written...  ...with cross-functional teams to enhance legal... 
    Suggested
    Hourly pay
    Contract work
    Part time
    For contractors
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $161.8k - $184.6k

     ...Associate, Data Scientist - LLM Customization Team At Capital One, we think...  ...for the Open Banking future. AI is transforming every industry...  ...the power of Large Language Models (LLMs), adapt and finetune...  ...from design through training, evaluation, and validation; partnering... 
    Suggested
    Full time
    Part time
    Local area
    Flexible hours

    Capital One

    McLean, VA
    4 days ago
  • $125 per hour

     ...QGIS specialists leverage their expertise in geographic information...  ...analysis to support AI research through...  ...work. This role involves evaluating AI-generated content...  ...prompts and evaluating LLM responses. Contribute...  ...asynchronously with AI research teams. Work Terms This... 
    Remote work
    Flexible hours

    SaidGig

    United States
    2 days ago
  • $80 - $110 per hour

     ...Join a cutting-edge GenAI team at a leading AI lab, where your expertise will be pivotal in developing advanced AI models. This role focuses on designing and evaluating machine learning and natural language processing tasks that will help identify and address capability... 
    Hourly pay
    Part time
    Remote work

    SaidGig

    United States
    2 days ago
  • $80 - $110 per hour

     ...Join a cutting-edge GenAI team at a leading AI lab, where your expertise will be pivotal in developing advanced AI models. This role focuses on building and evaluating frontier models, requiring experienced computer vision practitioners to serve as ground-truth experts... 
    Hourly pay
    Part time
    Remote work

    SaidGig

    United States
    8 days ago
  • $30 - $90 per hour

     ...Developer, you will play a crucial role in evaluating and training next-generation AI coding tools during their highly...  .... Test and evaluate alpha AI models in Cursor over multiple 4-day, 5+...  .... Collaborate with the research team via Slack, providing real-time feedback... 
    Remote job
    Hourly pay
    Contract work
    Part time

    SaidGig

    United States
    5 days ago
  • $220k

     ...This role focuses on advancing the evaluation and development of cutting-edge...  ...will operate at the intersection of AI research, software engineering, and model evaluation, designing the benchmarks...  ..., engineers, and applied AI teams to design experiments and evaluate... 
    Full time
    Remote work

    SaidGig

    United States
    1 day ago
  • $80 - $105 per hour

     ...in shaping the future of legal AI. This part-time contractor...  ...influence how advanced AI is trained, evaluated, and utilized in real-world...  ...expert feedback to improve model performance and output precision...  ...with product and research teams to refine data, guidelines, and... 
    Hourly pay
    Contract work
    Part time
    For contractors
    Remote work

    SaidGig

    United States
    11 days ago
  • $80 - $105 per hour

     ...influence the development of advanced AI systems in the legal field....  ...expert feedback to enhance model performance and output precision. Create objective evaluation frameworks and grading criteria...  ...Collaborate with product and research teams to refine data, guidelines, and... 
    Hourly pay
    Contract work
    Part time
    For contractors
    Remote work

    SaidGig

    United States
    a month ago
  • $70 per hour

     ...Position: AI Model Assessment Specialist Type: Contract Compensation: $22 - $70/hour Location: Remote Commitment...  ...10-40 hrs/week Role Responsibilities Evaluate and critique the performance and...  ...keen observation. Collaborate with team members by communicating findings and... 
    Contract work
    Remote work

    Crossing Hurdles

    New York, NY
    4 days ago
  •  ...Radiology professionals can apply their expertise to evaluate and enhance AI models in the medical imaging field. This role involves assessing AI-...  ...diagnostic findings, and coordinating care across medical teams. Commitment to maintaining safety, quality, and professional... 
    Remote work
    Flexible hours

    SaidGig

    United States
    1 day ago
  • $40 per hour

     ...experienced cybersecurity professionals to join our team to help train AI models. In this role, you will evaluate AI-generated security content, solve technical...  ...experience in cybersecurity (e.g., penetration testing, red teaming, incident response, detection engineering,... 
    Hourly pay
    Full time
    Part time
    Remote work

    DataAnnotation

    Virginia, MN
    4 days ago
  • $40 per hour

    A leading AI training firm in the United States is seeking an R&D Biologist to join their team. In this remote position, you will evaluate AI chatbots and enhance their models while ensuring the biological accuracy of their outputs. The ideal candidate should have an expert... 
    Hourly pay
    Remote work

    DataAnnotation

    Hartford, CT
    4 days ago
  • $30 - $90 per hour

     ...collaborating with cutting-edge AI research. As an experienced...  ...alongside a high-caliber engineering team. Key Responsibilities:...  ...Actively test new AI-powered models in Cursor, providing actionable...  ...). Experience designing or evaluating experimental tooling and developer... 
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    1 day ago
  • $40 per hour

    A data solutions company is seeking a Process Development Chemist to join their team remotely. In this role, you will train AI models by evaluating their performance on complex chemistry questions. The position offers flexibility to work on chosen projects at an hourly... 
    Hourly pay
    Remote work

    DataAnnotation

    Jackson, MS
    4 days ago
  • $40 per hour

    A tech company specializing in AI training is looking for a Statistician to join their team. In this remote role, you'll train AI models by providing complex math problems and evaluating their outputs for quality and correctness. The ideal candidate will have strong mathematical... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Oregon, WI
    4 days ago
  • $40 per hour

    A leading data annotation company is seeking a Statistician to join their team. This remote role involves training AI models by posing complex mathematical problems, evaluating outputs, and assessing the model's performance. Candidates must be detail-oriented and proficient... 
    Hourly pay
    Remote work

    DataAnnotation

    New York, NY
    4 days ago
  •  ...Medical professionals can apply their expertise to contribute to AI research projects that enhance the understanding of workplace tasks and language in their field. This role involves evaluating AI model outputs, assessing content related to your profession, and providing... 
    Remote work
    Flexible hours

    SaidGig

    United States
    1 day ago
  • A technology company in the United States is looking for an R&D Biologist to join their team and train AI models. You will be responsible for evaluating the logic and outputs of AI chatbots, requiring an expert level of biology. The position is either full-time or part... 
    Hourly pay
    Full time
    Part time
    Remote work
    Flexible hours

    DataAnnotation

    Louisiana, MO
    4 days ago
  • $121.9k - $197.1k

     ...'ll Do: The Adversarial Red Team Associate Principal is responsible...  ...capabilities. Leverage AI/LLM tooling to accelerate...  ...owners. Perform threat modeling, security risk assessments, and...  ...operate - and the ability to evaluate how these technologies can accelerate... 
    Work at office
    Remote work
    2 days per week

    The Options Clearing Corporation

    Chicago, IL
    3 days ago
  • $40 per hour

     ...We are looking for a Process Development Chemist to join our team to train AI models. You will measure the progress of these AI chatbots, evaluate their logic, and solve problems to improve the quality of each model. In this role you will need to hold an expert level... 
    Hourly pay
    Full time
    Contract work
    Part time
    Remote work

    DataAnnotation

    Iowa, LA
    4 days ago
  • $60 per hour

    Join the DataAnnotation team and contribute to developing cutting-edge AI systems, while enjoying the flexibility of...  ...help advance AI development. AI models are increasingly capable of performing...  ...-the-art AI models on tasks like evaluating AI-generated quantitative... 
    Remote job
    Hourly pay
    Full time
    Flexible hours

    DataAnnotation

    Sioux Falls, SD
    3 days ago
  •  ...Adversarial Red Team Associate Principal *****THIS POSITION IS NOT...  ...capabilities. Leverage AI/LLM tooling to accelerate offensive...  ...IT owners. Perform threat modeling, security risk assessments, and...  ...operate – and the ability to evaluate how these technologies can... 
    Work at office

    Options Clearing Corporation

    Chicago, IL
    4 days ago
  • $40 per hour

    A leading AI development company is seeking experienced quantitative professionals to join their remote team. The role involves evaluating AI-generated work, designing quantitative problems, and providing impactful feedback. Candidates with at least 2 years of quantitative... 
    Remote job
    Hourly pay

    DataAnnotation

    New York, NY
    1 day ago
  • $40 per hour

    A leading AI development firm is seeking experienced quantitative professionals to evaluate AI-generated quantitative analysis and provide impactful feedback. This fully remote...  ...field and be comfortable with coding. Join a team that shapes the future of AI systems while... 
    Remote job
    Hourly pay
    Flexible hours

    DataAnnotation

    Honolulu, HI
    4 days ago
  • $40 per hour

    A forward-thinking AI team is seeking quantitative professionals to evaluate and improve cutting-edge AI systems. The role involves working on AI-generated analyses and providing critical feedback for model enhancement. Candidates should have a strong quantitative background... 
    Remote job
    Hourly pay
    Flexible hours

    DataAnnotation

    Helena, MT
    4 days ago
  • $40 per hour

    A data science team is seeking experienced quantitative professionals to evaluate AI-generated work and contribute to the development of cutting-edge AI systems. This fully remote position offers flexible scheduling and competitive hourly pay starting at $40+. Ideal candidates... 
    Remote job
    Hourly pay
    Flexible hours

    DataAnnotation

    Madison, WI
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to LLM Red Team Specialist for AI Model Evaluation. Be the first to apply!