LLM Red Team Specialist for AI Model Evaluation
$60 - $90 per hourSaidGig
Join a leading AI lab''s cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. Overview
A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models and needs specialists who can find where those models break. Working in a red-teaming setup, you will design and probe complex, multi-step tasks to expose vulnerabilities, edge cases, and failure modes in frontier AI systems, the places where a model looks competent but is quietly wrong.
Each task represents one to two days of continuous, focused effort and spans multiple technical skills: coding, experimentation, and careful analysis. You will work in a tight feedback loop with the lab''s researchers, turning the failure modes you find into stronger benchmark tasks. This role is fully remote within the United States, at approximately 35 hours per week.
Key Responsibilities- Probe models: Explore how frontier AI models behave on coding, ML, and analysis tasks, and find the spots where they quietly get things wrong.
- Design challenges: Turn the weaknesses you find into well-crafted tasks that are hard for models but fair to grade.
- Document findings: Write up what you discover clearly, with evidence and steps others can reproduce.
- Strengthen tasks: Team up with task authors to close loopholes, shortcuts, and grading gaps.
- Work as a team: Share insights with researchers and fellow experts so the benchmark keeps getting better.
- MSc or PhD in a STEM field, or equivalent practical experience in a research-heavy domain requiring data analysis and coding.
- 1+ years of experience in a research, research-engineering, security, or AI-evaluation role.
- Demonstrated ability to identify vulnerabilities, edge cases, or failure modes in LLMs or ML systems, through red teaming, adversarial testing, security research, or rigorous model evaluation.
- Working proficiency in Python and Git, with the ability to script your own probes and analyses.
- Strong familiarity with LLM capabilities, limitations, and evaluation techniques.
- Past experience in AI training, model evaluation, or benchmark/task authoring is preferred.
- A perfectionist mindset: high attention to detail, creativity in finding what others missed, strong written communication, and the ability to work independently through ambiguous, open-ended problems.
- Ability to engage reliably for approximately 35 hours per week.
This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce.
CompensationHourly compensation ranges from $60 to $90.
EligibilityThis role is fully remote within the United States.
Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives.
Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client''s internal teams, and integration into standard enterprise workflows.
Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, disability, or any other characteristic protected by applicable law.]]><
$40 - $65 per hour
...high-impact project focused on the evaluation and enhancement of frontier language models as an LLM Red-Teamer. In this role, you will... ...to train next-generation AI systems, shaping how models learn... ...specification. Maintain calibration with team leads and quality control...SuggestedRemote jobHourly payFor contractorsImmediate start$208k - $300k
...Machine Learning Engineer - Model Evaluations, Public Sector The Public Sector ML team at Scale deploys advanced AI systems—including LLMs, agentic... ...safety metrics, including LLM-judge–based evaluations.... ...and execute stress tests and red-teaming workflows to uncover...SuggestedFull time$350k
...Join a dynamic research team as a Member of Technical... ...in shaping the future of AI-powered legal reasoning.... ...of large language models, agentic systems, and legal... ...development of rigorous evaluation frameworks to measure and... ...Advanced degree in Law (JD, LLM, SJD, PhD in Law, or...SuggestedRemote jobFull time$15 - $25 per hour
...As a Legal Specialist, you will leverage your legal expertise... ...of next-generation AI systems. Your insights... ...role in shaping how these models learn, reason, and... ...support structured legal evaluations. Prepare concise written... ...with cross-functional teams to enhance legal...SuggestedHourly payContract workPart timeFor contractorsRemote work$161.8k - $184.6k
...Associate, Data Scientist - LLM Customization Team At Capital One, we think... ...for the Open Banking future. AI is transforming every industry... ...the power of Large Language Models (LLMs), adapt and finetune... ...from design through training, evaluation, and validation; partnering...SuggestedFull timePart timeLocal areaFlexible hours$125 per hour
...QGIS specialists leverage their expertise in geographic information... ...analysis to support AI research through... ...work. This role involves evaluating AI-generated content... ...prompts and evaluating LLM responses. Contribute... ...asynchronously with AI research teams. Work Terms This...Remote workFlexible hours$80 - $110 per hour
...Join a cutting-edge GenAI team at a leading AI lab, where your expertise will be pivotal in developing advanced AI models. This role focuses on designing and evaluating machine learning and natural language processing tasks that will help identify and address capability...Hourly payPart timeRemote work$80 - $110 per hour
...Join a cutting-edge GenAI team at a leading AI lab, where your expertise will be pivotal in developing advanced AI models. This role focuses on building and evaluating frontier models, requiring experienced computer vision practitioners to serve as ground-truth experts...Hourly payPart timeRemote work$30 - $90 per hour
...Developer, you will play a crucial role in evaluating and training next-generation AI coding tools during their highly... .... Test and evaluate alpha AI models in Cursor over multiple 4-day, 5+... .... Collaborate with the research team via Slack, providing real-time feedback...Remote jobHourly payContract workPart time$220k
...This role focuses on advancing the evaluation and development of cutting-edge... ...will operate at the intersection of AI research, software engineering, and model evaluation, designing the benchmarks... ..., engineers, and applied AI teams to design experiments and evaluate...Full timeRemote work$80 - $105 per hour
...in shaping the future of legal AI. This part-time contractor... ...influence how advanced AI is trained, evaluated, and utilized in real-world... ...expert feedback to improve model performance and output precision... ...with product and research teams to refine data, guidelines, and...Hourly payContract workPart timeFor contractorsRemote work$80 - $105 per hour
...influence the development of advanced AI systems in the legal field.... ...expert feedback to enhance model performance and output precision. Create objective evaluation frameworks and grading criteria... ...Collaborate with product and research teams to refine data, guidelines, and...Hourly payContract workPart timeFor contractorsRemote work$70 per hour
...Position: AI Model Assessment Specialist Type: Contract Compensation: $22 - $70/hour Location: Remote Commitment... ...10-40 hrs/week Role Responsibilities Evaluate and critique the performance and... ...keen observation. Collaborate with team members by communicating findings and...Contract workRemote work- ...Radiology professionals can apply their expertise to evaluate and enhance AI models in the medical imaging field. This role involves assessing AI-... ...diagnostic findings, and coordinating care across medical teams. Commitment to maintaining safety, quality, and professional...Remote workFlexible hours
$40 per hour
...experienced cybersecurity professionals to join our team to help train AI models. In this role, you will evaluate AI-generated security content, solve technical... ...experience in cybersecurity (e.g., penetration testing, red teaming, incident response, detection engineering,...Hourly payFull timePart timeRemote work$40 per hour
A leading AI training firm in the United States is seeking an R&D Biologist to join their team. In this remote position, you will evaluate AI chatbots and enhance their models while ensuring the biological accuracy of their outputs. The ideal candidate should have an expert...Hourly payRemote work$30 - $90 per hour
...collaborating with cutting-edge AI research. As an experienced... ...alongside a high-caliber engineering team. Key Responsibilities:... ...Actively test new AI-powered models in Cursor, providing actionable... ...). Experience designing or evaluating experimental tooling and developer...Hourly payContract workRemote work$40 per hour
A data solutions company is seeking a Process Development Chemist to join their team remotely. In this role, you will train AI models by evaluating their performance on complex chemistry questions. The position offers flexibility to work on chosen projects at an hourly...Hourly payRemote work$40 per hour
A tech company specializing in AI training is looking for a Statistician to join their team. In this remote role, you'll train AI models by providing complex math problems and evaluating their outputs for quality and correctness. The ideal candidate will have strong mathematical...Hourly payRemote workFlexible hours$40 per hour
A leading data annotation company is seeking a Statistician to join their team. This remote role involves training AI models by posing complex mathematical problems, evaluating outputs, and assessing the model's performance. Candidates must be detail-oriented and proficient...Hourly payRemote work- ...Medical professionals can apply their expertise to contribute to AI research projects that enhance the understanding of workplace tasks and language in their field. This role involves evaluating AI model outputs, assessing content related to your profession, and providing...Remote workFlexible hours
- A technology company in the United States is looking for an R&D Biologist to join their team and train AI models. You will be responsible for evaluating the logic and outputs of AI chatbots, requiring an expert level of biology. The position is either full-time or part...Hourly payFull timePart timeRemote workFlexible hours
$121.9k - $197.1k
...'ll Do: The Adversarial Red Team Associate Principal is responsible... ...capabilities. Leverage AI/LLM tooling to accelerate... ...owners. Perform threat modeling, security risk assessments, and... ...operate - and the ability to evaluate how these technologies can accelerate...Work at officeRemote work2 days per week$40 per hour
...We are looking for a Process Development Chemist to join our team to train AI models. You will measure the progress of these AI chatbots, evaluate their logic, and solve problems to improve the quality of each model. In this role you will need to hold an expert level...Hourly payFull timeContract workPart timeRemote work$60 per hour
Join the DataAnnotation team and contribute to developing cutting-edge AI systems, while enjoying the flexibility of... ...help advance AI development. AI models are increasingly capable of performing... ...-the-art AI models on tasks like evaluating AI-generated quantitative...Remote jobHourly payFull timeFlexible hours- ...Adversarial Red Team Associate Principal *****THIS POSITION IS NOT... ...capabilities. Leverage AI/LLM tooling to accelerate offensive... ...IT owners. Perform threat modeling, security risk assessments, and... ...operate – and the ability to evaluate how these technologies can...Work at office
$40 per hour
A leading AI development company is seeking experienced quantitative professionals to join their remote team. The role involves evaluating AI-generated work, designing quantitative problems, and providing impactful feedback. Candidates with at least 2 years of quantitative...Remote jobHourly pay$40 per hour
A leading AI development firm is seeking experienced quantitative professionals to evaluate AI-generated quantitative analysis and provide impactful feedback. This fully remote... ...field and be comfortable with coding. Join a team that shapes the future of AI systems while...Remote jobHourly payFlexible hours$40 per hour
A forward-thinking AI team is seeking quantitative professionals to evaluate and improve cutting-edge AI systems. The role involves working on AI-generated analyses and providing critical feedback for model enhancement. Candidates should have a strong quantitative background...Remote jobHourly payFlexible hours$40 per hour
A data science team is seeking experienced quantitative professionals to evaluate AI-generated work and contribute to the development of cutting-edge AI systems. This fully remote position offers flexible scheduling and competitive hourly pay starting at $40+. Ideal candidates...Remote jobHourly payFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to LLM Red Team Specialist for AI Model Evaluation. Be the first to apply!
- amazon specialist United States
- specialist United States
- associate specialist United States
- ammunition specialist United States
- correctional behavioral specialist United States
- entry level safety specialist United States
- disclosure specialist United States
- erp specialist United States
- infection control specialist United States
- cash reconciliation specialist United States



