AI Evaluation Specialist
Weekday
Join a pioneering AI initiative focused on building next-generation evaluation benchmarks for frontier AI models. We are seeking analytical and technically skilled professionals to identify where advanced AI systems fail in subtle, real-world scenarios. Working in a red-teaming environment, you will design challenging, multi-step tasks that expose hidden vulnerabilities, reasoning gaps, and edge cases that traditional evaluations often miss.
In this role, you'll collaborate closely with AI researchers to transform discovered failure modes into high-quality benchmark tasks that improve the robustness, safety, and reasoning capabilities of state-of-the-art AI systems.
This is a fully remote, full-time engagement requiring approximately 35 hours per week .
Requirements
Key Responsibilities
- Investigate how frontier AI models perform across coding, machine learning, analytical reasoning, and complex problem-solving tasks.
- Identify hidden failure modes, edge cases, reasoning errors, and vulnerabilities that may not be apparent through standard testing.
- Design challenging evaluation tasks that accurately measure AI capabilities while remaining objective and reproducible.
- Document findings with clear technical explanations, supporting evidence, and reproducible methodologies.
- Collaborate with benchmark designers and AI researchers to refine evaluation tasks, eliminate loopholes, and strengthen grading criteria.
- Share insights and recommendations with cross-functional teams to continuously improve AI evaluation quality and benchmark coverage.
Required Qualifications
- Master's degree, PhD, or equivalent practical experience in a STEM discipline involving research, coding, or advanced data analysis.
- Minimum 1 year of experience in AI research, research engineering, security research, AI evaluation, or a related technical field.
- Demonstrated experience identifying vulnerabilities, adversarial behaviors, edge cases, or failure modes in Large Language Models or other machine learning systems.
- Strong proficiency in Python and Git , with the ability to build custom scripts for experimentation, testing, and analysis.
- Solid understanding of modern Large Language Models, their strengths, limitations, and evaluation methodologies.
- Experience with AI benchmarking, model evaluation, adversarial testing, prompt engineering, or dataset creation is highly desirable.
- Excellent analytical thinking, creativity, and attention to detail, with the ability to solve ambiguous, open-ended problems independently.
- Outstanding written communication skills for documenting technical findings clearly and accurately.
- Ability to commit approximately 35 hours per week on a consistent basis.
Preferred Qualifications
- Experience with AI safety, red teaming, adversarial machine learning, or security research.
- Background in benchmark design, evaluation framework development, or AI quality assurance.
- Experience creating reproducible technical experiments and documenting complex failure analyses.
- Familiarity with frontier AI research methodologies and model capability assessments.
Why Join
- Help shape the future of AI evaluation by identifying critical weaknesses before they reach production.
- Work on cutting-edge AI systems alongside researchers developing next-generation language models.
- Apply your technical expertise to improve AI reliability, reasoning, and robustness.
- Contribute directly to benchmark development that influences the evolution of advanced AI technologies.
- Enjoy the flexibility of a fully remote engagement while working on impactful research initiatives.
Equal Opportunity
We are committed to fostering an inclusive and diverse environment where all qualified applicants receive equal consideration. Reasonable accommodations are available throughout the application and engagement process.
Contract & Engagement Details
- Independent contractor engagement.
- Fully remote with flexible working hours.
- Expected commitment of approximately 35 hours per week .
- Project duration may be extended, shortened, or concluded based on project requirements and individual performance.
- Work does not require access to confidential or proprietary information from any current or former employer.
- Payments are issued weekly based on approved work completed.
- At this time, we are unable to support H1-B or STEM OPT candidates.
- ...To support AI development, the part-time AI Evaluation Specialist will review and assess AI-generated outputs for quality and usability while collaborating with teams to refine evaluation standards in a remote contract role. Key responsibilities Review and critically...SuggestedContract workPart timeRemote work
$130 - $180 per hour
Role Description Join a leading AI research initiative focused on advancing healthcare-focused artificial intelligence. We are seeking... ...to contribute their clinical expertise toward training and evaluating next-generation AI models capable of sophisticated medical reasoning...SuggestedWeekly payContract workPart timeFor contractorsRemote workFlexible hours- ...Join a pioneering AI initiative focused on building next-generation evaluation benchmarks for frontier AI models. We are seeking analytical and technically skilled professionals to identify where advanced AI systems fail in subtle, real-world scenarios. Working in a red...SuggestedFull timeContract workFor contractorsRemote workFlexible hours
- ...Join a fast-paced AI evaluation initiative supporting one of the world's leading AI research organizations. We are seeking detail-oriented professionals to evaluate AI-generated outputs by applying structured grading rubrics with precision and consistency. This is...SuggestedTemporary workImmediate start
$25 - $30 per hour
...Bilingual Traditional Chinese AI Evaluation Specialist is a remote Chinese specialist track for evaluating chinese evaluation outputs against native-speaker standards. Reviewers spot fluency, register, and cultural-context errors that automated checks miss, and write...SuggestedFor contractorsRemote work10 hours per week- ...create role-play scenarios across domains such as travel, financial services, telecoms and technical support, contributing to diverse evaluation datasets. Responsibilities include evaluating performance with metrics on task completion, naturalness, audio comprehension, and...Remote jobContract work
- AuraOne is seeking a Physical Sciences Research Assistant for a Remote AI Evaluation track. You will review AI outputs in physics, reproduce key derivations, and document correct methods to help train modeling systems. This contractor-style role emphasizes rigorous reasoning...Remote jobPart timeFor contractors
- A leading AI research accelerator is seeking Geospatial Experts to enhance AI systems through evaluations and real-world applications. This entry-level contractor position is fully remote with flexible hours, primarily focusing on geospatial reasoning tasks. Responsibilities...Remote jobFor contractorsFlexible hours
- A tech company specializing in AI projects is seeking skilled LibreSprite users to assist in evaluating AI-generated visual content. As an independent contractor, you can work flexibly from anywhere, contributing around 5-20 hours per week depending on project needs. Ideal...Remote jobFor contractors
$80 per hour
...Science Expert with Python experience for part-time, remote projects. This role focuses on tasks related to AI systems, including designing problems, evaluating solutions, and validating calculations. Ideal candidates will have a degree in Computer Science, proficiency...Remote jobHourly payPart time- OpenTrain AI seeks a Video Game AI Evaluation Expert on a contractor, part-time basis. You will design prompts for game development, esports, platforms, and communities; assess AI responses for accuracy, completeness and nuance; and develop evaluation datasets with clear...Remote jobPart timeFor contractors
$20 - $26 per hour
Prolific is seeking fluent Kannada speakers to act as evaluators who compare text and voice samples to assess naturalness and authenticity. You will listen to audio clips, rate quality, and flag any mismatches in tone or pronunciation, with emphasis on cultural context...Remote jobFlexible hours- A leading research accelerator is seeking a Geospatial Expert to enhance AI systems through advanced geospatial analysis. This entry-level, remote role involves evaluating geospatial datasets and supporting tasks aligned with crisis management and agriculture. Candidates...Remote jobContract work
- Dorado is seeking Speech AI Evaluation Specialists to support AI content improvement. This freelance, part-time role is based remotely from Malaysia, with 10+ hours per week and a starting date immediately. You will evaluate Vietnamese-language responses and provide structured...Remote jobPart timeFreelanceImmediate start10 hours per week
$150k - $250k
...About Distyl AI Distyl is an applied AI technology company partnering with the world’s most ambitious institutions to rearchitect... ...What We Are Looking For At Distyl, we build AI systems using Evaluation-Driven Development —an approach where evaluation is not an afterthought...Full timeWork at officeFlexible hours3 days per week- ...Prolific is seeking Product Designers and UX Specialists to join our Expert Network, contributing to the training and evaluation of cutting-edge AI models. This role requires expertise in usability, design systems, and user research, providing essential feedback on design...Remote workWork from homeFlexible hours
- ...Prolific is seeking talented Product Designers and UX Specialists to join our Expert Network for training and evaluation of cutting-edge AI models. Candidates should have a relevant educational background and at least one year of experience in design-related fields. Responsibilities...Remote workWork from homeFlexible hours
$60 - $85 per hour
...About the job Remote | Licensed Chemical Engineer & AI Evaluation Specialist - $60-$85/hour We are sharing a specialised part-time consulting opportunity for licensed US chemical engineers with professional experience in process design, process safety, plant operations...Hourly payFull timeContract workPart timeRemote work10 hours per weekFlexible hours- A leading AI research platform is seeking an AI Trainer with expertise in Graphic and Visual Design. The role involves evaluating designs, ensuring they meet professional standards, and providing valuable insights for AI development. Candidates should have formal qualifications...Remote jobFlexible hours
- ...Human Resources Expert for a remote contract in the United States. You will design realistic HR scenarios, craft model prompts, and evaluate AI-generated HR outputs while ensuring alignment with employment laws and best practices. You will collaborate with the team to...Remote jobContract work
$80 per hour
A technology firm is seeking QAs for autonomous AI agents to validate and improve task structures within a new project. Candidates... ...analytical thinkers with strong attention to detail and experience in evaluating scenarios. This flexible, project-based role offers competitive...Remote jobFlexible hours- Rex.zone is seeking a remote, full-time Senior AI Data Annotation role supporting training-data quality for modern AI/ML systems. You will deliver high-precision human feedback across RLHF, LLM evaluation, prompt evaluation, and QA evaluation to drive measurable model performance...Remote jobFull time
$20 per hour
Prolific is seeking a Fluent Gujarati Speaker for an AI Trainer role, where you will evaluate AI-generated content in Gujarati. This remote freelance position offers competitive pay of up to $20/hr. Ideal candidates will possess linguistic nuance and fluent English communication...Remote jobFreelanceFlexible hours- A forward-thinking tech company is seeking an AI Trainer specializing in visual and graphic design to evaluate AI outputs. Responsibilities include assessing design quality and providing feedback to ensure high professional standards. Applicants should have formal education...Remote jobFlexible hours
$80 per hour
...part-time opportunity focused on quality assurance for autonomous AI agents. You will analyze complex systems, review tasks for logic... ...and detail-oriented skills, with experience in policy evaluation or logic puzzles preferred. Compensation can reach up to $80/hour...Remote jobPart timeFlexible hours$60 per hour
...seeking contributors for a part-time QA project focused on autonomous AI agents. This flexible remote opportunity requires strong... ...familiarity with structured data formats. Candidates will review evaluation tasks, identify inconsistencies, and help define expected AI behaviors...Remote jobPart timeFlexible hours$80 - $120 per hour
Role Description ~Evaluate AI-generated artifacts against domain-specific quality rubrics. ~Identify factual, aesthetic, and presentation errors in documents, spreadsheets, and slide decks. ~Provide clear, structured written feedback to improve AI outputs. ~Apply...Part timeWork at officeRemote work- ...Summary This is a fully remote, hourly contractor role supporting AI data and language projects on a project-based, flexible hour... ..., and other content to support AI training datasets. LLM evaluation: reviewing AI-generated responses for accuracy, reasoning quality...Hourly payFor contractorsRemote workFlexible hours
- **Job Title: AI Trainer || Image Quality Evaluator || English** **Location**: Remote | Work from Home **Employment Type:** Project-based | Contract... ...a multilingual AI data Annotation and Transcription Specialists with strong proficiency in English. In this role, you will...Contract workRemote workWork from homeMonday to FridayDay shift
$20 - $80 per hour
...Role Overview Train and evaluate next-generation AI systems by scoring model outputs, annotating real-world content, and delivering clear, actionable feedback that improves model accuracy and reasoning across diverse domains. About the company micro1 is an AI data...Hourly payFor contractorsRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Evaluation Specialist. Be the first to apply!





