AI Evaluation Specialist
Weekday
Join a pioneering AI initiative focused on building next-generation evaluation benchmarks for frontier AI models. We are seeking analytical and technically skilled professionals to identify where advanced AI systems fail in subtle, real-world scenarios. Working in a red-teaming environment, you will design challenging, multi-step tasks that expose hidden vulnerabilities, reasoning gaps, and edge cases that traditional evaluations often miss.
In this role, you'll collaborate closely with AI researchers to transform discovered failure modes into high-quality benchmark tasks that improve the robustness, safety, and reasoning capabilities of state-of-the-art AI systems.
This is a fully remote, full-time engagement requiring approximately 35 hours per week .
Requirements
Key Responsibilities
- Investigate how frontier AI models perform across coding, machine learning, analytical reasoning, and complex problem-solving tasks.
- Identify hidden failure modes, edge cases, reasoning errors, and vulnerabilities that may not be apparent through standard testing.
- Design challenging evaluation tasks that accurately measure AI capabilities while remaining objective and reproducible.
- Document findings with clear technical explanations, supporting evidence, and reproducible methodologies.
- Collaborate with benchmark designers and AI researchers to refine evaluation tasks, eliminate loopholes, and strengthen grading criteria.
- Share insights and recommendations with cross-functional teams to continuously improve AI evaluation quality and benchmark coverage.
Required Qualifications
- Master's degree, PhD, or equivalent practical experience in a STEM discipline involving research, coding, or advanced data analysis.
- Minimum 1 year of experience in AI research, research engineering, security research, AI evaluation, or a related technical field.
- Demonstrated experience identifying vulnerabilities, adversarial behaviors, edge cases, or failure modes in Large Language Models or other machine learning systems.
- Strong proficiency in Python and Git , with the ability to build custom scripts for experimentation, testing, and analysis.
- Solid understanding of modern Large Language Models, their strengths, limitations, and evaluation methodologies.
- Experience with AI benchmarking, model evaluation, adversarial testing, prompt engineering, or dataset creation is highly desirable.
- Excellent analytical thinking, creativity, and attention to detail, with the ability to solve ambiguous, open-ended problems independently.
- Outstanding written communication skills for documenting technical findings clearly and accurately.
- Ability to commit approximately 35 hours per week on a consistent basis.
Preferred Qualifications
- Experience with AI safety, red teaming, adversarial machine learning, or security research.
- Background in benchmark design, evaluation framework development, or AI quality assurance.
- Experience creating reproducible technical experiments and documenting complex failure analyses.
- Familiarity with frontier AI research methodologies and model capability assessments.
Why Join
- Help shape the future of AI evaluation by identifying critical weaknesses before they reach production.
- Work on cutting-edge AI systems alongside researchers developing next-generation language models.
- Apply your technical expertise to improve AI reliability, reasoning, and robustness.
- Contribute directly to benchmark development that influences the evolution of advanced AI technologies.
- Enjoy the flexibility of a fully remote engagement while working on impactful research initiatives.
Equal Opportunity
We are committed to fostering an inclusive and diverse environment where all qualified applicants receive equal consideration. Reasonable accommodations are available throughout the application and engagement process.
Contract & Engagement Details
- Independent contractor engagement.
- Fully remote with flexible working hours.
- Expected commitment of approximately 35 hours per week .
- Project duration may be extended, shortened, or concluded based on project requirements and individual performance.
- Work does not require access to confidential or proprietary information from any current or former employer.
- Payments are issued weekly based on approved work completed.
- At this time, we are unable to support H1-B or STEM OPT candidates.
- ...To support AI development, the part-time AI Evaluation Specialist will review and assess AI-generated outputs for quality and usability while collaborating with teams to refine evaluation standards in a remote contract role. Key responsibilities Review and critically...SuggestedContract workPart timeRemote work
- A tech company specializing in AI projects is seeking skilled LibreSprite users to assist in evaluating AI-generated visual content. As an independent contractor, you can work flexibly from anywhere, contributing around 5-20 hours per week depending on project needs. Ideal...SuggestedFor contractorsRemote work
- ...Join a fast-paced AI evaluation initiative supporting one of the world's leading AI research organizations. We are seeking detail-oriented professionals to evaluate AI-generated outputs by applying structured grading rubrics with precision and consistency. This is...SuggestedTemporary workImmediate start
- ...Join a pioneering AI initiative focused on building next-generation evaluation benchmarks for frontier AI models. We are seeking analytical and technically skilled professionals to identify where advanced AI systems fail in subtle, real-world scenarios. Working in a red...SuggestedFull timeContract workFor contractorsRemote workFlexible hours
$25 - $30 per hour
...Bilingual German AI Evaluation Specialist is a remote German specialist track for evaluating german evaluation outputs against native-speaker standards. Reviewers spot fluency, register, and cultural-context errors that automated checks miss, and write structured rationale...SuggestedFor contractorsRemote work10 hours per week- Role Description Join a leading AI research initiative focused on advancing healthcare-focused artificial intelligence. We are seeking... ...to contribute their clinical expertise toward training and evaluating next-generation AI models capable of sophisticated medical reasoning...Weekly payContract workPart timeFor contractorsRemote workFlexible hours
- ...create role-play scenarios across domains such as travel, financial services, telecoms and technical support, contributing to diverse evaluation datasets. Responsibilities include evaluating performance with metrics on task completion, naturalness, audio comprehension, and...Remote jobContract work
- AuraOne is seeking a Physical Sciences Research Assistant for a Remote AI Evaluation track. You will review AI outputs in physics, reproduce key derivations, and document correct methods to help train modeling systems. This contractor-style role emphasizes rigorous reasoning...Remote jobPart timeFor contractors
$20 - $26 per hour
Prolific is seeking fluent Kannada speakers to act as evaluators who compare text and voice samples to assess naturalness and authenticity. You will listen to audio clips, rate quality, and flag any mismatches in tone or pronunciation, with emphasis on cultural context...Remote jobFlexible hours- A leading research accelerator is seeking a Geospatial Expert to enhance AI systems through advanced geospatial analysis. This entry-level, remote role involves evaluating geospatial datasets and supporting tasks aligned with crisis management and agriculture. Candidates...Remote jobContract work
- AuraOne is seeking a Turkish Bilingual AI Evaluation Expert (Remote) to review Turkish prompts and outputs against our quality rubric. This independent contractor role requires assessing model responses, labeling issues, and providing structured feedback for retraining....Remote jobFor contractors
- AuraOne seeks a Hebrew Bilingual Expert — AI Evaluation & Annotation (Remote) to review Hebrew prompts and responses against AuraOne's quality rubric, labeling edge cases and providing structured feedback to retrain models. In this remote contractor role, you will compare...Remote jobFor contractors
- OpenTrain AI seeks a Video Game AI Evaluation Expert on a contractor, part-time basis. You will design prompts for game development, esports, platforms, and communities; assess AI responses for accuracy, completeness and nuance; and develop evaluation datasets with clear...Remote jobPart timeFor contractors
$80 per hour
...Science Expert with Python experience for part-time, remote projects. This role focuses on tasks related to AI systems, including designing problems, evaluating solutions, and validating calculations. Ideal candidates will have a degree in Computer Science, proficiency...Remote jobHourly payPart time- A leading AI research accelerator is seeking Geospatial Experts to enhance AI systems through evaluations and real-world applications. This entry-level contractor position is fully remote with flexible hours, primarily focusing on geospatial reasoning tasks. Responsibilities...Remote jobFor contractorsFlexible hours
$60 per hour
...seeking contributors for a part-time QA project focused on autonomous AI agents. This flexible remote opportunity requires strong... ...familiarity with structured data formats. Candidates will review evaluation tasks, identify inconsistencies, and help define expected AI behaviors...Part timeRemote workFlexible hours$80 per hour
...part-time opportunity focused on quality assurance for autonomous AI agents. You will analyze complex systems, review tasks for logic... ...and detail-oriented skills, with experience in policy evaluation or logic puzzles preferred. Compensation can reach up to $80/hour...Part timeRemote workFlexible hours- Dorado is seeking Speech AI Evaluation Specialists to support AI content improvement. This freelance, part-time role is based remotely from Malaysia, with 10+ hours per week and a starting date immediately. You will evaluate Vietnamese-language responses and provide structured...Remote jobPart timeFreelanceImmediate start10 hours per week
- ...Prolific is seeking Product Designers and UX Specialists to join our Expert Network, contributing to the training and evaluation of cutting-edge AI models. This role requires expertise in usability, design systems, and user research, providing essential feedback on design...Remote workWork from homeFlexible hours
- ...Prolific is seeking talented Product Designers and UX Specialists to join our Expert Network for training and evaluation of cutting-edge AI models. Candidates should have a relevant educational background and at least one year of experience in design-related fields. Responsibilities...Remote workWork from homeFlexible hours
$20 - $60 per hour
...Role Overview Help train and evaluate next-generation AI systems by creating rigorous, real-world assessments that test how advanced models learn, reason, and perform. This remote contract opportunity is open to recent graduates and other researchers and writers with...Hourly payContract workFor contractorsRemote work$60 - $85 per hour
...About the job Remote | Licensed Chemical Engineer & AI Evaluation Specialist - $60-$85/hour We are sharing a specialised part-time consulting opportunity for licensed US chemical engineers with professional experience in process design, process safety, plant operations...Hourly payFull timeContract workPart timeRemote work10 hours per weekFlexible hours- AuraOne is seeking an Evaluation Harness Model Evaluation Specialist to work remotely as an independent contractor. You will review evaluation harness model outputs, label issues, and provide structured feedback to retrain the model using AuraOne's quality rubric. Strong...Remote jobFor contractors
- A leading AI research platform is seeking an AI Trainer with expertise in Graphic and Visual Design. The role involves evaluating designs, ensuring they meet professional standards, and providing valuable insights for AI development. Candidates should have formal qualifications...Remote jobFlexible hours
$20 per hour
Prolific is seeking a Fluent Gujarati Speaker for an AI Trainer role, where you will evaluate AI-generated content in Gujarati. This remote freelance position offers competitive pay of up to $20/hr. Ideal candidates will possess linguistic nuance and fluent English communication...Remote jobFreelanceFlexible hours$80 per hour
A technology firm is seeking QAs for autonomous AI agents to validate and improve task structures within a new project. Candidates... ...analytical thinkers with strong attention to detail and experience in evaluating scenarios. This flexible, project-based role offers competitive...Remote jobFlexible hours$60 per hour
A tech innovation company is seeking QAs for autonomous AI agents to validate and improve task structures and evaluate logic. The role requires excellent analytical thinking, attention to detail, and the ability to assess complex scenarios. Successful candidates can work...Remote jobFlexible hours- ....zone is seeking a Remote Data Labeling Specialist to work from anywhere within the United... ...will label and review multi-modal data for AI training, including text, images, audio,... ..., content safety labeling, and RLHF-style evaluation tasks. #J-18808-Ljbffr Rex.zoneRemote job
- Rex.zone is seeking a remote, full-time Senior AI Data Annotation role supporting training-data quality for modern AI/ML systems. You will deliver high-precision human feedback across RLHF, LLM evaluation, prompt evaluation, and QA evaluation to drive measurable model performance...Remote jobFull time
- ...Human Resources Expert for a remote contract in the United States. You will design realistic HR scenarios, craft model prompts, and evaluate AI-generated HR outputs while ensuring alignment with employment laws and best practices. You will collaborate with the team to...Remote jobContract work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Evaluation Specialist. Be the first to apply!


