AI Evaluation Specialist
Crossing Hurdles
A leading AI organization in Australia is seeking individuals with strong writing and analytical skills to evaluate and improve AI outputs. The ideal candidate must possess the ability to assess emotional nuances and detail while adhering to structured guidelines. Responsibilities include ensuring clarity and empathy in evaluations, identifying weaknesses in AI reasoning, and contributing to a safe AI environment. Immediate availability post-onboarding is required for this role. #J-18808-Ljbffr Crossing Hurdles
- Dorado is seeking Speech AI Evaluation Specialists to support AI content improvement. This freelance, part-time role is based remotely from Malaysia, with 10+ hours per week and a starting date immediately. You will evaluate Vietnamese-language responses and provide structured...SuggestedRemote jobPart timeFreelanceImmediate start10 hours per week
- Mercor is seeking experienced music producers and audio engineers to evaluate generative music AI models. You will assess AI-generated tracks, rate musicality, creativity, and production quality, and label songs by genre and instruments. The role requires native Norwegian...SuggestedImmediate startFlexible hours
- Dorado is seeking language professionals for a six‑month, independent contractor engagement to contribute to an AI evaluation project. Review AI‑generated content, annotate language data, and provide feedback to improve accuracy and reliability. The role requires reliability...SuggestedFor contractors
- AuraOne is seeking a Physical Sciences Research Assistant for a Remote AI Evaluation track. You will review AI outputs in physics, reproduce key derivations, and document correct methods to help train modeling systems. This contractor-style role emphasizes rigorous reasoning...SuggestedRemote jobPart timeFor contractors
- ...create role-play scenarios across domains such as travel, financial services, telecoms and technical support, contributing to diverse evaluation datasets. Responsibilities include evaluating performance with metrics on task completion, naturalness, audio comprehension, and...SuggestedRemote jobContract work
- A tech company specializing in AI projects is seeking skilled LibreSprite users to assist in evaluating AI-generated visual content. As an independent contractor, you can work flexibly from anywhere, contributing around 5-20 hours per week depending on project needs. Ideal...Remote jobFor contractors
$60 - $85 per hour
...About the job Remote | Licensed Chemical Engineer & AI Evaluation Specialist - $60-$85/hour We are sharing a specialised part-time consulting opportunity for licensed US chemical engineers with professional experience in process design, process safety, plant operations...Hourly payFull timeContract workPart timeRemote work10 hours per weekFlexible hours- Mercor is seeking experienced music producers and audio engineers to evaluate generative music AI models in partnership with a leading AI lab. You will assess AI-generated music across a wide range of genres, rate it against detailed quality standards, and perform labeling...Flexible hours
- ...Human Resources Expert for a remote contract in the United States. You will design realistic HR scenarios, craft model prompts, and evaluate AI-generated HR outputs while ensuring alignment with employment laws and best practices. You will collaborate with the team to...Remote jobContract work
- Welo Data is hiring Data Labeling Associates in New York City for Project Perseus. The role involves evaluating Arabic language AI outputs and ensuring AI safety, requiring professional proficiency in Arabic and experience in writing and AI safety. You’ll critique models...
- Welo Data is seeking a Data Labeling Associate in New York City to evaluate AI model outputs, providing structured feedback and insights. This role demands native-level Portuguese proficiency and at least a bachelor's degree, along with strong critical thinking and communication...Full timeRemote work
$60 per hour
Prolific seeks Computer Science Specialists to join their Expert Network for evaluating and training AI models. Tasks include comparing AI-generated responses, reviewing scientific papers, and fact-checking technical data. Candidates should have a minimum of a BSc in Computer...Hourly pay- CNTXT AI is seeking a fully remote, hourly contractor to support AI data and language projects on a flexible, project-based schedule. The role involves content generation, data annotation, LLM evaluation, and localization QA across diverse topics. Ideal candidates are native...Remote jobHourly payFor contractorsFlexible hours
- About the Opportunity A leading AI research organization is seeking advanced LLM power users with strong experience using MCP and (... ...connectors for real-world personal life tasks. This project focuses on evaluating how well AI systems handle personalized, multi-step life tasks...Trial period
- Cincinnatus LLC, the employer of record, is seeking a seasoned Retail SME to help train AI models by evaluating outputs against rubrics and applying real-world retail judgment. The role requires 8+ years in retail and the ability to commit at least 35 hours per week, with...
- ...Job Title AI Evaluation Engineer Location Hybrid / Remote Employment Type Full-time Job Summary We are seeking an AI Evaluation Engineer to design, implement, and maintain evaluation frameworks for AI and machine learning...Full timeRemote work
$80 per hour
...CV in English and indicate your level of English proficiency. Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment...Permanent employmentTemporary work- A technology consulting firm is seeking a Supply Chain & Data Systems AI Trainer to evaluate AI-generated responses for supply-chain scenarios. This remote, part-time role requires 5-8 years of experience in supply-chain operations and strong analytical writing skills....Remote jobContract workPart time
$80 per hour
...human intelligence to ethically shape the future of AI. What We Do The Mindrift platform connects specialists with AI projects from major tech innovators. Our... ...for someone who can design realistic and structured evaluation scenarios for LLM-based agents. You'll create test...Part timeFreelanceRemote workFlexible hours$229.9k - $262.4k
...Senior Lead AI Engineer (SDK's: Gen AI Evaluation and MCP) Overview At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized...Full timePart timeLocal area$140 - $150 per hour
...-time consulting opportunity for experienced corporate and M&A lawyers qualified to practise in France. The role supports an AI legal evaluation project benchmarking AI responses to French law questions. You will review AI-generated answers, assess legal reasoning, and...Remote jobPart time$50 - $100 per hour
...progressive technology company is looking for expert biologists to join their team part-time. Responsibilities include annotating and evaluating AI tool outputs, generating evaluation tasks, and conducting literature searches. Candidates should have a PhD or Masters in...Hourly payPart time- ...Member of Technical Staff - Evals located in New York, NY, where you will ensure our AI-powered features are high quality and reliable. Your responsibilities include designing evaluation frameworks, building automated tests, and developing tools for seamless evaluations....
- Scale AI, Inc. is looking for a Research Scientist focused on Frontier Risk Evaluations to develop evaluation measures and datasets for assessing AI risks. As a key team member, you will engage with government agencies and publish methodologies that influence AI safety...
- Omaze is looking for a Senior Applied AI Scientist to join their highly skilled Tech team. This role focuses on AI system evaluation, measurement, and optimization, combining data science with product engineering. You will work on evaluation frameworks for LLM features...
$50 - $70 per hour
United States Digital Space LLC is seeking a Data Scientist for AI Evaluation Analytics to work as a remote contractor on long-term projects. You will develop evaluation metrics, validate datasets, analyze signals, and apply statistical methods to measure AI system performance...Remote jobHourly payFor contractors$152k - $241.5k
...us! We believe open-weight models are foundational to American AI leadership and cybersecurity, and that trust in AI grows through... ...scrutiny. Our AI Safety & Security Engineering team builds and evaluates AI-powered tooling that helps find, validate, and patch software...Full timeRemote work$30 - $50 per hour
A tech company is seeking an AI Researcher to support end-to-end research for modern AI systems. This remote role involves designing experiments, defining evaluation protocols, and improving evaluation rigor for large language models. Key responsibilities include developing...Remote jobHourly pay- Rex.zone is seeking an AI Research Scientist to lead applied AI research projects for US-based customers, translating open-ended questions into measurable experiments in LLM evaluation and RLHF data design. You will evaluate prompts, design datasets, and work with cross...Remote jobHourly payFlexible hours
- Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models. You will assess AI-generated music across genres and rate it against detailed quality standards, working in Hindi and English. Responsibilities include head-to-head...Remote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Evaluation Specialist. Be the first to apply!


