AI Evaluation Specialist
Crossing Hurdles
A leading AI organization in Australia is seeking individuals with strong writing and analytical skills to evaluate and improve AI outputs. The ideal candidate must possess the ability to assess emotional nuances and detail while adhering to structured guidelines. Responsibilities include ensuring clarity and empathy in evaluations, identifying weaknesses in AI reasoning, and contributing to a safe AI environment. Immediate availability post-onboarding is required for this role. #J-18808-Ljbffr Crossing Hurdles
- ....zone is seeking a Remote Data Labeling Specialist to work from anywhere within the United... ...will label and review multi-modal data for AI training, including text, images, audio,... ...segmentation, content safety labeling, and RLHF-style evaluation tasks. #J-18808-Ljbffr REXSuggestedRemote job
- Welo Data is hiring Data Labeling Associates in New York City for Project Perseus. The role involves evaluating Arabic language AI outputs and ensuring AI safety, requiring professional proficiency in Arabic and experience in writing and AI safety. You’ll critique models...Suggested
$11 - $30.65 per hour
Meridial is seeking contractors to evaluate advanced agentic audio models by simulating realistic customer service interactions across multiple domains. You will contribute to developing diverse datasets and assess model performance using various metrics. The role requires...SuggestedRemote jobHourly payFor contractors- Obsidian is collaborating with AI labs to find experienced health insurance professionals to enhance AI systems related to coverage... ...assess AI performance on health insurance tasks. The role includes evaluating AI outputs, creating health insurance scenarios, and providing...Suggested
$80 per hour
...human intelligence to ethically shape the future of AI. What We Do The Mindrift platform connects specialists with AI projects from major tech innovators. Our... ...for someone who can design realistic and structured evaluation scenarios for LLM-based agents. You'll create test...SuggestedPart timeFreelanceRemote workFlexible hours- YO AI Labs is seeking an experienced Management Consultant to support AI training and evaluation projects. This contractor role emphasizes structured problem-solving, business analysis, and consulting expertise to assess and improve AI model performance. Remote work enables...Remote jobFor contractors
- We are looking for an AI Evaluation Scientist to design and execute evaluation processes that ensure our predictive and generative AI systems are accurate, reliable, safe, and aligned with mission requirements. This role is essential for establishing trust in AI solutions...
$30 - $50 per hour
A tech company is seeking an AI Researcher to support end-to-end research for modern AI systems. This remote role involves designing experiments, defining evaluation protocols, and improving evaluation rigor for large language models. Key responsibilities include developing...Remote jobHourly pay$204.44k - $324.99k
...test Prompt Builder templates and grounded AI experiences using Salesforce data, Data... ...) patterns Implement agent testing, evaluation, observability, and guardrails,... ...Developer I and Salesforce Certified Agentforce Specialist are required. Platform Developer II, Health...Full timeH1bLocal area$229.9k - $262.4k
...Overview Senior Manager, AI Engineer (Gen AI Platform Services: Agentic AI, Guardrails, Evaluation) Overview : At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader...Full timePart timeLocal area$400 per month
Obsidian is seeking contributors for a Frontier Code Agents project, focused on evaluating AI coding models in fraud and risk engineering. Candidates will use AI coding tools to handle complex tasks and provide technical assessments. The role requires 2+ years of experience...$161.8k - $184.6k
...Principal Data Scientist - AI Foundations, Specialist Models Data is at the center of everything we do. As a startup, we disrupted the credit... ...all phases of development, from design through training, evaluation, validation, and implementation Leverage Agentic AI tools...Full timePart timeLocal areaImmediate startFlexible hours- ...time role, requiring strong analytical skills and independent work. The position involves creating historically relevant prompts, evaluating AI outputs, and contributing to research initiatives. Ideal candidates will hold a PhD in History or a related discipline and...Part timeRemote workFlexible hours
- ...Remote Job Summary We are seeking experienced Senior Software Engineers to support an AI training project by creating reinforcement learning environments that evaluate AI models on complex software engineering tasks using Model Context Protocol (MCP) tools....Remote jobFor contractors
$18 per hour
A tech-focused annotation company is seeking annotators to review and label data for AI development projects. This role involves carefully evaluating various types of content based on specific guidelines. Applicants should have a bachelor's degree and advanced English...Remote jobPart timeFreelanceFlexible hours- ...to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core computer science domains, evaluate solution quality, and help establish gold-standard benchmarks used to...Remote work
$50 per hour
A leading AI training company is seeking experienced software engineers to help train generative AI models remotely. The role involves crafting questions related to computer science and evaluating AI-generated code. This freelance opportunity offers flexible hours and...Remote jobHourly payFreelanceFlexible hours- An educational organization is seeking a skilled professional for a remote contract role focused on creating and evaluating AI-generated academic content. This position requires a Master's degree in a relevant field, solid academic writing skills, and excellent analytical...Remote jobContract work
- Feedinkoo is seeking a Copywriting & Content Subject Matter Expert to provide expert evaluation of AI-generated content and create high-quality written material. This remote contractor position requires strong copywriting skills and the ability to give precise feedback...Remote jobContract workFor contractors
- Meridial is seeking a Danish Voice Acting Specialist to support AI training by providing recorded speech samples and expert feedback. You will... ...recording setup. You will work collaboratively with a team to evaluate AI output and improve voice design. The position offers a...Remote jobHourly pay
- Supporting ongoing AI and legal technology initiatives, the part-time AI Integration Specialist will work remotely on a project basis, focusing on chatbot development, integration... ...to identify workflow improvements and evaluate AI-driven legal processes Required...Contract workPart timeRemote workFlexible hours
$18 per hour
...intelligence. What We Do Toloka connects individuals with Generative AI projects from leading tech innovators. Our mission is to unlock... ...part in online projects such as rating AI-generated content, evaluating factual accuracy, or comparing responses — when projects are...Part timeFreelanceRemote work$190k - $200k
...Job Description Job Description Job Title: AI Engineer Location: New York, NY Department : Technology/Revenue Strategy... ...property and commercial teams What You'll Do Design, train, evaluate, and deploy machine learning models on real booking, rate, and...Full timeWork at officeRemote workShift workNight shift- ...Overview We are looking for Generalists with strong analytical and critical-thinking skills to support AI and data-focused projects. The role involves reviewing, evaluating, and annotating information across different topics while maintaining high accuracy and quality....For contractorsRemote work
- ...Contributor Network Welo Data, part of Welocalize, is a global AI data company with 500,000+ contributors delivering high-quality,... ...be first in line for flexible, remote projects in annotation, evaluation, and prompt creationalways on your terms. Note: This is not an...Remote workFlexible hours
$18 - $50 per hour
...Overview: Siemens Digital Industries Software is seeking an AI & Engineering Data Intern to support emerging Artificial Intelligence... ...engineering information into reusable data models. Evaluate and document AI-assisted engineering workflows and use cases....Remote jobHourly payFull timeInternshipLocal area- Prolific is seeking AI Trainers fluent in advanced Dutch to train and evaluate AI models. Candidates will join a participant pool and be compensated for AI tasks which take around one hour to complete. Responsibilities include analyzing, editing, and writing in Dutch,...
- Meridial is seeking a Politics & Government Specialist to contribute to a Freelance AI Trainer Project. This remote position offers a dynamic role where... ...civic contexts. Compensation ranges from $6 to $65 per hour, evaluated based on expertise. #J-18808-Ljbffr MeridialRemote jobHourly payFreelance
- A leading AI solutions provider is seeking AI Data Specialists to improve AI-generated content in English. This freelance, part-time position allows for a... ...per week. Responsibilities include data collection, evaluation, annotation, and labeling across various content types...Remote jobExtra incomePart timeFreelance10 hours per weekFlexible hours
- ...Perseus. This role requires professional-level Turkish proficiency and a background in writing, along with AI safety expertise. Responsibilities include evaluating AI's handling of Arabic nuances and providing feedback to improve model accuracy. Join a collaborative...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Evaluation Specialist. Be the first to apply!





