AI Workflow Evaluator & Personal Task Trainer
Obsidian
Obsidian is seeking advanced LLM power users to evaluate AI systems on real-world personal-life workflows. You will use MCP and plugins/connectors like Google Drive, Notion, and Expedia to perform multi-step tasks such as travel planning, health research, home services, and career planning. The role requires documenting decisions, recording your screen, and articulating why AI outputs are strong, weak, or unrealistic. #J-18808-Ljbffr Obsidian
- Obsidian is seeking experienced evaluators to test AI systems on complex personal workflows across health, travel, planning, home services, and career search. You will create realistic prompts, execute tasks with screen recording, and use your own plugins to complete actions...Suggested
- About the Opportunity A leading AI research organization is seeking advanced... .../connectors for real-world personal life tasks. This project focuses on evaluating how well AI systems handle... ...evaluate AI systems on complex personal workflows, including tasks across: Personal...SuggestedTrial period
- Productive Playhouse seeks a Portugal Portuguese AI Product Evaluator for an upcoming project testing and evaluating leading AI chatbots. You... ...This is an independent contractor engagement with batches of tasks; you set your hours and may work with other clients....SuggestedRemote jobFor contractorsFreelanceFlexible hours
- YO AI Labs is seeking an experienced Adobe Marketing Technology Expert for a remote contractor role. The candidate will test marketing workflows and evaluate AI-generated outcomes across Adobe Workfront, AEM, CJA, Analytics, and Experience Cloud. You will identify workflow...SuggestedRemote jobFor contractors
$150 per hour
Rise Data Labs is seeking experienced finance professionals for a short-term AI training and evaluation project. You’ll apply real-world finance expertise to develop complex finance tasks and evaluate AI-generated work for accuracy and quality. Rate: $150/hour. Start date...SuggestedTemporary workImmediate start- Mindrift is offering a project-based opportunity to craft and evaluate AI coding tasks. You will build a realistic developer environment—a virtual company with a codebase, infrastructure, tickets, docs, and conversations—to test AI agents. You design tasks from intermediate...
$400 per month
...is partnering with a leading AI research lab to support a Frontier... ...project. Contributors help evaluate and improve frontier AI coding... ...on realistic data engineering workflows and model evaluation. Spots... ...evaluate complex data engineering tasks. Review model-generated...$30 per hour
...: 10-40 hours/week Role Responsibilities Evaluate outputs from large language models and autonomous... ...standards. Review multi-step agent workflows, including screenshots and reasoning... ...Requirements Strong experience in LLM evaluation, AI output analysis, QA/testing, UX research,...Remote jobHourly payContract work$20 per hour
...professionals based in the United States to support AI model improvement through high-quality data... ...strong accuracy across repetitive review tasks. In this role, you will work on structured AI data tasks used to train and evaluate the Gemini model. CONTRACT: Short-term...Hourly payContract workTemporary workFor contractorsFreelanceRemote work$20 - $30 per hour
A remote-focused technology firm is looking for an individual proficient in evaluating large language model outputs. The role involves assessing AI systems, reviewing workflows, and providing actionable feedback to enhance product quality. Ideal candidates will have strong...Remote jobHourly pay- ...professionals to join a project that defines what excellent AI-assisted marketing work looks like. You will design task-specific grading criteria and rigorously score... ...strategy, and familiarity with AI training or evaluation. This is a highly collaborative, fast-moving...
- ...(English & Vietnamese) to help make advanced AI models safer. You’ll apply language fluency and cultural judgment to evaluate how models handle sensitive topics in Vietnamese... ...is required—we’ll train you on the workflow. The role involves writing expert-level prompts...
- Productive Playhouse is building an Albanian-speaking evaluator pool for AI chatbot testing. As an AI Evaluator, you’ll interact with AI models... ..., safety, and usefulness, providing data to improve models. Tasks come in batches with flexible hours and project-based...Remote jobFreelanceFlexible hours
- Productive Playhouse is building a talent pool of Armenian speakers to test and evaluate leading AI chatbots. This freelance, project-based opportunity lets you choose tasks, set your own hours, and work with other clients as needed. Open to freelancers outside the U.S....Remote jobFreelanceFlexible hours
$24 per hour
Prolific is seeking an AI Trainer with advanced Tamil fluency to evaluate AI models' understanding of the Tamil language's emotional and cultural nuances. Responsibilities... ...compensation is offered at up to $24/hr for tasks, with flexible hours and remote work options. #J-1880...Remote jobFlexible hours$100 - $125 per hour
Crossing Hurdles is looking for freelance Insurance Experts to translate insurance workflows into AI tasks. Responsibilities include evaluating AI outputs, documenting processes, and working on insurance operations. Candidates should have at least 5 years of experience...Remote jobHourly payFreelance10 hours per weekFlexible hours- Mercor is hiring musicians to evaluate generative music AI models in partnership with a leading AI lab. You will assess AI-generated lyrics across... ...week; compensation is hourly, with a potential shift to per-task structure while keeping the effective rate. #J-18808-Ljbffr...Hourly payImmediate startShift work
- Mercor is seeking experienced Medical and Health Services Managers to evaluate and improve AI-generated healthcare operations content and workflows. You will leverage your expertise in directing clinical services, personnel, budgets, and compliance across healthcare facilities...
$100 - $125 per hour
...Immediate | 24-hour fast-track onboarding required Key Responsibilities Translate real-world insurance workflows into structured tasks for AI systems Evaluate AI-generated outputs for accuracy, logical reasoning, and business relevance Work on use cases including...Remote jobHourly payFor contractorsFreelanceWork at officeImmediate startFlexible hours$30 - $50 per hour
...A leading AI training company is seeking a Remote Annotator to support human-in-the-loop AI training workflows for large language models. This role involves reviewing labeled datasets, performing evaluations for helpfulness and safety, and ensuring quality in training...Hourly payRemote work$75 per hour
...Medical Coder AI Content Evaluator (Remote) This role focuses on applying medical coding and healthcare operations expertise to assess and... ...Test AI performance against custom-built scenarios and refine tasks based on quality reviews, performance results, and...Remote jobContract workTemporary workWork at officeFlexible hours- Mercor partners with a leading AI research lab to support a Frontier Code Agents... ...on realistic data engineering workflows and model evaluation. Contributors help evaluate and improve... ...distributed data systems. Compensation is task-based and distributed upon acceptance....
- ...are seeking detail-oriented and motivated AI Data Annotators (Canadian French) to support... ...systems. In this role, you will review, evaluate, and validate AI-generated content to help... ...enhance model performance. Complete assigned tasks accurately and within established...Full timeContract workImmediate startRemote workFlexible hours
$30 per hour
...not just another player in the AI space - we are building the... ...Domain Experts for a high-level AI evaluation project. AI models are... ...and now tackle complex design tasks such as layout, typography, and... ...that Prolific may collect your personal data for recruiting and global...Remote jobWork from homeFlexible hours- Mercor is partnering with a leading AI research organization to engage experienced UI... ...product designers for a project focused on evaluating how well AI systems perform real-world... ...what excellent work looks like: designing task-specific grading criteria and scoring completed...
$50 - $60 per hour
A leading company in AI research is seeking candidates with a PhD or Master's degree in Chemistry to evaluate and improve large language models (LLMs). The position is part-time, around 30 hours a week, and operates in a remote and flexible environment. Responsibilities...Remote jobHourly payPart timeFlexible hours- Mercor is seeking expert Evaluators in Clinical / biomedical / pharma to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. This is a remote, hourly engagement. You will apply deep subject-matter...Remote jobHourly pay
- Mercor is hiring expert Evaluators in Product management / roadmap / PRD to review AI-generated work products for accuracy, rigor, and domain quality. This is a remote, hourly engagement. You will apply deep subject-matter expertise to grade outputs, evaluate artifacts...Remote jobHourly pay
- A global data analytics firm is seeking an Evaluator - Political Science to assess AI-generated responses for quality and relevance. This remote position requires a PhD or Masters in Political Science along with at least two years of research experience. Successful candidates...Remote job
- Obsidian is hiring expert Evaluators based in New York, United States for a remote role focusing on reviewing AI-generated work products like documents, spreadsheets, and slide decks. The ideal candidate will have over 5 years of experience in BI dashboards and performance...Hourly payWork at officeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Workflow Evaluator & Personal Task Trainer. Be the first to apply!
- education evaluator New York, NY
- program evaluator New York, NY
- social media evaluator New York, NY
- work from home web search evaluator New York, NY
- ai evaluator New York, NY
- clinical evaluator New York, NY
- evaluator New York, NY
- work from home social media evaluator New York, NY
- quality evaluator New York, NY
- personal trainers New York, NY

