Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Workflow Evaluator & Personal Task Trainer

Obsidian

Obsidian is seeking advanced LLM power users to evaluate AI systems on real-world personal-life workflows. You will use MCP and plugins/connectors like Google Drive, Notion, and Expedia to perform multi-step tasks such as travel planning, health research, home services, and career planning. The role requires documenting decisions, recording your screen, and articulating why AI outputs are strong, weak, or unrealistic. #J-18808-Ljbffr Obsidian

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the AI Workflow Evaluator & Personal Task Trainer in New York, NY vacancy
  • Obsidian is seeking experienced evaluators to test AI systems on complex personal workflows across health, travel, planning, home services, and career search. You will create realistic prompts, execute tasks with screen recording, and use your own plugins to complete actions... 
    Suggested

    Obsidian

    New York, NY
    3 days ago
  • About the Opportunity A leading AI research organization is seeking advanced...  .../connectors for real-world personal life tasks. This project focuses on evaluating how well AI systems handle...  ...evaluate AI systems on complex personal workflows, including tasks across: Personal... 
    Suggested
    Trial period

    Obsidian

    New York, NY
    4 days ago
  • Productive Playhouse seeks a Portugal Portuguese AI Product Evaluator for an upcoming project testing and evaluating leading AI chatbots. You...  ...This is an independent contractor engagement with batches of tasks; you set your hours and may work with other clients.... 
    Suggested
    Remote job
    For contractors
    Freelance
    Flexible hours

    United States Digital Space LLC

    New York, NY
    4 days ago
  • YO AI Labs is seeking an experienced Adobe Marketing Technology Expert for a remote contractor role. The candidate will test marketing workflows and evaluate AI-generated outcomes across Adobe Workfront, AEM, CJA, Analytics, and Experience Cloud. You will identify workflow... 
    Suggested
    Remote job
    For contractors

    YO AI Labs

    New York, NY
    3 days ago
  • $150 per hour

    Rise Data Labs is seeking experienced finance professionals for a short-term AI training and evaluation project. You’ll apply real-world finance expertise to develop complex finance tasks and evaluate AI-generated work for accuracy and quality. Rate: $150/hour. Start date... 
    Suggested
    Temporary work
    Immediate start

    BAM Ventures

    New York, NY
    5 days ago
  • Mindrift is offering a project-based opportunity to craft and evaluate AI coding tasks. You will build a realistic developer environment—a virtual company with a codebase, infrastructure, tickets, docs, and conversations—to test AI agents. You design tasks from intermediate... 

    Dorado

    New York, NY
    2 days ago
  • $400 per month

     ...is partnering with a leading AI research lab to support a Frontier...  ...project. Contributors help evaluate and improve frontier AI coding...  ...on realistic data engineering workflows and model evaluation. Spots...  ...evaluate complex data engineering tasks. Review model-generated... 

    Mercor

    New York, NY
    5 days ago
  • $30 per hour

     ...: 10-40 hours/week Role Responsibilities Evaluate outputs from large language models and autonomous...  ...standards. Review multi-step agent workflows, including screenshots and reasoning...  ...Requirements Strong experience in LLM evaluation, AI output analysis, QA/testing, UX research,... 
    Remote job
    Hourly pay
    Contract work

    Crossing Hurdles

    New York, NY
    3 days ago
  • $20 per hour

     ...professionals based in the United States to support AI model improvement through high-quality data...  ...strong accuracy across repetitive review tasks. In this role, you will work on structured AI data tasks used to train and evaluate the Gemini model. CONTRACT: Short-term... 
    Hourly pay
    Contract work
    Temporary work
    For contractors
    Freelance
    Remote work

    Gramian Consulting Group

    New York, NY
    6 days ago
  • $20 - $30 per hour

    A remote-focused technology firm is looking for an individual proficient in evaluating large language model outputs. The role involves assessing AI systems, reviewing workflows, and providing actionable feedback to enhance product quality. Ideal candidates will have strong... 
    Remote job
    Hourly pay

    Crossing Hurdles

    New York, NY
    3 days ago
  •  ...professionals to join a project that defines what excellent AI-assisted marketing work looks like. You will design task-specific grading criteria and rigorously score...  ...strategy, and familiarity with AI training or evaluation. This is a highly collaborative, fast-moving... 

    Obsidian

    New York, NY
    3 days ago
  •  ...(English & Vietnamese) to help make advanced AI models safer. You’ll apply language fluency and cultural judgment to evaluate how models handle sensitive topics in Vietnamese...  ...is required—we’ll train you on the workflow. The role involves writing expert-level prompts... 

    HumanitApp

    New York, NY
    4 days ago
  • Productive Playhouse is building an Albanian-speaking evaluator pool for AI chatbot testing. As an AI Evaluator, you’ll interact with AI models...  ..., safety, and usefulness, providing data to improve models. Tasks come in batches with flexible hours and project-based... 
    Remote job
    Freelance
    Flexible hours

    Productive Playhouse

    New York, NY
    3 days ago
  • Productive Playhouse is building a talent pool of Armenian speakers to test and evaluate leading AI chatbots. This freelance, project-based opportunity lets you choose tasks, set your own hours, and work with other clients as needed. Open to freelancers outside the U.S.... 
    Remote job
    Freelance
    Flexible hours

    Productive Playhouse

    New York, NY
    3 days ago
  • $24 per hour

    Prolific is seeking an AI Trainer with advanced Tamil fluency to evaluate AI models' understanding of the Tamil language's emotional and cultural nuances. Responsibilities...  ...compensation is offered at up to $24/hr for tasks, with flexible hours and remote work options. #J-1880... 
    Remote job
    Flexible hours

    Prolific

    New York, NY
    5 days ago
  • $100 - $125 per hour

    Crossing Hurdles is looking for freelance Insurance Experts to translate insurance workflows into AI tasks. Responsibilities include evaluating AI outputs, documenting processes, and working on insurance operations. Candidates should have at least 5 years of experience... 
    Remote job
    Hourly pay
    Freelance
    10 hours per week
    Flexible hours

    Crossing Hurdles

    New York, NY
    3 days ago
  • Mercor is hiring musicians to evaluate generative music AI models in partnership with a leading AI lab. You will assess AI-generated lyrics across...  ...week; compensation is hourly, with a potential shift to per-task structure while keeping the effective rate. #J-18808-Ljbffr... 
    Hourly pay
    Immediate start
    Shift work

    Mercor

    New York, NY
    2 days ago
  • Mercor is seeking experienced Medical and Health Services Managers to evaluate and improve AI-generated healthcare operations content and workflows. You will leverage your expertise in directing clinical services, personnel, budgets, and compliance across healthcare facilities... 

    Mercor

    New York, NY
    5 days ago
  • $100 - $125 per hour

     ...Immediate | 24-hour fast-track onboarding required Key Responsibilities Translate real-world insurance workflows into structured tasks for AI systems Evaluate AI-generated outputs for accuracy, logical reasoning, and business relevance Work on use cases including... 
    Remote job
    Hourly pay
    For contractors
    Freelance
    Work at office
    Immediate start
    Flexible hours

    Crossing Hurdles

    New York, NY
    2 days ago
  • $30 - $50 per hour

     ...A leading AI training company is seeking a Remote Annotator to support human-in-the-loop AI training workflows for large language models. This role involves reviewing labeled datasets, performing evaluations for helpfulness and safety, and ensuring quality in training... 
    Hourly pay
    Remote work

    Rex USA

    New York, NY
    3 days ago
  • $75 per hour

     ...Medical Coder AI Content Evaluator (Remote) This role focuses on applying medical coding and healthcare operations expertise to assess and...  ...Test AI performance against custom-built scenarios and refine tasks based on quality reviews, performance results, and... 
    Remote job
    Contract work
    Temporary work
    Work at office
    Flexible hours

    Aston Carter

    New York, NY
    3 days ago
  • Mercor partners with a leading AI research lab to support a Frontier Code Agents...  ...on realistic data engineering workflows and model evaluation. Contributors help evaluate and improve...  ...distributed data systems. Compensation is task-based and distributed upon acceptance.... 

    Mercor

    New York, NY
    5 days ago
  •  ...are seeking detail-oriented and motivated AI Data Annotators (Canadian French) to support...  ...systems. In this role, you will review, evaluate, and validate AI-generated content to help...  ...enhance model performance. Complete assigned tasks accurately and within established... 
    Full time
    Contract work
    Immediate start
    Remote work
    Flexible hours

    Centraprise

    New York, NY
    5 days ago
  • $30 per hour

     ...not just another player in the AI space - we are building the...  ...Domain Experts for a high-level AI evaluation project. AI models are...  ...and now tackle complex design tasks such as layout, typography, and...  ...that Prolific may collect your personal data for recruiting and global... 
    Remote job
    Work from home
    Flexible hours

    Prolific

    New York, NY
    4 days ago
  • Mercor is partnering with a leading AI research organization to engage experienced UI...  ...product designers for a project focused on evaluating how well AI systems perform real-world...  ...what excellent work looks like: designing task-specific grading criteria and scoring completed... 

    Obsidian

    New York, NY
    4 days ago
  • $50 - $60 per hour

    A leading company in AI research is seeking candidates with a PhD or Master's degree in Chemistry to evaluate and improve large language models (LLMs). The position is part-time, around 30 hours a week, and operates in a remote and flexible environment. Responsibilities... 
    Remote job
    Hourly pay
    Part time
    Flexible hours

    CloudDevs

    New York, NY
    3 days ago
  • Mercor is seeking expert Evaluators in Clinical / biomedical / pharma to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. This is a remote, hourly engagement. You will apply deep subject-matter... 
    Remote job
    Hourly pay

    Mercor

    New York, NY
    3 days ago
  • Mercor is hiring expert Evaluators in Product management / roadmap / PRD to review AI-generated work products for accuracy, rigor, and domain quality. This is a remote, hourly engagement. You will apply deep subject-matter expertise to grade outputs, evaluate artifacts... 
    Remote job
    Hourly pay

    Mercor

    New York, NY
    1 day ago
  • A global data analytics firm is seeking an Evaluator - Political Science to assess AI-generated responses for quality and relevance. This remote position requires a PhD or Masters in Political Science along with at least two years of research experience. Successful candidates... 
    Remote job

    Straive

    New York, NY
    3 days ago
  • Obsidian is hiring expert Evaluators based in New York, United States for a remote role focusing on reviewing AI-generated work products like documents, spreadsheets, and slide decks. The ideal candidate will have over 5 years of experience in BI dashboards and performance... 
    Hourly pay
    Work at office
    Remote work

    Obsidian

    New York, NY
    6 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Workflow Evaluator & Personal Task Trainer. Be the first to apply!