Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Evaluation Scenario Writer - AI Agent Testing Specialist

$80 per hour

Mindrift

Please submit your CV in English and indicate your level of English proficiency. Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment. What This Opportunity Involves Review and refine realistic coding tasks based on provided production codebases with realistic scope, requirements and information sources Write comprehensive functional tests that validate actual end-to-end behavior and edge‑cases, not just superficial checks Craft "fair but hard" challenges where the AI has all the context it needs, but has to work for it (information scattered across files and external sources, complex reasoning required) Analyze AI failures to understand what the model struggles with vs. what it masters Iterate based on feedback from expert QA reviewers who score your work on 7 quality criteria What We Look For Degree in Computer Science, Software Engineering or related fields 5+ years in software development, primarily Python (pytest, async/await, subprocess, file operations) Background in Full‑Stack development, with an equal focus on building React‑based interfaces and robust Back‑end systems Experience writing tests (functional, integration - not just running them) Docker containers (running evaluations locally in containers) CI/CD understanding (GitHub Actions as a user: triggers, labels, reading results) English proficiency - B2 How It Works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid Effort estimate Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted. Payment Paid contributions, with rates up to $80/hour* Fixed project rate or individual rates, depending on the project Some projects include incentive payments Note: Rates vary based on expertise, skills assessment, location, project needs, and other factors. Higher rates may be offered to highly specialized experts. Lower rates may apply during onboarding or non‑core project phases. Payment details are shared per project #J-18808-Ljbffr Mindrift

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Evaluation Scenario Writer - AI Agent Testing Specialist in New York, NY vacancy
  • $80 per hour

     ...intelligence to ethically shape the future of AI. What We Do The Mindrift platform connects specialists with AI projects from major tech...  ...can design realistic and structured evaluation scenarios for LLM-based agents. You'll create test cases that simulate human-performed... 
    Suggested
    Part time
    Freelance
    Remote work
    Flexible hours

    Mindrift

    New York, NY
    1 day ago
  • $55 per hour

     ...domain experts with cutting-edge AI projects. We are seeking an Evaluation Scenario Writer - QA for a project focused on ensuring...  ...scenarios created for LLM agents. This is a flexible, project-based...  ...manual scenario validation, automated test thinking, and collaboration with... 
    Suggested
    Part time
    Freelance
    Internship
    Remote work
    Flexible hours

    Mindrift

    New York, NY
    2 days ago
  • Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation...  ...is solvable by an AI agent Write tests that verify...  ...models fail and what scenarios reveal the difference... 
    Suggested
    Hourly pay
    Permanent employment
    Temporary work
    Freelance

    Socket.dev

    New York, NY
    3 days ago
  • Mercor is hiring experienced musicians to evaluate generative music AI models in partnership with a leading AI lab. You will assess AI-generated music across genres and rate it against detailed quality standards, working in Vietnamese and English. You will compare lyrics... 
    Suggested

    Mercor

    New York, NY
    4 days ago
  •  ...Technical Communicator for our Apps and Agents team. We are developing a portfolio of AI-powered applications and...  ...across multiple products Writing testing instructions, FAQs, setup guides,...  ...designers, testers, and communication specialists Exposure to practical frameworks... 
    Suggested

    Crossing Party Lines

    New York, NY
    2 days ago
  •  ...Software Engineer, Agent Evaluation and Quality Engineering · Full-time · San Francisco; New York Our mission is to automate coding. The...  ...You'll Work On Designing and building best-in-class AI evaluation system: curated datasets, offline replay, scorers /... 
    Full time
    Work at office

    Anysphere

    New York, NY
    21 hours ago
  • Datadog is seeking a Staff Applied Scientist in New York to own evaluation strategies for AI agent integrations. The role involves defining metrics, building datasets, and collaborating with AI engineers to enhance tool selection and retrieval relevance. The ideal candidate... 

    Datadog

    New York, NY
    1 day ago
  • Mercor offers a fully remote contractor role for drafting Danish law MCQs with worked solutions. You will create exam-style questions from your professional expertise, with flexible scheduling and weekly payments via Stripe or Wise. The assignment may extend or end early...
    Remote job
    Weekly pay
    For contractors
    Freelance
    Flexible hours

    Moonlight

    New York, NY
    2 days ago
  • $205k - $257k

    Scale has been the leading AI data foundry, helping fuel the most...  ...Finance vertical within our Agents Data & Reinforcement Learning...  ...workflows that labs use to train and evaluate agents) and the “data as a...  ...agents against real world scenarios. Partner with ML and Operations... 
    Full time

    Scale AI

    New York, NY
    1 day ago
  • CNTXT AI is seeking a remote contractor to evaluate AI-generated financial content and develop test cases that probe analytical reasoning. You will help improve how AI models handle financial information with clear explanations and rigorous checks. Responsibilities include... 
    Remote job
    For contractors

    CNTXT AI

    New York, NY
    21 hours ago
  • $180k - $225k

    About Scale AIScale AI is the data foundation for AI, helping organizations build and...  ...telecommunications to build production AI agents that automate complex workflows, help...  ...directly with enterprise customers to design, evaluate, and deploy intelligent systems that... 
    Full time

    Scale AI

    New York, NY
    3 days ago
  • $180k - $210k

    About TrabaTraba is building the AI operating layer for the industrial supply chain.We...  ...to join as a founding member of the Agents team and help us build the next layer of...  ...meaningful workDefine product behavior, evaluation systems, and quality frameworks that improve... 
    Temporary work
    Local area
    Flexible hours
    Shift work

    Traba

    New York, NY
    1 day ago
  • Finance & Accounting Task Author (AI Training) About the Role...  ...expert tasks that train and evaluate advanced AI agents on accounting reconciliation...  ...rigorous accounting scenarios — billing versus revenue reconciliation...  ...mindset: willing to test and refine tasks until difficulty... 
    Hourly pay
    Ongoing contract
    Contract work
    Freelance
    Remote work
    Flexible hours

    Dorado

    New York, NY
    3 days ago
  • $216k - $270k

     ...Scale’s mission is to develop reliable AI systems for the world’s most important decisions...  ...when it matters most, combining rigorous evaluation with full-stack deployment so our...  ...Role As a Machine Learning Engineer on Agent Oversight, you will drive the end-to-end... 
    Full time

    Scale Ai, Inc.

    New York, NY
    21 hours ago
  • $218k - $273k

     ...Scale AI is the data foundation for AI, helping organizations build and deploy reliable...  ...AI capabilities. About the General Agents Team The General Agents team, part of...  ...lifecycle—from model and system design to evaluation, deployment, and iteration—bridging... 
    Full time

    Scale Ai

    New York, NY
    21 hours ago
  •  ...human. You will define how AI agents navigate, operate, and succeed...  ..., trust specifications, and evaluation criteria, deliverables that...  ...regression sets, and adversarial scenarios that gate every capability...  ..., ideation, prototyping, testing, iteration, and delivery in... 
    Contract work
    Remote work
    Worldwide

    United States Digital Space LLC

    New York, NY
    4 days ago
  •  ...instincts could directly improve the AI systems millions of people use...  ...with cutting-edge AI systems, evaluate their responses, and provide...  ...a wide variety of topics and scenarios Evaluate responses for...  ..., technical, or professional testing experience required Nice... 
    Hourly pay
    Ongoing contract
    Contract work
    Freelance
    Remote work
    Flexible hours

    Alignerr

    New York, NY
    21 hours ago
  • $161.3k - $241.9k

     ...end-to-end. By combining frontier agentic AI, an enterprise-grade platform, and deep...  ...Role Overview As a Software Engineer, Agents, you'll build the systems that make our AI...  ..., and are experienced in using practical evaluations to drive task completion quality and customer... 
    Full time

    Harvey

    New York, NY
    21 hours ago
  •  ...human customer experiences with AI. We are primarily an in-...  ...fastest-growing industries—our agents already reach over 50% of U.S...  ...help members find in-network specialists, check availability, and locate...  ...description. We strive to evaluate all applicants consistently without... 
    Full time
    Flexible hours

    Sierra

    New York, NY
    21 hours ago
  • A mobile app insights company is looking for a Freelance Editorial Writer to test and review mobile apps from the comfort of your home. In this remote role, you will install apps, evaluate their performance, and write brief summaries of your experience. This position offers... 
    Fixed term contract
    Freelance
    Remote work

    Review Pays

    New York, NY
    1 day ago
  • $250k - $330k

     ...Decagon Decagon is the leading conversational AI platform empowering every brand to...  ...Health, and Hunter Douglas to deploy AI agents that power personalized, deeply satisfying...  ...research efforts Experiment with and run evaluations on the latest text and voice models, then... 
    Full time
    Work at office

    Decagon

    New York, NY
    21 hours ago
  •  ...The hardest problems in both AI and biology are being solved here...  .... You will build the agent infrastructure that turns model...  ...from experiment management to evaluation pipelines You will establish...  ...engineering practices across the team: testing, CI/CD, code review, and... 
    Full time

    Output Biosciences

    New York, NY
    21 hours ago
  • $152k - $240k

     ..., and control spend effortlessly. Brex’s AI-native automation and world-class service...  .... What you’ll do We're building AI agents to automate and augment internal functions...  ..., APIs, and data sources. Define evaluation frameworks, success metrics, and feedback... 
    Full time
    Work at office
    Remote work
    Work from home

    Brex Inc.

    New York, NY
    21 hours ago
  • $200k - $240k

    Traba is the AI operating layer for the industrial supply chain.We started in workforce...  ...and General Catalyst.You'll build the AI agents themselves—the harnesses, evals,...  ...hypothesis and a deadline, you scope, build, evaluate, and ship without waiting for a spec.Sweat... 
    Temporary work
    Local area
    Shift work
    Day shift

    Traba

    New York, NY
    21 hours ago
  • $1,500 - $2,000 per month

    Builder Lead Converter AI Agent Specialist Builder Lead Converter is hiring an AI Agent Specialist...  ...AI agent logic and workflows Deploy, test, document, and improve intelligent systems...  ...prompt frameworks, guardrails, and evaluation logic Optimize AI agents for: Accuracy... 
    For contractors
    Work at office
    Remote work
    Monday to Friday

    Builder Lead Converter

    New York, NY
    2 days ago
  • $200k - $300k

    Real-Time Voice AI Agent Systems Engineer Company: HireNow Staffing...  ...prompting, speech performance, testing, reliability, and continuous...  ...Create scalable testing and evaluation frameworks incorporating A/B...  ...conversational coverage as new customer scenarios and operational requirements... 
    Permanent employment
    Full time
    Relocation
    Visa sponsorship

    HireNow Staffing

    New York, NY
    1 day ago
  • $115k - $190k

     ...new technology platform that continually tests fraud rules for implementation issues. They...  ...sources and building analytical models to evaluate the trade-offs between fraud rule/control...  ...and automations using tools like Generative AI, Python, SQL, and DataikuProject... 
    Temporary work
    Work at office

    Morgan Stanley

    New York, NY
    3 days ago
  •  ...a remote Copywriting & Content Subject Matter Expert to review AI-generated marketing content and create expert copy. This part-time...  ...will develop AI training content and optimize AI performance by evaluating responses. If qualified, you will be first in line for relevant... 
    Remote job
    Contract work
    Part time

    YO IT Consulting

    New York, NY
    21 hours ago
  • $190k - $270k

    Staff Software Engineer - Agent QualityP-1215At Databricks, we are...  ...running the world’s best data and AI platform so our customers can...  ...of a new team focused on evaluating and continuously improving Databricks...  ...devtools, CI/CD platforms, testing frameworks, observability... 
    Local area
    Worldwide

    DataBricks

    New York, NY
    1 day ago
  • $160,000 - $210,000 per week

     ...Recruiter Own the talent engine from day one. Confidential Enterprise AI Agent Startup · New York, NY New York, NY · On-site 5 days/week $16...  ...initial technical screens to assess candidates’ relevance — evaluating language experience, AI background, and fit before passing to... 
    H1b
    Live in
    Work at office
    Immediate start

    Aionia

    New York, NY
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Evaluation Scenario Writer - AI Agent Testing Specialist. Be the first to apply!