Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Evaluation Scenario Writer - AI Agent Testing Specialist

$80 per hour
Temporary

Mindrift

This opportunity is only for candidates currently residing in the specified country. Your location may affect eligibility and rates. Please submit your resume in English and indicate your level of English.

At Mindrift, innovation meets opportunity. We believe in using the power of collective human intelligence to ethically shape the future of AI.

What We Do

The Mindrift platform connects specialists with AI projects from major tech innovators. Our mission is to unlock the potential of Generative AI by tapping into real‑world expertise from across the globe.

About the Role

We’re looking for someone who can design realistic and structured evaluation scenarios for LLM‑based agents. You’ll create test cases that simulate human‑performed tasks and define gold‑standard behavior to compare agent actions against. You’ll work to ensure each scenario is clearly defined, well‑scored, and easy to execute and reuse. You’ll need a sharp analytical mindset, attention to detail, and an interest in how आरोच्छ in how AI agents make decisions.

Responsibilities
    \ wallet Create structured test cases that simulate complex human workflows
  • Define gold‑standard behavior and scoring logic to evaluate agent actions.
  • Analyze agent logs, failure modes, and decision paths.
  • Work with code repositories and test frameworks to validate your scenarios.
  • Iterate on prompts, instructions, and test cases to improve clarity and difficulty.
  • Ensure that scenarios are production‑ready, easy to run, and reusable.

How To Get Started

Simply apply to this post, qualify, виды and get the chance to contribute to projects aligned with your skills, on your own schedule. From creating training prompts to refining model responses, you’ll help shape the future of AI while ensuring technology benefits everyone.

Requirements

  • Bachelor’s and/or Master’s Degree in Computer Science, Software Engineering, Data Science / Data Analytics, Artificial Intelligence / Machine Learning, Computational Linguistics / Natural Language Processing (NLP), Information Systems, or 天天中彩票为什么 related fields.
  • Background in QA, software testing, data analysis, or NLP annotation.
  • Good understanding of test design principles (e.g., reproducibility, coverage, edge cases).
  • Strong written communication skills in English.
  • Comfortable with structured formats like JSON/YאםLAML for scenario description.
  • Can define expected agent behaviors (gold paths) and scoring logic.
  • Basic experience with Python and JS.
  • Curious and open to working with AI‑generated content, agent logs, and prompt‑based behavior.

Nice to Have

  • Experience in writing manual or automated test cases.
  • Familiarity with LLM capabilities and typical failure modes.
  • Understanding of scoring metrics (precision, recall, coverage, reward functions).

EEP Benefits

  • Get paid for your expertise, with rates that can go up to $80/hour depending on your skills, experience, and project needs.
  • Take part in a flexible, remote, freelance project that fits around your primary professional or academic commitments.
  • Participate in an advanced AI project and gain valuable experience to enhance your portfolio.
  • Influence how future AI models understand and communicate in your field of expertise.

Seniority Level

Entry level

Employment Type минуты

Part‑time

Job Function

Other

Industries

IT Services and IT Consulting

#J-18808-Ljbffr
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Evaluation Scenario Writer - AI Agent Testing Specialist in Texas vacancy
  • $80 per hour

    A leading AI consultancy is seeking a detail-oriented individual to design evaluation scenarios for LLM-based agents. This part-time, remote role allows you to create structured test cases while working flexibly around your commitments. The ideal candidate holds a relevant... 
    Suggested
    Part time
    Remote work

    Mindrift

    Houston, TX
    2 days ago
  • $80 per hour

     ...ethically shape the future of AI. What We Do The...  ...modern AI systems are tested and evaluated? This is a flexible,...  ...QAs for autonomous AI agents for a new project...  ...and thinking through scenarios, implications, and edge...  ...Working closely with QA, writers, or developers to suggest... 
    Suggested
    Permanent employment
    Part time
    Freelance
    Remote work
    Flexible hours

    Mindrift

    Houston, TX
    2 days ago
  • $55 per hour

    A leading AI innovation firm in Dallas is seeking QAs for autonomous AI agents to ensure the quality of complex systems and scenarios. This flexible, remote project is ideal for those with excellent...  .... Candidates should be adept at evaluating scenarios and documenting... 
    Suggested
    Remote job
    Flexible hours

    Mindrift

    Dallas, TX
    1 day ago
  • $80 per hour

    Get AI‑powered advice on this job and more exclusive features...  ...tools for running and evaluating agent behavior. You’ll implement...  ...check agent actions against scenario definitions Creating or extending tools that writers and QAs use to test agents Working closely with... 
    Suggested
    Freelance
    Remote work
    Flexible hours

    Mind Rift

    Dallas, TX
    2 days ago
  • $60 per hour

    A leading AI firm in Austin is looking for QA experts to validate and improve AI systems. This...  ...to detail. Candidates will review AI evaluation tasks, identify inconsistencies, and define expected behaviors for agents. Ideal applicants have experience in policy evaluation... 
    Suggested
    Remote job
    Freelance

    Mind Rift

    Austin, TX
    2 days ago
  • $45.5 - $65 per hour

     ...Cypress, Python, PyTest, JMeter, AI Testing - Mandatory.1-2 days onsite at client...  ...performance.Build automated test scenarios for AI-driven workflows, including...  ..., chatbots, and intelligent agents.Develop test datasets and evaluation frameworks for AI model validation... 

    Yoh

    Dallas, TX
    3 days ago
  • $30 per hour

     ...broadest and deepest suite of AI-powered cloud...  ...a highly motivated AI Agent Intern to join Oracle'...  ...workflows for logistics scenarios (e.g., shipment risk detection...  ...datasets for agent testing. Solution...  ...Research & Analysis Evaluate logistics AI use cases... 
    Hourly pay
    Temporary work
    Internship
    Flexible hours

    Oracle

    Austin, TX
    2 days ago
  •  ...Eleven is looking for a Staff Engineer - Agents to join the Enterprise AI Team!At 7-Eleven Enterprise AI Team,...  ...development, model training, testing, and deploymentConduct deep dives into...  ...continued performance and reliability.Evaluate and grow technical talent throughout... 
    Hourly pay

    7 eleven

    Irving, TX
    10 hours ago
  •  ...Opportunity for advancement Senior Product Manager, AI Agents Work Type: Full-Time, Onsite Location: Dallas, Texas...  ...technology partners and vendors in the AI ecosystem. Lead A/B testing, agent evaluation, and post-deployment analysis to drive continuous... 
    Full time

    Select Minds LLC

    Dallas, TX
    more than 2 months ago
  • $110k - $160k

    Hellopatient is seeking a technical AI Agent Product Manager in Austin, Texas, to lead the delivery of AI agents tailored for healthcare...  ...and continuously enhance agent performance through structured evaluation and real-world feedback. The position offers a competitive... 

    Hellopatient

    Austin, TX
    3 days ago
  • $80 - $100 per hour

     ...Role Overview Help train next-generation AI coding agents by contributing expert, real-world examples, technical walkthroughs, and structured assessments. In this contractor role you will document how AI coding tools support complex software development tasks, explain... 
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Real County, TX
    4 days ago
  • $66.3k - $84k

     ...Technical Writer At Jacobs, we're challenging today to...  ...will also involve leveraging AI software such as Co-Pilot Agent and Power BI, working...  ...including inspection and test plans, reports, templates,...  ...Apply critical thinking to evaluate technical content, identify... 
    For subcontractor

    Jacobs Solutions

    Houston, TX
    10 hours ago
  • $60 - $64 per hour

     ...SolutionsRole: Senior UI Engineer with AI agentsLocation: Austin, TXWe...  ...ways of building UI with AI agents. The Frameworks team builds...  ...-of-concept applications to evaluate their features and limitations...  ...to deliver high-quality, well-tested code.Ideal Candidate Profile:Deep... 
    Hourly pay
    Full time

    SRI Tech

    Austin, TX
    2 days ago
  •  ...Consulting is looking for a part-time remote Copywriting & Content Subject Matter Expert (SME) to review and create AI-generated content. You will evaluate the quality of AI outputs and provide expertise to enhance persuasive and accurate writing. The ideal candidate has... 
    Remote job
    Part time
    Immediate start

    YO IT Consulting

    Houston, TX
    1 day ago
  •  ...in Dallas, TX is seeking an experienced AI Architect to act as a trusted advisor to...  ...technology leaders while leading the design, evaluation, and implementation of enterprise-scale...  ...business opportunities, design custom AI agents, evaluate emerging AI products, and establish... 

    TechDigital Group

    Dallas, TX
    10 hours ago
  • $110k - $160k

    About the Role Hello Patient is hiring a technical, high-agency AI Agent Product Manager to own the end-to-end delivery of AI agents...  ...continuously iterate on agent performance using structured evaluation, testing, and real customer feedback. What You'll Do Own customer... 

    Hellopatient

    Austin, TX
    10 hours ago
  • $151.28k - $190k

     ...next generation of enterprise AI governance capabilities to enable...  ...responsible adoption of AI agents. The team is evolving governance...  ...capabilities, continuously evaluating and adopting new frameworks, features...  ...design patterns, automated testing, and system integration.Strong... 
    Local area
    Worldwide
    Flexible hours

    Ping Identity

    Austin, TX
    1 day ago
  •  ...Power is standing up a dedicated AI build team to make Quote-to-...  ...a versioned prompt and evaluation library to support consistent,...  ...Context Protocol) or similar agent-integration patterns. Experience...  ...harnesses or systematic prompt testing. Exposure to CRM (HubSpot) or... 

    Maverick Power

    Plano, TX
    4 days ago
  •  ...stakeholders and iterate the wireframe screens. To test with sample users to ensure that the...  ...create a persona (user description) and scenario (interaction with product). To utilize the current product (heuristic evaluation) or findings obtained, competitor product analysis... 

    Omni Inclusive

    San Antonio, TX
    3 days ago
  •  ...build proof-of-concept applications to evaluate their features and limitations. · Implement...  .... · Architect and develop multi-step AI agents that integrate with UI component...  ...independently to deliver high-quality, well-tested code. Ideal Candidate Profile: ·... 

    Cloud Analytics Technologies LLC

    Austin, TX
    4 days ago
  • $89.2k - $209.5k

     ...Description Oracle Health is seeking a Senior AI Agent Engineer to build production AI agents...  ..., balancing speed with safety, evaluation, and maintainability. This person should...  .... Responsibilities Design, build, test, and deploy production AI agents and AI-... 
    Temporary work
    Flexible hours

    Oracle

    Austin, TX
    1 day ago
  • $93.75k - $133.2k

     ...civilian hiring freeze. As an FBI special agent, your career is defined by how you think...  ..., and operate across a range of complex scenarios. Every day brings new challenges that...  ...accounting and finance reflects how you evaluate information, navigate complexity, and operate... 
    Work at office
    Local area
    Shift work

    Federal Bureau of Investigation (FBI)

    Dallas, TX
    2 days ago
  •  ...Job Description The Role: We are seeking an AI Agent Engineer to design, build, and operationalize AI-powered agents...  ...improve agent performance. Establish and maintain evaluation frameworks, metrics, and test harnesses for AI agent behavior and output quality.... 
    Local area
    Work from home
    Relocation
    Relocation package

    General Motors

    Austin, TX
    1 day ago
  • $262k - $364k

     ...strategy and architecture for AI-ready data foundations, including...  ...search and summarization agents, alongside modular frameworks...  ...time queries.Design automated evaluation frameworks to improve agent accuracy...  ...design and architecture; and testing/launching software products.... 

    Google

    Sunnyvale, TX
    4 days ago
  • $207k - $301k

     ...scalability and efficiency of Orcas Agent Runtime Platform.Work cross-...  ....Architect and develop novel evaluation and customization solutions,...  ...assessment ( Cloud Lifecycle AI Tools).Mentor and provide technical...  ....5 years of experience testing, and launching software products... 

    Google

    Sunnyvale, TX
    4 days ago
  •  ...has developed an artificial intelligence (AI) powered technology stack purpose-built...  ...commercial self-driving software to develop, test and deploy autonomous capabilities for...  ...screening of applications. As part of the evaluation process, we provide Endorsed with job... 
    Visa sponsorship
    Shift work
    Night shift
    Weekend work

    Kodiak Robotics

    Dallas, TX
    10 hours ago
  • Who is Sonar? Sonar is driving the future of agent-centric software development. As the leader in AI code verification and governance, we solve a critical problem: ensuring that software generated by AI-assisted developers or autonomous agents is reliable, secure, and maintainable... 
    Full time

    Sonar Source

    Austin, TX
    11 days ago
  •  ...Job Description Job Description Who is Sonar? Sonar is driving the future of agent-centric software development. As the leader in AI code verification and governance, we solve a critical problem: ensuring that software generated by AI-assisted developers or autonomous... 
    Work at office
    Relocation

    Sonar

    Austin, TX
    22 days ago
  •  ...AI Technical Writer Location: San Antonio, TX FTE Only Experience Required: 7+ Years Must Have Technical/Functional Skills: Content Creation on AI/GEN AI: Writing, editing, and updating various materials, including user manuals, installation guides, release... 

    AceStack LLC

    San Antonio, TX
    19 hours ago
  •  ...build proof‑of‑concept applications to evaluate their features and limitations. Implement...  .... Architect and develop multi‑step AI agents that integrate with UI component libraries...  ...independently to deliver high‑quality, well‑tested code. Ideal Candidate Profile: Deep... 
    Permanent employment
    Contract work
    Local area

    Cloud Hybrid Technologies LLC

    Austin, TX
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Evaluation Scenario Writer - AI Agent Testing Specialist. Be the first to apply!