Evaluation Scenario Writer - AI Agent Testing Specialist
$80 per hourMindrift
This opportunity is only for candidates currently residing in the specified country. Your location may affect eligibility and rates. Please submit your resume in English and indicate your level of English.
At Mindrift, innovation meets opportunity. We believe in using the power of collective human intelligence to ethically shape the future of AI.
What We Do
The Mindrift platform connects specialists with AI projects from major tech innovators. Our mission is to unlock the potential of Generative AI by tapping into real‑world expertise from across the globe.
About the Role
We’re looking for someone who can design realistic and structured evaluation scenarios for LLM‑based agents. You’ll create test cases that simulate human‑performed tasks and define gold‑standard behavior to compare agent actions against. You’ll work to ensure each scenario is clearly defined, well‑scored, and easy to execute and reuse. You’ll need a sharp analytical mindset, attention to detail, and an interest in how आरोच्छ in how AI agents make decisions.
Responsibilities- \ wallet Create structured test cases that simulate complex human workflows
- Define gold‑standard behavior and scoring logic to evaluate agent actions.
- Analyze agent logs, failure modes, and decision paths.
- Work with code repositories and test frameworks to validate your scenarios.
- Iterate on prompts, instructions, and test cases to improve clarity and difficulty.
- Ensure that scenarios are production‑ready, easy to run, and reusable.
How To Get Started
Simply apply to this post, qualify, виды and get the chance to contribute to projects aligned with your skills, on your own schedule. From creating training prompts to refining model responses, you’ll help shape the future of AI while ensuring technology benefits everyone.
Requirements
- Bachelor’s and/or Master’s Degree in Computer Science, Software Engineering, Data Science / Data Analytics, Artificial Intelligence / Machine Learning, Computational Linguistics / Natural Language Processing (NLP), Information Systems, or 天天中彩票为什么 related fields.
- Background in QA, software testing, data analysis, or NLP annotation.
- Good understanding of test design principles (e.g., reproducibility, coverage, edge cases).
- Strong written communication skills in English.
- Comfortable with structured formats like JSON/YאםLAML for scenario description.
- Can define expected agent behaviors (gold paths) and scoring logic.
- Basic experience with Python and JS.
- Curious and open to working with AI‑generated content, agent logs, and prompt‑based behavior.
Nice to Have
- Experience in writing manual or automated test cases.
- Familiarity with LLM capabilities and typical failure modes.
- Understanding of scoring metrics (precision, recall, coverage, reward functions).
EEP Benefits
- Get paid for your expertise, with rates that can go up to $80/hour depending on your skills, experience, and project needs.
- Take part in a flexible, remote, freelance project that fits around your primary professional or academic commitments.
- Participate in an advanced AI project and gain valuable experience to enhance your portfolio.
- Influence how future AI models understand and communicate in your field of expertise.
Seniority Level
Entry level
Employment Type минуты
Part‑time
Job Function
Other
Industries
IT Services and IT Consulting
#J-18808-Ljbffr$80 per hour
A leading AI consultancy is seeking a detail-oriented individual to design evaluation scenarios for LLM-based agents. This part-time, remote role allows you to create structured test cases while working flexibly around your commitments. The ideal candidate holds a relevant...SuggestedPart timeRemote work$80 per hour
...ethically shape the future of AI. What We Do The... ...modern AI systems are tested and evaluated? This is a flexible,... ...QAs for autonomous AI agents for a new project... ...and thinking through scenarios, implications, and edge... ...Working closely with QA, writers, or developers to suggest...SuggestedPermanent employmentPart timeFreelanceRemote workFlexible hours$55 per hour
A leading AI innovation firm in Dallas is seeking QAs for autonomous AI agents to ensure the quality of complex systems and scenarios. This flexible, remote project is ideal for those with excellent... .... Candidates should be adept at evaluating scenarios and documenting...SuggestedRemote jobFlexible hours$80 per hour
Get AI‑powered advice on this job and more exclusive features... ...tools for running and evaluating agent behavior. You’ll implement... ...check agent actions against scenario definitions Creating or extending tools that writers and QAs use to test agents Working closely with...SuggestedFreelanceRemote workFlexible hours$60 per hour
A leading AI firm in Austin is looking for QA experts to validate and improve AI systems. This... ...to detail. Candidates will review AI evaluation tasks, identify inconsistencies, and define expected behaviors for agents. Ideal applicants have experience in policy evaluation...SuggestedRemote jobFreelance$45.5 - $65 per hour
...Cypress, Python, PyTest, JMeter, AI Testing - Mandatory.1-2 days onsite at client... ...performance.Build automated test scenarios for AI-driven workflows, including... ..., chatbots, and intelligent agents.Develop test datasets and evaluation frameworks for AI model validation...$30 per hour
...broadest and deepest suite of AI-powered cloud... ...a highly motivated AI Agent Intern to join Oracle'... ...workflows for logistics scenarios (e.g., shipment risk detection... ...datasets for agent testing. Solution... ...Research & Analysis Evaluate logistics AI use cases...Hourly payTemporary workInternshipFlexible hours- ...Eleven is looking for a Staff Engineer - Agents to join the Enterprise AI Team!At 7-Eleven Enterprise AI Team,... ...development, model training, testing, and deploymentConduct deep dives into... ...continued performance and reliability.Evaluate and grow technical talent throughout...Hourly pay
- ...Opportunity for advancement Senior Product Manager, AI Agents Work Type: Full-Time, Onsite Location: Dallas, Texas... ...technology partners and vendors in the AI ecosystem. Lead A/B testing, agent evaluation, and post-deployment analysis to drive continuous...Full time
$110k - $160k
Hellopatient is seeking a technical AI Agent Product Manager in Austin, Texas, to lead the delivery of AI agents tailored for healthcare... ...and continuously enhance agent performance through structured evaluation and real-world feedback. The position offers a competitive...$80 - $100 per hour
...Role Overview Help train next-generation AI coding agents by contributing expert, real-world examples, technical walkthroughs, and structured assessments. In this contractor role you will document how AI coding tools support complex software development tasks, explain...Remote jobHourly payFor contractors$66.3k - $84k
...Technical Writer At Jacobs, we're challenging today to... ...will also involve leveraging AI software such as Co-Pilot Agent and Power BI, working... ...including inspection and test plans, reports, templates,... ...Apply critical thinking to evaluate technical content, identify...For subcontractor$60 - $64 per hour
...SolutionsRole: Senior UI Engineer with AI agentsLocation: Austin, TXWe... ...ways of building UI with AI agents. The Frameworks team builds... ...-of-concept applications to evaluate their features and limitations... ...to deliver high-quality, well-tested code.Ideal Candidate Profile:Deep...Hourly payFull time- ...Consulting is looking for a part-time remote Copywriting & Content Subject Matter Expert (SME) to review and create AI-generated content. You will evaluate the quality of AI outputs and provide expertise to enhance persuasive and accurate writing. The ideal candidate has...Remote jobPart timeImmediate start
- ...in Dallas, TX is seeking an experienced AI Architect to act as a trusted advisor to... ...technology leaders while leading the design, evaluation, and implementation of enterprise-scale... ...business opportunities, design custom AI agents, evaluate emerging AI products, and establish...
$110k - $160k
About the Role Hello Patient is hiring a technical, high-agency AI Agent Product Manager to own the end-to-end delivery of AI agents... ...continuously iterate on agent performance using structured evaluation, testing, and real customer feedback. What You'll Do Own customer...$151.28k - $190k
...next generation of enterprise AI governance capabilities to enable... ...responsible adoption of AI agents. The team is evolving governance... ...capabilities, continuously evaluating and adopting new frameworks, features... ...design patterns, automated testing, and system integration.Strong...Local areaWorldwideFlexible hours- ...Power is standing up a dedicated AI build team to make Quote-to-... ...a versioned prompt and evaluation library to support consistent,... ...Context Protocol) or similar agent-integration patterns. Experience... ...harnesses or systematic prompt testing. Exposure to CRM (HubSpot) or...
- ...stakeholders and iterate the wireframe screens. To test with sample users to ensure that the... ...create a persona (user description) and scenario (interaction with product). To utilize the current product (heuristic evaluation) or findings obtained, competitor product analysis...
- ...build proof-of-concept applications to evaluate their features and limitations. · Implement... .... · Architect and develop multi-step AI agents that integrate with UI component... ...independently to deliver high-quality, well-tested code. Ideal Candidate Profile: ·...
$89.2k - $209.5k
...Description Oracle Health is seeking a Senior AI Agent Engineer to build production AI agents... ..., balancing speed with safety, evaluation, and maintainability. This person should... .... Responsibilities Design, build, test, and deploy production AI agents and AI-...Temporary workFlexible hours$93.75k - $133.2k
...civilian hiring freeze. As an FBI special agent, your career is defined by how you think... ..., and operate across a range of complex scenarios. Every day brings new challenges that... ...accounting and finance reflects how you evaluate information, navigate complexity, and operate...Work at officeLocal areaShift work- ...Job Description The Role: We are seeking an AI Agent Engineer to design, build, and operationalize AI-powered agents... ...improve agent performance. Establish and maintain evaluation frameworks, metrics, and test harnesses for AI agent behavior and output quality....Local areaWork from homeRelocationRelocation package
$262k - $364k
...strategy and architecture for AI-ready data foundations, including... ...search and summarization agents, alongside modular frameworks... ...time queries.Design automated evaluation frameworks to improve agent accuracy... ...design and architecture; and testing/launching software products....$207k - $301k
...scalability and efficiency of Orcas Agent Runtime Platform.Work cross-... ....Architect and develop novel evaluation and customization solutions,... ...assessment ( Cloud Lifecycle AI Tools).Mentor and provide technical... ....5 years of experience testing, and launching software products...- ...has developed an artificial intelligence (AI) powered technology stack purpose-built... ...commercial self-driving software to develop, test and deploy autonomous capabilities for... ...screening of applications. As part of the evaluation process, we provide Endorsed with job...Visa sponsorshipShift workNight shiftWeekend work
- Who is Sonar? Sonar is driving the future of agent-centric software development. As the leader in AI code verification and governance, we solve a critical problem: ensuring that software generated by AI-assisted developers or autonomous agents is reliable, secure, and maintainable...Full time
- ...Job Description Job Description Who is Sonar? Sonar is driving the future of agent-centric software development. As the leader in AI code verification and governance, we solve a critical problem: ensuring that software generated by AI-assisted developers or autonomous...Work at officeRelocation
- ...AI Technical Writer Location: San Antonio, TX FTE Only Experience Required: 7+ Years Must Have Technical/Functional Skills: Content Creation on AI/GEN AI: Writing, editing, and updating various materials, including user manuals, installation guides, release...
- ...build proof‑of‑concept applications to evaluate their features and limitations. Implement... .... Architect and develop multi‑step AI agents that integrate with UI component libraries... ...independently to deliver high‑quality, well‑tested code. Ideal Candidate Profile: Deep...Permanent employmentContract workLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Evaluation Scenario Writer - AI Agent Testing Specialist. Be the first to apply!




