AI Agent Evaluation Analyst (Freelance)
$60 per hourMindrift
Location & Eligibility This opportunity is only for candidates currently residing in the specified country. Your location may affect eligibility and rates. Please submit your resume in English and indicate your level of English proficiency. About Mindrift At Mindrift, innovation meets opportunity. We believe in using the power of collective human intelligence to ethically shape the future of AI. What We Do The Mindrift platform, launched and powered by Toloka, connects domain experts with cutting‑edge AI projects from innovative tech clients. Our mission is to unlock the potential of GenAI by tapping into real‑world expertise from across the globe. Who We’re Looking For We’re looking for curious and intellectually proactive contributors—people who double‑check assumptions and play devil’s advocate. If you thrive in ambiguity, enjoy remote asynchronous work, and want to learn how modern AI systems are tested and evaluated, we want to hear from you. Project Overview We are seeking QA experts for autonomous AI agents in a project focused on validating and improving complex task structures, policy logic, and agent evaluation frameworks. Throughout the project, you will balance quality assurance, research, and logical problem‑solving. Responsibilities Review evaluation tasks and scenarios for logic, completeness, and realism. Identify inconsistencies, missing assumptions, or unclear decision points. Define clear expected behaviours (gold standards) for AI agents. Annotate cause‑effect relationships, reasoning paths, and plausible alternatives. Think through complex systems and policies as a human would to ensure agents are tested properly. Collaborate with QA, writers, or developers to suggest refinements or edge‑case coverage. Requirements Excellent analytical thinking: ability to reason about complex systems, scenarios, and logical implications. Strong attention to detail: spot contradictions, ambiguities, and vague requirements. Familiarity with structured data formats: read (not necessarily write) JSON/YAML. Ability to assess scenarios holistically: identify what’s missing, unrealistic, or potentially breaking. Good communication and clear writing (in English) to document findings. We also value applicants who have: Experience with policy evaluation, logic puzzles, case studies, or structured scenario design. Background in consulting, academia, olympiads (e.g. logic/math/informatics), or research. Exposure to LLMs, prompt engineering, or AI‑generated content. Familiarity with QA or test‑case thinking (edge cases, failure modes, "what could go wrong"). Some understanding of how scoring or evaluation works in agent testing (precision, coverage, etc.). Benefits Competitive pay up to $60/hour depending on skills, experience, and project needs. Flexible, remote, freelance project that fits around your primary professional or academic commitments. Advanced AI project experience to enhance your portfolio. Opportunity to influence how future AI models understand and communicate in your field of expertise. #J-18808-Ljbffr Mindrift
$60 per hour
A leading AI firm in Austin is looking for QA experts to validate... ...AI systems. This remote, freelance role requires strong analytical... .... Candidates will review AI evaluation tasks, identify... ...define expected behaviors for agents. Ideal applicants have experience...FreelanceRemote job- Join to apply for the Online Data Analyst Odia role at TELUS Digital AI Data Solutions Are you a detail-oriented... ...national and local geography? This freelance opportunity allows you to work at... ...worldwide Completing research and evaluation tasks in a web-based environment...FreelancePart timeLocal areaWorldwide
$60 per hour
...firm is seeking legal consultants with US law experience for part-time, project-based opportunities. You will generate prompts for AI, evaluate solutions, and improve reasoning standards. Ideal candidates have a law degree and 2+ years of legal experience. Strong written...FreelancePart time- ...collaborate with researchers on improving AI model performance in finance. You will apply... ...statistical analysis, and financial engineering to evaluate and train AI systems, with no prior AI experience required. This flexible, freelance-like role offers 10-30 hours per week,...FreelanceRemote jobHourly pay10 hours per weekFlexible hours
- We are seeking a skilled Data Analyst to perform complex data analysis supporting both... ...for data collection and presentation, evaluating data quality, and delivering actionable... ...commuter benefits to our employees, including freelancers - which sets us apart in the industries...FreelanceHourly payContract workWork at office
- Central Health is seeking a Compensation Analyst - Core Compensation in Austin, Texas. This role focuses on supporting compensation programs through market analysis and job evaluation, ensuring competitive and equitable pay practices across the organization. The ideal...
$112.5k - $147.5k
...looking for an experienced Senior Analyst, IT Internal Controls & SOX... ...role will be responsible for evaluating the design and operating effectiveness... ...assess risks associated with AI-enabled processes and... ...use of AI-enabled solutions, agents, automations, or productivity...Flexible hours- Strategic Data Analyst, Growth & PlacementAustin, TX | Grand Rapids, MI (Onsite 4 days per... ...a core expectation of this role.Leverage AI tools (Claude, Palantir AI FDE, Gemini, and... ...work, while applying rigorous critical evaluation to any AI-generated output before it informs...Full timeTemporary workImmediate startVisa sponsorshipFlexible hours
$26 - $28 per hour
...Data Labeling Analyst Welo Data is looking for detail-oriented and reliable individuals... ...Labeling Analysts, supporting speech and voice AI systems. This is a high-impact... ...this role is more execution-focused than evaluation-heavy roles, it still requires strong judgment...Full timeWork experience placementRemote workVisa sponsorship$34 per hour
...Data Labeling Analyst Welo Data is looking for detail-oriented and reliable individuals... ...Labeling Analysts, supporting speech and voice AI systems. This is a high-impact... ...this role is more execution-focused than evaluation-heavy roles, it still requires strong judgment...Work experience placementRemote work- Data Analyst, Data Intelligence Austin, TX (Onsite 4 days per week) Note: This is a full-time... ...and translate them to data and create/evaluate associated KPIs. Experience with SQL, Python... ...Demonstrated experience leveraging AI tools for natural language querying, including...Full timeTemporary workImmediate startVisa sponsorshipFlexible hours
$110k - $160k
Hellopatient is seeking a technical AI Agent Product Manager in Austin, Texas, to lead the delivery of AI agents tailored for healthcare... ...and continuously enhance agent performance through structured evaluation and real-world feedback. The position offers a competitive...$55 per hour
...intelligence to ethically shape the future of AI. What we do The Mindrift platform... ...Define comprehensive scoring criteria to evaluate the accuracy of the AI’s answers.... ...challenging, complex guidelines. Our freelance role is fully remote so, you just need a...FreelanceRemote jobFull timePart time$55 per hour
...and indicate your level of English proficiency. Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment. What This...FreelancePermanent employmentTemporary workPart time10 hours per week$73 per hour
A leading AI consultancy is seeking a Quantitative Statistics Expert to work flexibly as a freelance AI Trainer. This remote role requires a Bachelor's degree in Statistics and... ...include generating AI prompts and evaluating model accuracy, allowing you to impact the...FreelanceRemote job$50 per hour
...ethically shape the future of AI. What we do The Mindrift... ...generation and code review Prompt evaluation and complex data annotation... ...models Benchmarking and agent-based code execution in... ...complex guidelines. Our freelance role is fully remote so, you just...FreelanceRemote jobFull timePart time$60 per hour
...indicate your level of English proficiency. Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment. What This...FreelancePermanent employmentTemporary workPart time10 hours per week$55 per hour
A leading AI firm is seeking a Freelance Biology Expert with proficiency in Python to contribute to advanced AI projects from the comfort of your home. The role involves generating prompts, evaluating AI responses, and leveraging your expertise in Biology. Candidates should...FreelanceRemote job$60 - $64 per hour
...Tech SolutionsRole: Senior UI Engineer with AI agentsLocation: Austin, TXWe are looking... ...enable new ways of building UI with AI agents. The Frameworks team builds and maintains... ...and build proof-of-concept applications to evaluate their features and limitations.Implement...Hourly payFull time$110k - $160k
About the Role Hello Patient is hiring a technical, high-agency AI Agent Product Manager to own the end-to-end delivery of AI agents in... ...ll continuously iterate on agent performance using structured evaluation, testing, and real customer feedback. What You'll Do Own...$30 per hour
...industry's broadest and deepest suite of AI-powered cloud applications. The following... ...We are seeking a highly motivated AI Agent Intern to join Oracle's Supply Chain Applications... ...demo scripts. Research & Analysis Evaluate logistics AI use cases (forecasting, resilience...Hourly payTemporary workInternshipFlexible hours$55 per hour
A leading innovative tech firm seeks a Freelance AI Trainer specializing in Civil Engineering and Python. This part-time, fully remote role involves designing and evaluating AI models on civil engineering challenges. Candidates should have a Bachelor’s, Master’s, or PhD...FreelanceRemote jobHourly payPart time- A global leader in customer experience is seeking a Freelance Luxury Brand Evaluator in Austin, TX. In this role, you will assess customer experiences with high-end brands by visiting stores or evaluating online. Enjoy flexible assignments and compensation based on your...FreelanceFlexible hours
$80.9k - $115.5k
...informed strategy and execution. The Data Analyst, Data Analytics will work with customer,... ...functional delivery practices. Exposure to AI-enabled analytics, prompt engineering,... ...education, experience and skills and an evaluation of internal pay equity. Candidates who...Temporary workWork experience placementLocal areaImmediate startFlexible hours- ...Work with management to prioritize business and information needs Evaluate, analyze and conceptualize strategic risks and threats to trust... ..., submissions to this position are subject to the use of AI to perform preliminary candidate screenings, focused on ensuring...Work experience placementWork at officeLocal areaWork from homeFlexible hours
- ...etc.)Responsible for the maintenance, training and testing of upgrades to the Oracle Cloud ERP module.Evaluate, customize, and implement Oracle embedded AI and AI Agent Studio capabilities across Finance and Procurement modules.Assist in execution of changes to the...Full timeWork experience placement
$151.28k - $190k
...building the next generation of enterprise AI governance capabilities to enable secure,... ...scalable, and responsible adoption of AI agents. The team is evolving governance... ...iterate on AI agent capabilities, continuously evaluating and adopting new frameworks, features, and...Local areaWorldwideFlexible hours$23 per hour
Toloka connects curious people from around the world with freelance online tasks that train and improve artificial intelligence. Annotators review data, label content, and evaluate material according to project guidelines. This is a part-time, remote, freelance opportunity...FreelanceRemote jobHourly payPart time- RWS is seeking AI Data Specialists for flexible remote work in Texas. This freelance, part-time role focuses on improving AI-generated content in English, with immediate... ...commitment. Tasks include data collection, evaluation, annotation, and labeling across media types,...FreelanceRemote jobPart timeImmediate startFlexible hours
$90.9k - $130.7k
...proficient Senior Business Systems Analyst to spearhead the... ...primarily Gainsight and Staircase AI. In this role, you will act as... ...and implement future-state AI agents and automated workflows for customer... ...opportunity employer. We evaluate qualified applicants without regard...Local area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Agent Evaluation Analyst (Freelance). Be the first to apply!
- agent assistant Austin, TX
- tsa agent Austin, TX
- state farm agent Austin, TX
- import export agent Austin, TX
- remote chat agent Austin, TX
- freight agent no experience Austin, TX
- agent Austin, TX
- executive protection agent Austin, TX
- work from home chat agent Austin, TX
- telemarketer - state farm agent team member Austin, TX



