AI Agent Evaluation Analyst (Freelance)
$60 per hourMindrift
Location & Eligibility This opportunity is only for candidates currently residing in the specified country. Your location may affect eligibility and rates. Please submit your resume in English and indicate your level of English proficiency. About Mindrift At Mindrift, innovation meets opportunity. We believe in using the power of collective human intelligence to ethically shape the future of AI. What We Do The Mindrift platform, launched and powered by Toloka, connects domain experts with cutting‑edge AI projects from innovative tech clients. Our mission is to unlock the potential of GenAI by tapping into real‑world expertise from across the globe. Who We’re Looking For We’re looking for curious and intellectually proactive contributors—people who double‑check assumptions and play devil’s advocate. If you thrive in ambiguity, enjoy remote asynchronous work, and want to learn how modern AI systems are tested and evaluated, we want to hear from you. Project Overview We are seeking QA experts for autonomous AI agents in a project focused on validating and improving complex task structures, policy logic, and agent evaluation frameworks. Throughout the project, you will balance quality assurance, research, and logical problem‑solving. Responsibilities Review evaluation tasks and scenarios for logic, completeness, and realism. Identify inconsistencies, missing assumptions, or unclear decision points. Define clear expected behaviours (gold standards) for AI agents. Annotate cause‑effect relationships, reasoning paths, and plausible alternatives. Think through complex systems and policies as a human would to ensure agents are tested properly. Collaborate with QA, writers, or developers to suggest refinements or edge‑case coverage. Requirements Excellent analytical thinking: ability to reason about complex systems, scenarios, and logical implications. Strong attention to detail: spot contradictions, ambiguities, and vague requirements. Familiarity with structured data formats: read (not necessarily write) JSON/YAML. Ability to assess scenarios holistically: identify what’s missing, unrealistic, or potentially breaking. Good communication and clear writing (in English) to document findings. We also value applicants who have: Experience with policy evaluation, logic puzzles, case studies, or structured scenario design. Background in consulting, academia, olympiads (e.g. logic/math/informatics), or research. Exposure to LLMs, prompt engineering, or AI‑generated content. Familiarity with QA or test‑case thinking (edge cases, failure modes, "what could go wrong"). Some understanding of how scoring or evaluation works in agent testing (precision, coverage, etc.). Benefits Competitive pay up to $60/hour depending on skills, experience, and project needs. Flexible, remote, freelance project that fits around your primary professional or academic commitments. Advanced AI project experience to enhance your portfolio. Opportunity to influence how future AI models understand and communicate in your field of expertise. #J-18808-Ljbffr Mindrift
$60 per hour
...A leading AI firm in Austin is looking for QA experts to validate... ...AI systems. This remote, freelance role requires strong analytical... .... Candidates will review AI evaluation tasks, identify... ...define expected behaviors for agents. Ideal applicants have experience...FreelanceRemote work- Join to apply for the Online Data Analyst Odia role at TELUS Digital AI Data Solutions Are you a detail-oriented... ...national and local geography? This freelance opportunity allows you to work at... ...worldwide Completing research and evaluation tasks in a web-based environment...FreelancePart timeLocal areaWorldwide
$60 per hour
...firm is seeking legal consultants with US law experience for part-time, project-based opportunities. You will generate prompts for AI, evaluate solutions, and improve reasoning standards. Ideal candidates have a law degree and 2+ years of legal experience. Strong written...FreelancePart time- ...technology company for Bitcoin mining and AI cloud. Bitdeer is committed to... ...responsible for: ~ This role builds the evaluation and decision systems that make agentic inference... .... You will own and extend our LLM and agent evaluation pipeline, develop representative...SuggestedFull time
- The Texas Health and Human Services Commission (HHSC) invites applications for a Research Analyst (Research Specialist V) on the Applied Statistics & Evaluation Team. The role requires designing and conducting advanced research, data collection and reporting to inform...SuggestedRemote work
- ...Social Factor is looking to find US-based, qualified, available freelance Data Analysts to jump in on projects or new business opportunities that are currently in the pipeline. The role will be focused primarily on Sprinklr Implementation projects like building dashboards...FreelanceTemporary workCasual workRemote work
$73 per hour
A leading AI consultancy is seeking a Quantitative Statistics Expert to work flexibly as a freelance AI Trainer. This remote role requires a Bachelor's degree in Statistics and... ...include generating AI prompts and evaluating model accuracy, allowing you to impact the...FreelanceRemote job$55 per hour
...and indicate your level of English proficiency. Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment. What This...FreelancePermanent employmentTemporary workPart time10 hours per week- ...are public servants committed to improving opportunities for students and supporting those who serve them. Job Description The Evaluation and Research Scientist works in the Evaluation Activities Unit within the Division of Research and Analysis. This role performs complex...Contract workImmediate startFlexible hours
$60 per hour
...indicate your level of English proficiency. Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment. What This...FreelancePermanent employmentTemporary workPart time10 hours per week$110k - $160k
Hellopatient is seeking a technical AI Agent Product Manager in Austin, Texas, to lead the delivery of AI agents tailored for healthcare... ...and continuously enhance agent performance through structured evaluation and real-world feedback. The position offers a competitive...$55 per hour
A leading AI firm is seeking a Freelance Biology Expert with proficiency in Python to contribute to advanced AI projects from the comfort of your home. The role involves generating prompts, evaluating AI responses, and leveraging your expertise in Biology. Candidates should...FreelanceRemote job$55 per hour
...intelligence to ethically shape the future of AI. What we do The Mindrift platform... ...Define comprehensive scoring criteria to evaluate the accuracy of the AI’s answers.... ...challenging, complex guidelines. Our freelance role is fully remote so, you just need a...FreelanceRemote jobFull timePart time$30 per hour
...industry's broadest and deepest suite of AI-powered cloud applications. The following... ...We are seeking a highly motivated AI Agent Intern to join Oracle's Supply Chain Applications... ...demo scripts. Research & Analysis Evaluate logistics AI use cases (forecasting, resilience...Hourly payTemporary workInternshipFlexible hours- ...experienced Senior Software Engineers to support an AI training project by creating reinforcement learning environments that evaluate AI models on complex software engineering... ...golden reference solutions. Evaluate AI agents' ability to reason through complex codebases...Remote jobFor contractors
- A global leader in customer experience is seeking a Freelance Luxury Brand Evaluator in Austin, TX. In this role, you will assess customer experiences with high-end brands by visiting stores or evaluating online. Enjoy flexible assignments and compensation based on your...FreelanceFlexible hours
$55 per hour
A leading innovative tech firm seeks a Freelance AI Trainer specializing in Civil Engineering and Python. This part-time, fully remote role involves designing and evaluating AI models on civil engineering challenges. Candidates should have a Bachelor’s, Master’s, or PhD...FreelanceRemote jobHourly payPart time$100k - $130k
...depends on it. We are looking for a Data Analyst II, Ecommerce to own analytics for Greenlight... ...Amazon Ads and Shopify-driven channels. Evaluate ROAS, CAC, and payback period at the... ...Help define and share best practices for AI-assisted analytics across the team. Partner...Work at officeLocal areaRemote workWork from homeFlexible hoursShift workDay shift$60 - $64 per hour
...Tech SolutionsRole: Senior UI Engineer with AI agentsLocation: Austin, TXWe are looking... ...enable new ways of building UI with AI agents. The Frameworks team builds and maintains... ...and build proof-of-concept applications to evaluate their features and limitations.Implement...Hourly payFull time$97.5k - $209.5k
...Services team is hiring a Principal Data Analyst, where you will analyze, prepare, and process... ...other pertinent data health measures; evaluate complex data sets for analytical... ...innovations to life-saving care. And with AI embedded across our products and services...Contract workTemporary workWork experience placementLocal areaFlexible hours$60 per hour
...Freelance Software Developer (Ruby) - AI Trainer Location : Remote, full‑time, freelance Overview At Mindrift, we connect AI projects with specialists... ...AI models. Define comprehensive scoring criteria to evaluate AI responses. Correct model outputs based on domain...FreelanceFull timePart timeRemote workWorldwide- ...operational performance. As a Strategic Data Analyst on the Growth & Placement analytics pod,... ...core expectation of this role. Leverage AI tools (Claude, Palantir AI FDE, Gemini,... ...engineering work, while applying rigorous critical evaluation to any AI‑generated output before it...Full timeTemporary workVisa sponsorship
- ...etc.)Responsible for the maintenance, training and testing of upgrades to the Oracle Cloud ERP module.Evaluate, customize, and implement Oracle embedded AI and AI Agent Studio capabilities across Finance and Procurement modules.Assist in execution of changes to the...Full timeWork experience placement
- RWS is seeking AI Data Specialists for flexible remote work in Texas. This freelance, part-time role focuses on improving AI-generated content in English, with immediate... ...commitment. Tasks include data collection, evaluation, annotation, and labeling across media types,...FreelanceRemote jobPart timeImmediate startFlexible hours
$34 per hour
...Data Labeling Analyst Welo Data is looking for detail-oriented and reliable individuals... ...Labeling Analysts, supporting speech and voice AI systems. This is a high-impact... ...this role is more execution-focused than evaluation-heavy roles, it still requires strong judgment...Full timeWork experience placementRemote workVisa sponsorship$73 per hour
Quantitative Statistics Expert - Freelance AI Trainer This opportunity is only for candidates currently residing in the specified country... ...that challenge AI. Define comprehensive scoring criteria to evaluate the accuracy of the AI's answers. Correct the model's responses...FreelancePart timeRemote work$23 per hour
...curious people from around the world with freelance online tasks that train and improve... ...Annotators connects individuals with Generative AI projects from leading tech innovators.... ...projects such as rating AI-generated content, evaluating factual accuracy, or comparing responses...FreelancePart timeRemote work- ...Role Haven is looking for a contract analyst to do hands-on analytical work across the... ...dbt) rather than building infrastructure. AI tooling handles much of the routine pipeline... ..., program, and customer segment Evaluate channel and campaign performance against...Contract workLocal areaRemote workShift work
- ...Description Job Description Data Privacy Analyst Job Type: Contractor Location:... ...Analysts to support a data privacy and AI training project focused on protecting sensitive... ...improperly redacted sensitive data. Evaluate data against established privacy and...Contract workFor contractorsRemote work
- ...robotics and software, ACS brings together AI, computer vision, precision motion, and... ...are looking for a detail-oriented Data Analyst to support our rapidly growing Supply Chain... ...management review Survey suppliers to evaluate production capacity and readiness for...Work at officeLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Agent Evaluation Analyst (Freelance). Be the first to apply!
- executive protection agent Austin, TX
- cruise agent Austin, TX
- telemarketer - state farm agent team member Austin, TX
- state farm agent Austin, TX
- work from home chat agent Austin, TX
- agent Austin, TX
- agent assistant Austin, TX
- airport agent Austin, TX
- commissioning agent Austin, TX
- operations agent Austin, TX




