Remote AI Agent Evaluation Specialist
$80 per hourMindrift
- Remote job
A leading tech company is seeking contributors for a flexible part-time opportunity focused on quality assurance for autonomous AI agents. You will analyze complex systems, review tasks for logic, and ensure clear expected behaviors for AI. Ideal candidates possess excellent analytical and detail-oriented skills, with experience in policy evaluation or logic puzzles preferred. Compensation can reach up to $80/hour based on expertise and project needs. This role offers valuable experience in an advanced AI project and fits around your primary commitments. #J-18808-Ljbffr Mindrift
$60 per hour
...contributors for a part-time QA project focused on autonomous AI agents. This flexible remote opportunity requires strong analytical and critical... ...with structured data formats. Candidates will review evaluation tasks, identify inconsistencies, and help define expected...Remote workPart timeFlexible hours$120k - $145k
...Summary The Senior Systems Specialist is responsible for installing... ...technical expert, providing support remotely and, as needed, at production... ...routine analysis to include evaluation of hardware and software... ...Microsoft Copilot Administration - Agent building, Prompt Engineering...Remote workFull timeContract workWork at officeWorldwideHome office- ...technology company is looking for a detail-oriented individual to design structured evaluation scenarios for AI agents. This entry-level part-time role allows you to contribute remotely, creating test cases for LLM-based agents. The ideal candidate will have a background...Remote jobPart time
$80 per hour
A leading AI consultancy is seeking a detail-oriented individual to design evaluation scenarios for LLM-based agents. This part-time, remote role allows you to create structured test cases while working flexibly around your commitments. The ideal candidate holds a relevant...Remote workPart time$80 per hour
A technology firm is seeking QAs for autonomous AI agents to validate and improve task structures within a new project. Candidates will... ...thinkers with strong attention to detail and experience in evaluating scenarios. This flexible, project-based role offers competitive...Remote jobFlexible hours- AI Agent Development & Evaluation Primary Care, Acute Symptom Triage & Travel Health | Contract, up to 10 hrs/week About the Role We are seeking Nurses... ...adherence, and patient safety. This is a part-time, remote contract position (1099) requiring up to 10 hours per week...Remote workHourly payContract workPart time10 hours per weekFlexible hours
$80 per hour
...ethically shape the future of AI. The Mindrift platform, launched... ...complexity? Does an async, remote, flexible opportunity sound exciting... ...AI systems are tested and evaluated? Analysts, researchers or... ...hunt for QAs for autonomous AI agents for a new project focused on validating...Remote workPermanent employmentPart timeFreelanceFlexible hours- ESR Healthcare is looking for a Robotics Evaluation Specialist to work remotely, applying expertise to train AI systems. This role includes executing physical actions on camera, evaluating robotic performance, and providing structured feedback. The ideal candidate should...Remote job
$60 per hour
A leading AI firm in Austin is looking for QA experts to validate... ...and improve AI systems. This remote, freelance role requires... ...detail. Candidates will review AI evaluation tasks, identify inconsistencies... ...define expected behaviors for agents. Ideal applicants have experience...Remote jobFreelance$24 - $30 per hour
Mercor is seeking a Moonlight MCQA - Arabic Generalist to draft Arabic MCQs for an AI evaluation dataset. This contract role offers remote work with flexible hours and a process including AI interview based on your resume. Must hold a master's or PhD, 2-6 years of experience...Remote jobHourly payContract workFreelanceFlexible hours$80 per hour
...ethically shape the future of AI. What We Do The Mindrift platform connects specialists with AI projects from major... ...realistic and structured evaluation scenarios for LLM-based agents. You'll create test cases... ...Take part in a flexible, remote, freelance project that fits...Remote workPart timeFreelanceFlexible hours$55 per hour
A leading AI innovation firm in Dallas is seeking QAs for autonomous AI agents to ensure the quality of complex systems and scenarios. This flexible, remote project is ideal for those with excellent analytical... ...should be adept at evaluating scenarios and documenting findings...Remote jobFlexible hours- ...Mathematics and Statistics experts to support the training and evaluation of state-of-the-art AI systems. We need subject-matter experts who can apply... ...meaningful errors within their field. Flexible, remote contract work starting ASAP. You will work on challenging...Remote jobContract workImmediate startFlexible hours
- ...SWE Agent Evaluation Specialist is a remote engineering review track for evaluating production code, debugging traces, and developer-facing AI outputs against real-world correctness standards. Reviewers reproduce failures, write the unit test the model should have written...Remote jobHourly payFor contractors10 hours per week
- Dorado is seeking a Physics Specialist to contribute deep scientific expertise to AI model evaluation. You will craft and assess challenging physics problems, probe model reasoning at the frontier, and help identify where models fail under rigorous scientific scrutiny....Remote job
- ...Expert Codebase Evaluation Specialist - Coding Agent Review is a remote evaluation track for reviewing codebase evaluation evaluation prompts and responses against... ...team can use to retrain. Why this role matters AI data reviewers help turn codebase evaluation...Remote jobHourly payFor contractors10 hours per week
$80 per hour
...Be among the first 25 applicants Get AI‑powered advice on this job and more exclusive... ...and internal tools for running and evaluating agent behavior. You'll implement base methods... ...project needs Take part in a flexible, remote, freelance project that fits around your...Remote workTemporary workPart timeFreelanceFlexible hours$80 per hour
...technology company is seeking a part-time QA contributor for an AI project to validate autonomous agents. Candidates must have strong analytical thinking,... ...expected behaviors for AI agents. This flexible remote opportunity allows you to work on your own schedule while...Remote jobPart timeFlexible hours$60 - $75 per hour
...Applied Biology Benchmark Specialist - AI Evaluation is a remote review track for evaluating AI outputs across biology reasoning, calculations, and research workflows. Reviewers grade derivations and assumptions, reproduce key results, and document the correct method...Remote jobFor contractors10 hours per week$125 per hour
...Role Overview ParaView Specialists apply hands-on experience with scientific visualization... ..., and large dataset analysis to evaluate AI-generated outputs and guide AI systems... ...Comfortable working asynchronously with remote research teams and providing clear, detailed...Remote workHourly payFull timeContract workPart timeFor contractorsFlexible hours$1,750 - $2,150 per month
...connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our... ...,750–$2,150 per completed task Location: Remote Role Responsibilities Review and evaluate AI-generated outputs related to threat analysis,...Remote workHourly payFull timeContract workSummer work$152k - $240k
...control spend effortlessly. Brex's AI-native automation and world-... ...We're building AI agents to automate and augment internal... ...four weeks per year of fully remote work! Responsibilities... ...and data sources. Define evaluation frameworks, success metrics, and...Remote workFull timeWork at officeWork from home$200k - $320k
...technical depth and a passion for AI-driven product development.... .... This is a full-time remote opportunity, with preference for... ...production-ready LLM pipelines and AI agent systems Develop AI-driven... ...new product initiatives Evaluate and recommend AI architectures...Remote workFull timeVisa sponsorship$202.5k - $247.5k
...sharing localhost or running AI workloads in production. We... ...worth your time. About the Agent Team Our Agent team... ...AWS. Engineers develop by using remote development tools and/or ssh to... ...and actual compensation will be evaluated based on factors including,...Remote workPermanent employmentFull timeWork at officeLocal areaImmediate startHome officeFlexible hours$162.8k - $244.2k
...Fortune 500 company and a leading AI platform for managing people, money, and agents, we’re shaping the future of work... ...-on with agent architectures, evaluation frameworks, and customer workflows... ...both worlds: in-person time and remote. Our approach enables our teams to...Remote workFull timeWork at officeHome officeFlexible hoursWeekend work$170k - $225k
...Description As an Enterprise Agent Engineer, you'll build the system... ...it matters ~Instrument and evaluate agents by measuring... ...hands-on experience building AI agents or agentic automations... ...plus bonus and equity ~Fully remote with flexible hours (core hours...Remote workFull timeFlexible hours- ...is seeking a Vietnamese Voice Acting Specialist for a freelance AI Trainer project. The role is critical... ...emotional expression. This position is remote and designed for individuals with... ...voice acting credentials. You will evaluate AI outputs and support the advancement...Remote workHourly payFreelance
$30 per hour
...Description Mindrift connects specialists with project-based AI opportunities for leading... ..., focused on testing, evaluating, and improving AI systems.... ...dataset to evaluate AI coding agents - how well a model handles... ...Take part in a part-time, remote, freelance project that...Remote workHourly payPermanent employmentTemporary workPart timeFreelance$112.09k - $160.33k
...Location: Burlington, VT (hybrid) or remote About OhMD OhMD is an AI patient communication platform used... ...— and OhMD AI, our AI voice agent. It picks up on the first ring, handles... ...choices you made Familiarity with evaluating agent quality in a structured way...Remote workLive in- Zillow Group’s Agentic AI team seeks a Principal Applied Scientist to set the science... ...for advanced reasoning and long-running agent systems. You will partner with researchers... ...contributor role emphasizes leadership, rigorous evaluation, and cross-functional impact, including...Remote job
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Remote AI Agent Evaluation Specialist. Be the first to apply!






