Remote AI Agent Evaluation Specialist
$80 per hourMindrift
- Remote job
A technology firm is seeking QAs for autonomous AI agents to validate and improve task structures within a new project. Candidates will work on complex systems, ensuring quality assurance and logical problem-solving. Ideal candidates are analytical thinkers with strong attention to detail and experience in evaluating scenarios. This flexible, project-based role offers competitive pay up to $80/hour, allowing integration with academic or professional commitments. #J-18808-Ljbffr Mindrift
$120k - $145k
...Summary The Senior Systems Specialist is responsible for installing... ...technical expert, providing support remotely and, as needed, at production... ...routine analysis to include evaluation of hardware and software... ...Microsoft Copilot Administration - Agent building, Prompt Engineering...Remote workFull timeContract workWork at officeWorldwideHome office$80 per hour
A leading AI consultancy is seeking a detail-oriented individual to design evaluation scenarios for LLM-based agents. This part-time, remote role allows you to create structured test cases while working flexibly around your commitments. The ideal candidate holds a relevant...Remote workPart time$60 per hour
...A leading AI firm in Austin is looking for QA experts to validate... ...and improve AI systems. This remote, freelance role requires... ...detail. Candidates will review AI evaluation tasks, identify inconsistencies... ...define expected behaviors for agents. Ideal applicants have experience...Remote workFreelance- ...technology company is looking for a detail-oriented individual to design structured evaluation scenarios for AI agents. This entry-level part-time role allows you to contribute remotely, creating test cases for LLM-based agents. The ideal candidate will have a background...Remote jobPart time
$80 per hour
...-time opportunity focused on quality assurance for autonomous AI agents. You will analyze complex systems, review tasks for logic, and... ...analytical and detail-oriented skills, with experience in policy evaluation or logic puzzles preferred. Compensation can reach up to $80/...Remote jobPart timeFlexible hours$60 per hour
...contributors for a part-time QA project focused on autonomous AI agents. This flexible remote opportunity requires strong analytical and critical... ...with structured data formats. Candidates will review evaluation tasks, identify inconsistencies, and help define expected...Remote jobPart timeFlexible hours$23 - $26.45 per hour
Appen is seeking Independent Contractors for a remote Audio Evaluation Project. You will listen to two audio recordings per case and decide which is better based on the provided guidelines. The role emphasizes precision, consistency, and adherence to instructions. The project...Remote jobHourly payFor contractorsWork from homeFlexible hours$80 per hour
...ethically shape the future of AI. The Mindrift platform, launched... ...complexity? Does an async, remote, flexible opportunity sound exciting... ...AI systems are tested and evaluated? Analysts, researchers or... ...hunt for QAs for autonomous AI agents for a new project focused on validating...Remote workPermanent employmentPart timeFreelanceFlexible hours- ESR Healthcare is looking for a Robotics Evaluation Specialist to work remotely, applying expertise to train AI systems. This role includes executing physical actions on camera, evaluating robotic performance, and providing structured feedback. The ideal candidate should...Remote job
$24 - $30 per hour
Mercor is seeking a Moonlight MCQA - Arabic Generalist to draft Arabic MCQs for an AI evaluation dataset. This contract role offers remote work with flexible hours and a process including AI interview based on your resume. Must hold a master's or PhD, 2-6 years of experience...Remote jobHourly payContract workFreelanceFlexible hours$80 per hour
...ethically shape the future of AI. Our platform connects specialists with AI projects from major tech... ...realistic and structured evaluation scenarios for LLM-based agents. You'll create test cases that... ...needs Take part in a flexible, remote, freelance project that fits around...Remote workPart timeFreelanceFlexible hours$55 per hour
A leading AI innovation firm in Dallas is seeking QAs for autonomous AI agents to ensure the quality of complex systems and scenarios. This flexible, remote project is ideal for those with excellent analytical... ...should be adept at evaluating scenarios and documenting findings...Remote jobFlexible hours- ...Mathematics and Statistics experts to support the training and evaluation of state-of-the-art AI systems. We need subject-matter experts who can apply... ...meaningful errors within their field. Flexible, remote contract work starting ASAP. You will work on challenging...Remote jobContract workImmediate startFlexible hours
- ...SWE Agent Evaluation Specialist is a remote engineering review track for evaluating production code, debugging traces, and developer-facing AI outputs against real-world correctness standards. Reviewers reproduce failures, write the unit test the model should have written...Remote jobHourly payFor contractors10 hours per week
- Dorado is seeking a Physics Specialist to contribute deep scientific expertise to AI model evaluation. You will craft and assess challenging physics problems, probe model reasoning at the frontier, and help identify where models fail under rigorous scientific scrutiny....Remote job
- ...Expert Codebase Evaluation Specialist - Coding Agent Review is a remote evaluation track for reviewing codebase evaluation evaluation prompts and responses against... ...team can use to retrain. Why this role matters AI data reviewers help turn codebase evaluation...Remote jobHourly payFor contractors10 hours per week
$90 per hour
...creative and technical talent with leading AI research labs. Headquartered in San... ...Compensation: $90/hour Location: Remote Role Responsibilities Render a reference... ...RLHF , preference-labeling, model-evaluation, or structured code-review work. Pixel...Remote workContract workSummer workLocal area$80 - $100 per hour
...Role Overview Help train and evaluate AI coding agents by contributing real-world STEM workflows, problem examples, and candid feedback. In this... ...AI systems Work Terms Contractor position, fully remote. The engagement supports a customer project focused on improving...Remote workHourly payFor contractors$80 - $100 per hour
...help train and improve next-generation AI coding agents. In this contractor role you will provide... ...end feature development. Critically evaluate AI agents in diverse scenarios,... ...Engagement type: Contractor. Location: Remote. Work will involve contributing to a...Remote jobHourly payFor contractors$80 - $120 per hour
...elite creative and technical talent with leading AI research labs. Headquartered in San Francisco... .... Position: Process improvement / SOPs Evaluator Type: Contract Compensation: $80–$120/hour Location: Remote Role Responsibilities Evaluate AI-...Remote workContract workSummer workWork at office$80 per hour
Get AI‑powered advice on this job and more exclusive features. This opportunity is... ...servers and internal tools for running and evaluating agent behavior. You’ll implement base methods... ...project needs Take part in a flexible, remote, freelance project that fits around your...Remote workFreelanceFlexible hours$80 - $120 per hour
...elite creative and technical talent with leading AI research labs. Headquartered in San Francisco... .... Position: Humanities / arts / culture Evaluator Type: Contract Compensation: $80–$120/hour Location: Remote Role Responsibilities Evaluate AI-...Remote workContract workSummer workWork at office$40 - $50 per hour
...Overview Apply your linguistics expertise to evaluate large language model outputs and help train next-generation AI systems. You will analyze model-human conversations... ...contractor relationship. Location: Remote, applicants must be located in the United States...Remote jobHourly payFor contractorsImmediate start$80 per hour
...technology company is seeking a part-time QA contributor for an AI project to validate autonomous agents. Candidates must have strong analytical thinking,... ...expected behaviors for AI agents. This flexible remote opportunity allows you to work on your own schedule while...Remote jobPart timeFlexible hours$60 - $75 per hour
...Applied Biology Benchmark Specialist - AI Evaluation is a remote review track for evaluating AI outputs across biology reasoning, calculations, and research workflows. Reviewers grade derivations and assumptions, reproduce key results, and document the correct method...Remote jobFor contractors10 hours per week$1,750 - $2,150 per month
...connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our... ...,750–$2,150 per completed task Location: Remote Role Responsibilities Review and evaluate AI-generated outputs related to threat analysis,...Remote workHourly payFull timeContract workSummer work$152k - $240k
...control spend effortlessly. Brex’s AI-native automation and world-... ...’ll do We're building AI agents to automate and augment... ...four weeks per year of fully remote work! Responsibilities:... ..., and data sources. Define evaluation frameworks, success metrics, and...Remote workFull timeWork at officeWork from home$200k - $320k
...technical depth and a passion for AI-driven product development.... .... This is a full-time remote opportunity, with preference for... ...production-ready LLM pipelines and AI agent systems Develop AI-driven... ...new product initiatives Evaluate and recommend AI architectures...Remote workFull timeVisa sponsorship$202.5k - $247.5k
...sharing localhost or running AI workloads in production. We... ...worth your time. About the Agent Team Our Agent team... ...AWS. Engineers develop by using remote development tools and/or ssh to... ...and actual compensation will be evaluated based on factors including,...Remote workPermanent employmentFull timeWork at officeLocal areaImmediate startHome officeFlexible hours$186.1k - $292.81k
...needed to drive revenue performance. AI agents are becoming central to how CaptivateIQ... ...agent SDK, LLM orchestration layer, and the evaluation and observability infrastructure that... ...effectively with cross-functional partners in a remote-first environment. High sense of...Remote workFlexible hoursShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Remote AI Agent Evaluation Specialist. Be the first to apply!
- remote chat agent Providence, RI
- ai scientist Providence, RI
- senior devops engineer remote Providence, RI
- remote medical coding supervisor Providence, RI
- remote virtual Providence, RI
- remote contract attorney Providence, RI
- clinical data manager remote Providence, RI
- online remote Providence, RI
- remote work no experience Providence, RI
- remote accounting Providence, RI




