Remote AI Agent Evaluation Specialist
$80 per hourMindrift
- Remote job
A leading tech company is seeking contributors for a flexible part-time opportunity focused on quality assurance for autonomous AI agents. You will analyze complex systems, review tasks for logic, and ensure clear expected behaviors for AI. Ideal candidates possess excellent analytical and detail-oriented skills, with experience in policy evaluation or logic puzzles preferred. Compensation can reach up to $80/hour based on expertise and project needs. This role offers valuable experience in an advanced AI project and fits around your primary commitments. #J-18808-Ljbffr Mindrift
$125 per hour
...Role Overview QGIS Specialists apply QGIS and spatial analysis expertise to evaluate AI-generated GIS outputs, test spatial data workflows, and provide detailed feedback... ...full-time employment position. Work is remote and asynchronous, so you can complete tasks independently...Remote workHourly payFull timeContract workPart timeFlexible hours$60 per hour
...contributors for a part-time QA project focused on autonomous AI agents. This flexible remote opportunity requires strong analytical and critical... ...with structured data formats. Candidates will review evaluation tasks, identify inconsistencies, and help define expected...Remote workPart timeFlexible hours$80 per hour
A technology firm is seeking QAs for autonomous AI agents to validate and improve task structures within a new project. Candidates will... ...thinkers with strong attention to detail and experience in evaluating scenarios. This flexible, project-based role offers competitive...Remote workFlexible hours$80 per hour
...-time opportunity focused on quality assurance for autonomous AI agents. You will analyze complex systems, review tasks for logic, and... ...analytical and detail-oriented skills, with experience in policy evaluation or logic puzzles preferred. Compensation can reach up to $80/...Remote workPart timeFlexible hours$80 - $110 per hour
...Overview Work on the forefront of generative AI by designing and executing real-world... ...executable tests where applicable, run evaluations against a target model, and analyze failures... ...with client teams. Fully remote work within the United States, with an expected...Remote workHourly payPart timeFreelance- ...technology company is looking for a detail-oriented individual to design structured evaluation scenarios for AI agents. This entry-level part-time role allows you to contribute remotely, creating test cases for LLM-based agents. The ideal candidate will have a background...Remote jobPart time
$80 per hour
A leading AI consultancy is seeking a detail-oriented individual to design evaluation scenarios for LLM-based agents. This part-time, remote role allows you to create structured test cases while working flexibly around your commitments. The ideal candidate holds a relevant...Remote workPart time$60 per hour
A leading AI firm in Austin is looking for QA experts to validate... ...and improve AI systems. This remote, freelance role requires... ...detail. Candidates will review AI evaluation tasks, identify inconsistencies... ...define expected behaviors for agents. Ideal applicants have experience...Remote jobFreelance- ...SWE Agent Evaluation Specialist is a remote engineering review track for evaluating production code, debugging traces, and developer-facing AI outputs against real-world correctness standards. Reviewers reproduce failures, write the unit test the model should have written...Remote jobHourly payFor contractors10 hours per week
$60 - $90 per hour
...researchers to convert findings into robust evaluation benchmarks. Key Responsibilities... ...research, research engineering, security, or AI evaluation. Proven ability to identify... ...or task-based gig. Position is fully remote within the United States. Typical engagement...Remote workHourly payFull timeFreelance- ...Expert Codebase Evaluation Specialist - Coding Agent Review is a remote evaluation track for reviewing codebase evaluation evaluation prompts and responses against... ...team can use to retrain. Why this role matters AI data reviewers help turn codebase evaluation...Remote jobHourly payFor contractors10 hours per week
$20 - $30 per hour
...Role Overview Evaluate images to help train next generation AI systems by providing high quality, real world assessments that shape how models learn and reason. This remote, contractor role focuses on domain knowledge, visual judgment, and clear written reasoning. No...Remote jobHourly payFor contractors$70 - $90 per hour
...content for security vulnerabilities to help AI models recognize and classify threats.... ...in English. Work Terms Location: Remote. Engagement type: Hourly contract.... ...a short interview and a questionnaire to evaluate domain expertise. If hired, onboarding...Remote workHourly payContract workTemporary work$50 per hour
...Role Overview Math Specialists apply advanced mathematical training to evaluate AI-generated mathematics content, design domain-relevant questions, and give detailed... ...-based contract role that can be performed remotely alongside research, teaching, coursework, or industry...Remote workHourly payFull timeContract workPart timeFor contractorsFlexible hours- ...Excel Specialist - AI Workflow Evaluator is a remote review track for evaluating AI outputs across excel specialist operations workflows. Reviewers grade workflow correctness, policy adherence, and stakeholder fit; flag operational risk; and document the right next step...Remote jobHourly payFor contractorsWork experience placement10 hours per week
$85 per hour
...estimation, construction planning, and solar design experience to evaluate AI-generated content and create expert-level training data. In... ..., and ability to work asynchronously and independently with remote research teams. No prior AI experience required. Work Terms...Remote workHourly payContract workPart timeFlexible hours$1,750 - $2,150 per month
...connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our... ...,750–$2,150 per completed task Location: Remote Role Responsibilities Review and evaluate AI-generated outputs related to threat analysis,...Remote workHourly payFull timeContract workSummer work$10 - $20 per hour
...quality language data that trains and improves AI systems. This contractor role focuses on... ...comfortable working independently in a remote environment. Preferred Qualifications... ...data lab that creates training data and evaluations for frontier AI models. Experts...Remote workHourly payFor contractors$125 per hour
...Role Overview CAD Tool Specialist applies CAD software expertise to AI research projects, providing domain knowledge... ...design workflows. You will work remotely and asynchronously on short-term,... ...for CAD workflows and tooling. Evaluate and score large language model responses...Remote workHourly payFull timeTemporary workPart timeWork experience placementFlexible hours$80 - $120 per hour
...elite creative and technical talent with leading AI research labs. Headquartered in San Francisco... .... Position: Process improvement / SOPs Evaluator Type: Contract Compensation: $80–$120/hour Location: Remote Role Responsibilities Evaluate AI-...Remote workContract workSummer workWork at office$152k - $240k
...control spend effortlessly. Brex’s AI-native automation and world-... ...’ll do We're building AI agents to automate and augment... ...four weeks per year of fully remote work! Responsibilities:... ..., and data sources. Define evaluation frameworks, success metrics, and...Remote workFull timeWork at officeWork from home- ...is seeking a Vietnamese Voice Acting Specialist for a freelance AI Trainer project. The role is critical... ...emotional expression. This position is remote and designed for individuals with... ...strong voice acting credentials. You will evaluate AI outputs and support the...Remote workHourly payFreelance
$25 - $30 per hour
...DataAnnotation is seeking a Certified Coding Specialist (CCS) to join their team and help train AI models. This role requires expertise in healthcare to evaluate AI performance and improve model quality. Applicants should have fluency in English and a medical or healthcare...Remote workHourly payFor contractors$15 - $25 per hour
...Role Overview Apply your clinical knowledge to evaluate and structure healthcare information that helps train next-generation AI systems. This part-time contractor role... ...Ability to multitask and work efficiently in a remote setting, with strong analytical and organizational...Remote workHourly payPart timeFor contractors$200k - $320k
...technical depth and a passion for AI-driven product development.... .... This is a full-time remote opportunity, with preference for... ...production-ready LLM pipelines and AI agent systems Develop AI-driven... ...new product initiatives Evaluate and recommend AI architectures...Remote workFull timeVisa sponsorship$70 - $90 per hour
...that help train next-generation AI systems for a leading customer... ...and research domain. This remote contractor role focuses on converting... ...integrity. Critically evaluate experimental outcomes and propose... ...information for non-specialist audiences. Contribute domain...Remote jobHourly payFor contractors$202.5k - $247.5k
...sharing localhost or running AI workloads in production. We... ...worth your time. About the Agent Team Our Agent team... ...AWS. Engineers develop by using remote development tools and/or ssh to... ...and actual compensation will be evaluated based on factors including,...Remote workPermanent employmentFull timeWork at officeLocal areaImmediate startHome officeFlexible hours$100k - $115k
...is seeking a detail-oriented Program Evaluation Specialist to support program evaluation, implementation... ...-grade platforms and mission-ready AI to federal agencies at commercial speed... ...security clearances, due to the nature of the work. Job Locations US-RemoteRemote workContract work$60 - $90 per hour
...creative and technical talent with leading AI research labs. Headquartered in San... ...Compensation: $60–$90/hour Location: Remote Commitment: 15–40 hours/week... ...specialized cybersecurity topics. Evaluate and annotate model responses for technical...Remote workFull timeContract workSummer workImmediate start$80 - $150 per hour
...Prolific is seeking Medical Doctors to join as Domain Expert participants in training AI models. Responsibilities include evaluating AI responses for accuracy and writing feedback for improvement. A competitive pay of $80-$150 per hour based on skills is offered. Candidates...Remote workHourly payWork from home
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Remote AI Agent Evaluation Specialist. Be the first to apply!



