Remote AI Agent Evaluation Specialist
$80 per hourMindrift
- Remote job
A leading tech company is seeking contributors for a flexible part-time opportunity focused on quality assurance for autonomous AI agents. You will analyze complex systems, review tasks for logic, and ensure clear expected behaviors for AI. Ideal candidates possess excellent analytical and detail-oriented skills, with experience in policy evaluation or logic puzzles preferred. Compensation can reach up to $80/hour based on expertise and project needs. This role offers valuable experience in an advanced AI project and fits around your primary commitments. #J-18808-Ljbffr Mindrift
- ...technology company is looking for a detail-oriented individual to design structured evaluation scenarios for AI agents. This entry-level part-time role allows you to contribute remotely, creating test cases for LLM-based agents. The ideal candidate will have a background...Remote workPart time
$80 per hour
A leading AI consultancy is seeking a detail-oriented individual to design evaluation scenarios for LLM-based agents. This part-time, remote role allows you to create structured test cases while working flexibly around your commitments. The ideal candidate holds a relevant...Remote workPart time$80 per hour
A technology firm is seeking QAs for autonomous AI agents to validate and improve task structures within a new project. Candidates will... ...thinkers with strong attention to detail and experience in evaluating scenarios. This flexible, project-based role offers competitive...Remote jobFlexible hours$60 per hour
...contributors for a part-time QA project focused on autonomous AI agents. This flexible remote opportunity requires strong analytical and critical... ...with structured data formats. Candidates will review evaluation tasks, identify inconsistencies, and help define expected...Remote jobPart timeFlexible hours$60 per hour
...A leading AI firm in Austin is looking for QA experts to validate... ...and improve AI systems. This remote, freelance role requires... ...detail. Candidates will review AI evaluation tasks, identify inconsistencies... ...define expected behaviors for agents. Ideal applicants have experience...Remote workFreelance$60 per hour
...ethically shape the future of AI. What We Do The Mindrift platform... ...thrive in ambiguity, enjoy remote asynchronous work, and want to... ...modern AI systems are tested and evaluated, we want to hear from you.... ...QA experts for autonomous AI agents in a project focused on validating...Remote workFreelanceFlexible hours- ...Freelance · Remote · North America, LATAM, or India About Turing... ...companies, working with frontier AI labs to accelerate model... ...through high-quality training data, evaluations, and engineering talent.... ...applications, you’ll work with coding agents across real-world repositories...Remote workHourly payTemporary workFor contractorsFreelanceImmediate start
$90 per hour
...creative and technical talent with leading AI research labs. Headquartered in San... ...Compensation: $90/hour Location: Remote Role Responsibilities Render a reference... ...RLHF , preference-labeling, model-evaluation, or structured code-review work. Pixel...Remote workContract workSummer workLocal area$70 - $110 per hour
...Help evaluate and improve AI systems by applying practical project management expertise to AI-generated... ...of experience as a Project Management Specialist, Project Manager, Technical Project... ...communication skills. Work Terms Remote, hourly engagement. Immediate start...Remote workHourly payImmediate start- ...SWE Agent Evaluation Specialist is a remote engineering review track for evaluating production code, debugging traces, and developer-facing AI outputs against real-world correctness standards. Reviewers reproduce failures, write the unit test the model should have written...Remote jobHourly payFor contractors10 hours per week
- ...Expert Codebase Evaluation Specialist - Coding Agent Review is a remote evaluation track for reviewing codebase evaluation evaluation prompts and responses against... ...team can use to retrain. Why this role matters AI data reviewers help turn codebase evaluation...Remote jobHourly payFor contractors10 hours per week
$80 - $100 per hour
...help train and improve next-generation AI coding agents. In this contractor role you will provide... ...end feature development. Critically evaluate AI agents in diverse scenarios,... ...Engagement type: Contractor. Location: Remote. Work will involve contributing to a...Remote workHourly payFor contractors$40 - $50 per hour
...Overview Apply your linguistics expertise to evaluate large language model outputs and help train next-generation AI systems. You will analyze model-human conversations... ...contractor relationship. Location: Remote, applicants must be located in the United States...Remote jobHourly payFor contractorsImmediate start$20 - $60 per hour
...Help train next-generation AI systems by creating rigorous, real-world evaluations that test how well advanced models learn, reason, and perform. This remote contract opportunity is open to recent graduates, advanced-degree holders, and professionals from any background...Remote workHourly payContract workFor contractors$70 - $110 per hour
...technical talent with leading AI research labs. Headquartered... ...Position: Project Management Specialists Type: Contract... ...70–$110/hour Location: Remote Duration: 4–6 weeks Commitment... ...Role Responsibilities Evaluate AI-generated project plans...Remote workHourly payContract workSummer workImmediate start- ...(~15 hours/week) Location: Remote Job Summary We are seeking... ...Software Engineers to support an AI training project by creating... ...learning environments that evaluate AI models on complex software... ...reference solutions. Evaluate AI agents' ability to reason through...Remote jobFor contractors
- ...is seeking a Vietnamese Voice Acting Specialist for a freelance AI Trainer project. The role is critical... ...emotional expression. This position is remote and designed for individuals with... ...strong voice acting credentials. You will evaluate AI outputs and support the...Remote workHourly payFreelance
$80 - $120 per hour
...elite creative and technical talent with leading AI research labs. Headquartered in San Francisco... .../ philanthropy / community programs Evaluator Type: Contract Compensation: $80–$120/hour Location: Remote Role Responsibilities Evaluate AI-generated...Remote workContract workSummer workWork at office$80 - $150 per hour
...elite creative and technical talent with leading AI research labs. Headquartered in San Francisco... .... Position: MS Excel / Google Sheets Evaluator Type: Contract Compensation: $80–$150/hour Location: Remote Role Responsibilities Review and assess...Remote workContract workSummer work$70k - $140k
...Program Evaluation Specialist - Department of State The Position: We are seeking driven, thoughtful candidates with experience in program... ..., including paid vacation and sick leave, flexible and remote work opportunities, and tuition and training reimbursement....Remote workFull timeWork at officeWork from homeFlexible hoursNight shift- ...BairesDev is seeking a Principal AI Engineer to set the technical direction and own the architecture of a major enterprise AI assistant... ...You will shape the system from the ground up, ensure robust evaluation, and drive best practices in CI/CD, observability, and scalable...Remote job
- ...consulting firm in Colby, Kansas, is seeking a Nurse Practitioner or Physician Assistant to conduct examinations for veteran compensation evaluations. Responsibilities include conducting assessments, reviewing medical histories, and completing reports, all while collaborating...Remote jobWork from homeFlexible hours
- ...world's Top 1% of tech talent, works remotely on roles that drive significant impact... ...exceptional career development and success. AI Engineer (Agents) at BairesDev As an AI Engineer... ...systems by implementing rigorous evaluation and moderation guardrails. Collaborate...Remote workLocal areaWork from homeWorldwideFlexible hours
- ...for the planning, development, implementation, coordination, and evaluation of community focused initiatives that advance prevalence,... ...retirement savings Positions that are eligible for hybrid or mobile/remote work mode are at the discretion of the hiring department. Work...Remote work
$70k - $80k
...Job purpose The Program Evaluation Specialist will help lead program impact evaluation efforts at AFT. The Specialist will help construct... ...can be proud of Competitive compensation & benefits Remote work opportunity Timeline To be considered...Remote workFull timeCasual workLocal areaFlexible hoursWeekend workAfternoon shift$125 per hour
...Based Engagement (Corp-to-Corp) | Remote with International Travel | $1... ...that matter most — across AI and digital transformation strategy... ...is engaging a Principal AI Agent Architect to lead the most... ...recommendations, including comparative evaluations across the Enterprise tier (e....Remote workHourly payTemporary work- ...Emerging Technologies Specialist to join the Digital Humanities... ...intelligence (AI), large language models... ...commitment to exploring and evaluating new tools and methods.... ...AI tools, and coding agents such as Claude Code; awareness... ...with up to two remote days per week.About UVAThe...Remote workLocal areaVisa sponsorship2 days per week
- Obsidian is hiring expert Evaluators in Special education/IEP to review AI-generated work products for accuracy and rigor. This is a remote, hourly engagement where you will apply your deep subject-matter expertise to grade outputs. Candidates must possess at least 5 years...Remote workHourly payWork at office
- ...professional to join a GenAI red-team within a leading AI lab network. The role involves probing frontier models, designing robust evaluation tasks, and documenting findings for reproducibility. This full-time W-2 position is remote within the United States, requiring about 35...Remote workFull time
$112.09k - $160.33k
...Location: Burlington, VT (hybrid) or remote About OhMD OhMD is an AI patient communication platform... ...— and OhMD AI, our AI voice agent. It picks up on the first ring,... ...choices you made Familiarity with evaluating agent quality in a structured...Remote workLive in
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Remote AI Agent Evaluation Specialist. Be the first to apply!





