Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Remote AI Agent Evaluation Specialist

$80 per hour

Mindrift

Raleigh, NC
  • Remote job

A leading tech company is seeking contributors for a flexible part-time opportunity focused on quality assurance for autonomous AI agents. You will analyze complex systems, review tasks for logic, and ensure clear expected behaviors for AI. Ideal candidates possess excellent analytical and detail-oriented skills, with experience in policy evaluation or logic puzzles preferred. Compensation can reach up to $80/hour based on expertise and project needs. This role offers valuable experience in an advanced AI project and fits around your primary commitments. #J-18808-Ljbffr Mindrift

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Remote AI Agent Evaluation Specialist in Raleigh, NC vacancy
  •  ...technology company is looking for a detail-oriented individual to design structured evaluation scenarios for AI agents. This entry-level part-time role allows you to contribute remotely, creating test cases for LLM-based agents. The ideal candidate will have a background... 
    Remote work
    Part time

    Mind Rift

    United States
    3 days ago
  • $80 per hour

    A leading AI consultancy is seeking a detail-oriented individual to design evaluation scenarios for LLM-based agents. This part-time, remote role allows you to create structured test cases while working flexibly around your commitments. The ideal candidate holds a relevant... 
    Remote work
    Part time

    Mindrift

    Houston, TX
    4 days ago
  • $80 per hour

    A technology firm is seeking QAs for autonomous AI agents to validate and improve task structures within a new project. Candidates will...  ...thinkers with strong attention to detail and experience in evaluating scenarios. This flexible, project-based role offers competitive... 
    Remote job
    Flexible hours

    Mindrift

    Providence, RI
    4 days ago
  • $60 per hour

     ...contributors for a part-time QA project focused on autonomous AI agents. This flexible remote opportunity requires strong analytical and critical...  ...with structured data formats. Candidates will review evaluation tasks, identify inconsistencies, and help define expected... 
    Remote job
    Part time
    Flexible hours

    Mindrift

    Kansas City, MO
    4 days ago
  • $60 per hour

     ...A leading AI firm in Austin is looking for QA experts to validate...  ...and improve AI systems. This remote, freelance role requires...  ...detail. Candidates will review AI evaluation tasks, identify inconsistencies...  ...define expected behaviors for agents. Ideal applicants have experience... 
    Remote work
    Freelance

    Mind Rift

    United States
    3 days ago
  • $60 per hour

     ...ethically shape the future of AI. What We Do The Mindrift platform...  ...thrive in ambiguity, enjoy remote asynchronous work, and want to...  ...modern AI systems are tested and evaluated, we want to hear from you....  ...QA experts for autonomous AI agents in a project focused on validating... 
    Remote work
    Freelance
    Flexible hours

    Mindrift

    Austin, TX
    4 days ago
  •  ...Freelance · Remote · North America, LATAM, or India About Turing...  ...companies, working with frontier AI labs to accelerate model...  ...through high-quality training data, evaluations, and engineering talent....  ...applications, you’ll work with coding agents across real-world repositories... 
    Remote work
    Hourly pay
    Temporary work
    For contractors
    Freelance
    Immediate start

    Turing

    Remote
    21 days ago
  • $90 per hour

     ...creative and technical talent with leading AI research labs. Headquartered in San...  ...Compensation: $90/hour Location: Remote Role Responsibilities Render a reference...  ...RLHF , preference-labeling, model-evaluation, or structured code-review work. Pixel... 
    Remote work
    Contract work
    Summer work
    Local area

    Mercor

    San Francisco, CA
    2 days ago
  • $70 - $110 per hour

     ...Help evaluate and improve AI systems by applying practical project management expertise to AI-generated...  ...of experience as a Project Management Specialist, Project Manager, Technical Project...  ...communication skills. Work Terms Remote, hourly engagement. Immediate start... 
    Remote work
    Hourly pay
    Immediate start

    SaidGig

    United States
    5 days ago
  •  ...SWE Agent Evaluation Specialist is a remote engineering review track for evaluating production code, debugging traces, and developer-facing AI outputs against real-world correctness standards. Reviewers reproduce failures, write the unit test the model should have written... 
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    5 days ago
  •  ...Expert Codebase Evaluation Specialist - Coding Agent Review is a remote evaluation track for reviewing codebase evaluation evaluation prompts and responses against...  ...team can use to retrain. Why this role matters AI data reviewers help turn codebase evaluation... 
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    5 days ago
  • $80 - $100 per hour

     ...help train and improve next-generation AI coding agents. In this contractor role you will provide...  ...end feature development. Critically evaluate AI agents in diverse scenarios,...  ...Engagement type: Contractor. Location: Remote. Work will involve contributing to a... 
    Remote work
    Hourly pay
    For contractors

    SaidGig

    United States
    3 days ago
  • $40 - $50 per hour

     ...Overview Apply your linguistics expertise to evaluate large language model outputs and help train next-generation AI systems. You will analyze model-human conversations...  ...contractor relationship. Location: Remote, applicants must be located in the United States... 
    Remote job
    Hourly pay
    For contractors
    Immediate start

    SaidGig

    Europe
    2 days ago
  • $20 - $60 per hour

     ...Help train next-generation AI systems by creating rigorous, real-world evaluations that test how well advanced models learn, reason, and perform. This remote contract opportunity is open to recent graduates, advanced-degree holders, and professionals from any background... 
    Remote work
    Hourly pay
    Contract work
    For contractors

    SaidGig

    United States
    1 day ago
  • $70 - $110 per hour

     ...technical talent with leading AI research labs. Headquartered...  ...Position: Project Management Specialists Type: Contract...  ...70–$110/hour Location: Remote Duration: 4–6 weeks Commitment...  ...Role Responsibilities Evaluate AI-generated project plans... 
    Remote work
    Hourly pay
    Contract work
    Summer work
    Immediate start

    Mercor

    San Francisco, CA
    6 days ago
  •  ...(~15 hours/week) Location: Remote Job Summary We are seeking...  ...Software Engineers to support an AI training project by creating...  ...learning environments that evaluate AI models on complex software...  ...reference solutions. Evaluate AI agents' ability to reason through... 
    Remote job
    For contractors

    YO AI Labs

    Houston, TX
    5 days ago
  •  ...is seeking a Vietnamese Voice Acting Specialist for a freelance AI Trainer project. The role is critical...  ...emotional expression. This position is remote and designed for individuals with...  ...strong voice acting credentials. You will evaluate AI outputs and support the... 
    Remote work
    Hourly pay
    Freelance

    Meridial

    United States
    3 days ago
  • $80 - $120 per hour

     ...elite creative and technical talent with leading AI research labs. Headquartered in San Francisco...  .../ philanthropy / community programs Evaluator Type: Contract Compensation: $80–$120/hour Location: Remote Role Responsibilities Evaluate AI-generated... 
    Remote work
    Contract work
    Summer work
    Work at office

    Mercor

    Atlanta, GA
    24 days ago
  • $80 - $150 per hour

     ...elite creative and technical talent with leading AI research labs. Headquartered in San Francisco...  .... Position: MS Excel / Google Sheets Evaluator Type: Contract Compensation: $80–$150/hour Location: Remote Role Responsibilities Review and assess... 
    Remote work
    Contract work
    Summer work

    Mercor

    Miami, FL
    5 days ago
  • $70k - $140k

     ...Program Evaluation Specialist - Department of State The Position: We are seeking driven, thoughtful candidates with experience in program...  ..., including paid vacation and sick leave, flexible and remote work opportunities, and tuition and training reimbursement.... 
    Remote work
    Full time
    Work at office
    Work from home
    Flexible hours
    Night shift

    Censeo Consulting Group

    Washington DC
    1 day ago
  •  ...BairesDev is seeking a Principal AI Engineer to set the technical direction and own the architecture of a major enterprise AI assistant...  ...You will shape the system from the ground up, ensure robust evaluation, and drive best practices in CI/CD, observability, and scalable... 
    Remote job

    Jobleads-US

    Palo Alto, CA
    2 days ago
  •  ...consulting firm in Colby, Kansas, is seeking a Nurse Practitioner or Physician Assistant to conduct examinations for veteran compensation evaluations. Responsibilities include conducting assessments, reviewing medical histories, and completing reports, all while collaborating... 
    Remote job
    Work from home
    Flexible hours

    Rodgers and Rodgers Consulting

    Colby, KS
    2 days ago
  •  ...world's Top 1% of tech talent, works remotely on roles that drive significant impact...  ...exceptional career development and success. AI Engineer (Agents) at BairesDev As an AI Engineer...  ...systems by implementing rigorous evaluation and moderation guardrails. Collaborate... 
    Remote work
    Local area
    Work from home
    Worldwide
    Flexible hours

    BairesDev

    Palo Alto, CA
    2 days ago
  •  ...for the planning, development, implementation, coordination, and evaluation of community focused initiatives that advance prevalence,...  ...retirement savings Positions that are eligible for hybrid or mobile/remote work mode are at the discretion of the hiring department. Work... 
    Remote work

    University of Michigan Flint

    Ann Arbor, MI
    3 days ago
  • $70k - $80k

     ...Job purpose The Program Evaluation Specialist will help lead program impact evaluation efforts at AFT. The Specialist will help construct...  ...can be proud of Competitive compensation & benefits Remote work opportunity Timeline To be considered... 
    Remote work
    Full time
    Casual work
    Local area
    Flexible hours
    Weekend work
    Afternoon shift

    American Farmland Trust

    United States
    3 days ago
  • $125 per hour

     ...Based Engagement (Corp-to-Corp) | Remote with International Travel | $1...  ...that matter most — across AI and digital transformation strategy...  ...is engaging a Principal AI Agent Architect to lead the most...  ...recommendations, including comparative evaluations across the Enterprise tier (e.... 
    Remote work
    Hourly pay
    Temporary work

    Jobleads-US

    Los Angeles, CA
    1 day ago
  •  ...Emerging Technologies Specialist to join the Digital Humanities...  ...intelligence (AI), large language models...  ...commitment to exploring and evaluating new tools and methods....  ...AI tools, and coding agents such as Claude Code; awareness...  ...with up to two remote days per week.About UVAThe... 
    Remote work
    Local area
    Visa sponsorship
    2 days per week

    University of Virginia

    Charlottesville, VA
    1 day ago
  • Obsidian is hiring expert Evaluators in Special education/IEP to review AI-generated work products for accuracy and rigor. This is a remote, hourly engagement where you will apply your deep subject-matter expertise to grade outputs. Candidates must possess at least 5 years... 
    Remote work
    Hourly pay
    Work at office

    Obsidian

    United States
    5 days ago
  •  ...professional to join a GenAI red-team within a leading AI lab network. The role involves probing frontier models, designing robust evaluation tasks, and documenting findings for reproducibility. This full-time W-2 position is remote within the United States, requiring about 35... 
    Remote work
    Full time

    Mercor

    New York, NY
    4 days ago
  • $112.09k - $160.33k

     ...Location: Burlington, VT (hybrid) or remote About OhMD                       OhMD is an AI patient communication platform...  ...— and OhMD AI, our AI voice agent. It picks up on the first ring,...  ...choices you made Familiarity with evaluating agent quality in a structured... 
    Remote work
    Live in

    OhMD

    Burlington, VT
    29 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Remote AI Agent Evaluation Specialist. Be the first to apply!