Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Agent Evaluation Engineer - Remote

YO AI Labs

New York, NY
  • Remote job

Job Description

Job Description

Senior Software Engineer

Job Type: Contractor (~15 hours/week)
Location: Remote

Job Summary

We are seeking experienced Senior Software Engineers to support an AI training project by creating reinforcement learning environments that evaluate AI models on complex software engineering tasks using Model Context Protocol (MCP) tools.

You will design reproducible environments, deterministic verification, and reference solutions for tasks such as bug fixing, feature implementation, codebase refactoring, and performance optimization . No prior AI experience is required.

Key Responsibilities
  • Create reinforcement learning environments for software engineering tasks.
  • Design tasks involving bug fixing, feature development, refactoring, and performance optimization .
  • Build deterministic verification systems and golden reference solutions.
  • Evaluate AI agents' ability to reason through complex codebases and use MCP tools effectively.
  • Develop realistic, reproducible software engineering scenarios.
  • Ensure tasks accurately measure coding ability, problem-solving, and tool usage.
  • Document solutions and provide clear technical feedback.
Required Skills
  • Strong proficiency in Python 3, Java, Rust, C++, or TypeScript .
  • Strong understanding of algorithms and data structures .
  • Experience with bug fixing and debugging complex software issues .
  • Proven experience in feature implementation and codebase refactoring .
  • Strong knowledge of performance optimization and tuning .
  • Excellent written and verbal communication.
  • Strong attention to detail.
Preferred Qualifications
  • Experience working with large-scale or distributed codebases .
  • Familiarity with AI/ML systems is a plus but not required.
  • Experience with rigorous code reviews and software engineering best practices .
  • Experience working effectively in remote or cross-functional teams.
Hiring Process
  1. Submit an application and screening questions.
  2. Complete an AI interview (~30 minutes).
  3. Complete a technical assessment, if required.
  4. Hiring Manager review.
Compensation

Compensation is output-based , with payment provided per task that meets project specifications. Minimum weekly submission requirements may apply.

Availability

Selected experts should be prepared to begin their first tasks within 24–48 hours of completing onboarding .

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the AI Agent Evaluation Engineer - Remote in New York, NY vacancy
  • $80 per hour

     ...ethically shape the future of AI. What We Do The Mindrift...  ...Calling all security researchers, engineers, and penetration testers with...  ...tools for running and evaluating agent behavior. You’ll implement base...  ...needs Take part in a flexible, remote, freelance project that fits... 
    Remote work
    Hourly pay
    Part time
    Freelance
    Flexible hours

    Mindrift

    Kansas City, MO
    1 day ago
  • $80 per hour

     ...-time opportunity focused on quality assurance for autonomous AI agents. You will analyze complex systems, review tasks for logic, and...  ...analytical and detail-oriented skills, with experience in policy evaluation or logic puzzles preferred. Compensation can reach up to $80/... 
    Remote work
    Part time
    Flexible hours

    Mind Rift

    United States
    4 days ago
  • $60 per hour

     ...contributors for a part-time QA project focused on autonomous AI agents. This flexible remote opportunity requires strong analytical and critical...  ...with structured data formats. Candidates will review evaluation tasks, identify inconsistencies, and help define expected... 
    Remote work
    Part time
    Flexible hours

    Mind Rift

    Kansas City, MO
    4 days ago
  •  ...help design, develop, and operate MaaS platforms using LLMs and multimodal models, spanning training, inference, prompt engineering, and evaluation pipelines. The team builds end-to-end solutions, analyzes large-scale log data, and collaborates across research, platform... 
    Suggested

    ByteDance

    Seattle, WA
    3 days ago
  • $80 per hour

    A leading AI consultancy is seeking a detail-oriented individual to design evaluation scenarios for LLM-based agents. This part-time, remote role allows you to create structured test cases while working flexibly around your commitments. The ideal candidate holds a relevant... 
    Remote work
    Part time

    Mindrift

    Houston, TX
    1 day ago
  • $60 per hour

     ...innovation company is seeking QAs for autonomous AI agents to validate and improve task structures and evaluate logic. The role requires excellent analytical thinking...  .... Successful candidates can work flexibly and remotely, earning rates up to $60/hour. This position is... 
    Remote job
    Flexible hours

    Mindrift

    Raleigh, NC
    1 day ago
  • $80 per hour

    A technology firm is seeking QAs for autonomous AI agents to validate and improve task structures within a new project. Candidates will...  ...thinkers with strong attention to detail and experience in evaluating scenarios. This flexible, project-based role offers competitive... 
    Remote job
    Flexible hours

    Mindrift

    Providence, RI
    1 day ago
  • $140 - $190 per hour

     ...safety data while helping assess AI-authored pharmacovigilance...  ...Reports and Periodic Benefit Risk Evaluation Reports, including their...  ...evaluation rubrics to measure an AI agent''s ability to produce...  ...PBRER reporting. Work Terms Remote, hourly engagement. Applicants... 
    Remote work
    Hourly pay

    SaidGig

    Canada
    8 days ago
  •  ...technology company is looking for a detail-oriented individual to design structured evaluation scenarios for AI agents. This entry-level part-time role allows you to contribute remotely, creating test cases for LLM-based agents. The ideal candidate will have a background... 
    Remote job
    Part time

    Mindrift

    Brooklyn, NY
    2 days ago
  • $60 per hour

     ...A leading AI firm in Austin is looking for QA experts to validate...  ...and improve AI systems. This remote, freelance role requires...  ...detail. Candidates will review AI evaluation tasks, identify inconsistencies...  ...define expected behaviors for agents. Ideal applicants have experience... 
    Remote work
    Freelance

    Mind Rift

    Austin, TX
    3 days ago
  •  ...security enablement with agentic AI. Our platform tightens the...  ...for a talented Product Engineer to join the Systems Agents engineering team. This...  ..., and have strong evaluations. This role is ideal for...  ...and vision insurance. ~Remote friendly with WeWork access... 
    Remote work
    Full time

    Vannevar

    Remote
    a month ago
  • $60 per hour

     ...ethically shape the future of AI. What We Do The...  ...thrive in ambiguity, enjoy remote asynchronous work, and...  ...systems are tested and evaluated, we want to hear from...  ...experts for autonomous AI agents in a project focused on...  ...to LLMs, prompt engineering, or AI‑generated content... 
    Remote work
    Freelance
    Flexible hours

    Mindrift

    Austin, TX
    1 day ago
  • $175k - $287k

     ...the business needs of the team. LinkedIn’s Core AI is building the Evaluation Operating System (EOS), a foundational Agent Evaluation platform that defines how all AI...  ...infrastructure for all LinkedIn AI Agents.As a Staff Engineer, you will own the end-to-end technical vision,... 
    For contractors
    Work experience placement
    Work at office
    Flexible hours

    Linkedin

    Mountain View, CA
    3 days ago
  • $200k - $275k

     ...Secure Every Identity, from AI to Human Identity is the key...  ...let's talk. About Okta for AI Agents Okta secures access for 20,...  ...identity, platform, and security engineering teams, write production code...  ...outcomes end to end. #LI-Remote P25257_3456702 Below is... 
    Remote work
    Local area
    Worldwide
    Flexible hours
    Shift work

    Okta, Inc.

    Chicago, IL
    5 days ago
  •  ...leading security-first enterprise AI company, is seeking a Member...  ...Staff in the Safety for Agents team to advance safer LLMs through...  ...post-training algorithms, and evaluation methods. You'll collaborate with...  ...across global offices, while remote-friendly policies support... 
    Remote work
    Flexible hours

    Cohere

    New York, NY
    2 days ago
  • $80 per hour

    A tech company specializing in AI is seeking QAs for autonomous AI agents to validate and enhance task structures...  ...allows contributors to work remotely while engaging in a complex AI project...  ...detail, facilitating AI testing and evaluation without needing a coding... 
    Remote job
    Flexible hours

    Mind Rift

    Houston, TX
    3 days ago
  • $55 per hour

    A leading AI innovation firm is seeking QAs for autonomous AI agents to improve evaluation frameworks. Candidates should possess excellent analytical thinking and attention...  ...tasks and define clear standards. This remote, flexible opportunity offers rates up to $55/... 
    Remote job
    Part time
    Flexible hours

    Mindrift

    Raleigh, NC
    1 day ago
  • $144.7k - $261.3k

    Job DescriptionAbout the TeamThe Evaluation Foundations team, part of Embodied AI's Scaling Foundations, solves critical...  ...vehicle development. We engineer high-performance tools that identify...  ...more.#GM-AV-1 This role is based remotely, but if the selected candidate... 
    Remote work
    Full time
    Local area
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Austin, TX
    1 day ago
  • $55 per hour

    A leading AI innovation firm in Dallas is seeking QAs for autonomous AI agents to ensure the quality of complex systems and scenarios. This flexible, remote project is ideal for those with excellent analytical...  ...should be adept at evaluating scenarios and documenting findings... 
    Remote job
    Flexible hours

    Mindrift

    Dallas, TX
    5 days ago
  • YO AI Labs is seeking experienced civil engineers to contribute to an AI evaluation project. You will apply engineering expertise to create realistic scenarios, develop solutions...  ...on code compliance and technical accuracy. Remote-friendly engagement with quick onboarding and... 
    Remote job

    YO AI Labs

    Raleigh, NC
    3 days ago
  • YO AI Labs is seeking experienced Civil Engineers for a remote, output-based evaluation project focused on AI-driven engineering tasks. You will apply engineering expertise to create realistic scenarios, develop accurate solutions, and assess AI-generated outputs. No prior... 
    Remote job

    YO AI Labs

    Los Angeles, CA
    3 days ago
  • $80 per hour

     ...ethically shape the future of AI. What We Do The Mindrift...  ...realistic and structured evaluation scenarios for LLM-based agents. You'll create test cases...  ...Science, Software Engineering, Data Science / Data Analytics...  ...Take part in a flexible, remote, freelance project that... 
    Remote work
    Part time
    Freelance
    Flexible hours

    Mindrift

    New York, NY
    1 day ago
  •  ...Cohere's cutting-edge AI workspace platform, designed...  ...that connects AI agents with workplace tools and...  ...You’ll work with product engineering, modelling and design...  ...users create, configure, evaluate, and improve workflows...  ...other offices if you are remote, plus an annual company... 
    Remote work
    Full time
    Local area

    Cohere

    Remote
    23 days ago
  • YO AI Labs seeks experienced Civil Engineers for a remote contractor role contributing to AI evaluation. You will craft realistic engineering tasks, generate supporting drawings and reports, and assess code compliance across structural, transportation, and water systems... 
    Remote job
    For contractors

    YO AI Labs

    Dallas, TX
    3 days ago
  • YO AI Labs is seeking experienced Civil Engineers to contribute to an AI evaluation project. You will apply your engineering expertise to create realistic technical scenarios, develop accurate solutions, and evaluate AI-generated outputs. No prior AI experience is required... 
    Remote job

    YO AI Labs

    Atlanta, GA
    3 days ago
  • YO AI Labs seeks an experienced Civil Engineer to contribute to AI evaluation projects by creating realistic civil tasks, applying structural design and regulatory codes, and developing precise solutions. As a contractor, you will produce engineering drawings, reports,... 
    Remote job
    For contractors

    YO AI Labs

    Miami, FL
    2 days ago
  • YO AI Labs in the United States is seeking experienced Civil Engineers to contribute to a high-impact AI evaluation project. You will apply your engineering expertise to create realistic scenarios, develop accurate solutions, and evaluate AI-generated engineering outputs... 
    Remote job
    For contractors

    YO AI Labs

    Annapolis, MD
    3 days ago
  • YO AI Labs is seeking experienced Civil Engineers to contribute to a high-impact AI evaluation project. You will apply engineering expertise to create realistic tasks, develop accurate solutions, and evaluate AI-generated outputs. No prior AI experience is required. Responsibilities... 
    Remote job

    YO AI Labs

    Washington DC
    3 days ago
  • $100 per hour

     ...Agent Engineer — Tool-Use and Multi-Step Workflow Evaluation is a remote engineering review track for evaluating production code, debugging traces, and developer-facing AI outputs against real-world correctness standards. Reviewers reproduce failures, write the unit test... 
    Remote job
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    16 days ago
  •  ...Tyndale is expanding our AI and automation capabilities...  ...seeking an experienced AI Agent & Automation Engineer to build production solutions...  ...support.   HYBRID/REMOTE: Tyndale supports a strong...  ...platforms, AI agent testing/evaluation, containerization and operational... 
    Remote work
    Casual work
    1 day per week

    Tyndale

    Pipersville, PA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Agent Evaluation Engineer - Remote. Be the first to apply!