Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Agent Evaluation Engineer - Remote

YO AI Labs

Austin, TX
  • Remote job

Job Description

Job Description

Senior Software Engineer

Job Type: Contractor (~15 hours/week)
Location: Remote

Job Summary

We are seeking experienced Senior Software Engineers to support an AI training project by creating reinforcement learning environments that evaluate AI models on complex software engineering tasks using Model Context Protocol (MCP) tools.

You will design reproducible environments, deterministic verification, and reference solutions for tasks such as bug fixing, feature implementation, codebase refactoring, and performance optimization . No prior AI experience is required.

Key Responsibilities
  • Create reinforcement learning environments for software engineering tasks.
  • Design tasks involving bug fixing, feature development, refactoring, and performance optimization .
  • Build deterministic verification systems and golden reference solutions.
  • Evaluate AI agents' ability to reason through complex codebases and use MCP tools effectively.
  • Develop realistic, reproducible software engineering scenarios.
  • Ensure tasks accurately measure coding ability, problem-solving, and tool usage.
  • Document solutions and provide clear technical feedback.
Required Skills
  • Strong proficiency in Python 3, Java, Rust, C++, or TypeScript .
  • Strong understanding of algorithms and data structures .
  • Experience with bug fixing and debugging complex software issues .
  • Proven experience in feature implementation and codebase refactoring .
  • Strong knowledge of performance optimization and tuning .
  • Excellent written and verbal communication.
  • Strong attention to detail.
Preferred Qualifications
  • Experience working with large-scale or distributed codebases .
  • Familiarity with AI/ML systems is a plus but not required.
  • Experience with rigorous code reviews and software engineering best practices .
  • Experience working effectively in remote or cross-functional teams.
Hiring Process
  1. Submit an application and screening questions.
  2. Complete an AI interview (~30 minutes).
  3. Complete a technical assessment, if required.
  4. Hiring Manager review.
Compensation

Compensation is output-based , with payment provided per task that meets project specifications. Minimum weekly submission requirements may apply.

Availability

Selected experts should be prepared to begin their first tasks within 24–48 hours of completing onboarding .

Vacancy posted 11 days ago
Similar jobs that could be interesting for youBased on the AI Agent Evaluation Engineer - Remote in Austin, TX vacancy
  • $80 per hour

    Get AI‑powered advice on this job and more exclusive features...  ...Calling all security researchers, engineers, and penetration testers with...  ...tools for running and evaluating agent behavior. You’ll implement base...  ...needs Take part in a flexible, remote, freelance project that fits... 
    Remote work
    Freelance
    Flexible hours

    Mind Rift

    Dallas, TX
    3 days ago
  •  ...technology company is looking for a detail-oriented individual to design structured evaluation scenarios for AI agents. This entry-level part-time role allows you to contribute remotely, creating test cases for LLM-based agents. The ideal candidate will have a background... 
    Remote work
    Part time

    Mind Rift

    United States
    2 days ago
  •  ...help design, develop, and operate MaaS platforms using LLMs and multimodal models, spanning training, inference, prompt engineering, and evaluation pipelines. The team builds end-to-end solutions, analyzes large-scale log data, and collaborates across research, platform... 
    Suggested

    ByteDance

    Seattle, WA
    23 hours ago
  • $80 per hour

     ...-time opportunity focused on quality assurance for autonomous AI agents. You will analyze complex systems, review tasks for logic, and...  ...analytical and detail-oriented skills, with experience in policy evaluation or logic puzzles preferred. Compensation can reach up to $80/... 
    Remote work
    Part time
    Flexible hours

    Mind Rift

    United States
    2 days ago
  • $60 per hour

     ...innovation company is seeking QAs for autonomous AI agents to validate and improve task structures and evaluate logic. The role requires excellent analytical thinking...  .... Successful candidates can work flexibly and remotely, earning rates up to $60/hour. This position is... 
    Remote job
    Flexible hours

    Mindrift

    Raleigh, NC
    3 days ago
  • $80 per hour

    A leading AI consultancy is seeking a detail-oriented individual to design evaluation scenarios for LLM-based agents. This part-time, remote role allows you to create structured test cases while working flexibly around your commitments. The ideal candidate holds a relevant... 
    Remote work
    Part time

    Mindrift

    Houston, TX
    3 days ago
  •  ...leading security-first enterprise AI company, is seeking a Member...  ...Staff in the Safety for Agents team to advance safer LLMs through...  ...post-training algorithms, and evaluation methods. You'll collaborate with...  ...across global offices, while remote-friendly policies support... 
    Remote work
    Flexible hours

    Cohere

    New York, NY
    4 days ago
  • $80 per hour

    A technology firm is seeking QAs for autonomous AI agents to validate and improve task structures within a new project. Candidates will...  ...thinkers with strong attention to detail and experience in evaluating scenarios. This flexible, project-based role offers competitive... 
    Remote job
    Flexible hours

    Mindrift

    Providence, RI
    3 days ago
  • $60 per hour

     ...contributors for a part-time QA project focused on autonomous AI agents. This flexible remote opportunity requires strong analytical and critical...  ...with structured data formats. Candidates will review evaluation tasks, identify inconsistencies, and help define expected... 
    Remote job
    Part time
    Flexible hours

    Mindrift

    Kansas City, MO
    3 days ago
  • $55 per hour

     ...first 25 applicants Get AI-powered advice on this...  ...complexity? Does an async, remote, flexible opportunity...  ...systems are tested and evaluated? This is a flexible,...  ...QAs for autonomous AI agents for a new project focused...  ...Exposure to LLMs, prompt engineering, or AI‑generated... 
    Remote work
    Permanent employment
    Part time
    Freelance
    Flexible hours

    Mindrift

    Dallas, TX
    3 days ago
  • $144.7k - $261.3k

    Job DescriptionAbout the TeamThe Evaluation Foundations team, part of Embodied AI's Scaling Foundations, solves critical...  ...vehicle development. We engineer high-performance tools that identify...  ...more.#GM-AV-1 This role is based remotely, but if the selected candidate... 
    Remote work
    Full time
    Local area
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Warren, MI
    3 days ago
  • $175k - $287k

     ...the business needs of the team. LinkedIn’s Core AI is building the Evaluation Operating System (EOS), a foundational Agent Evaluation platform that defines how all AI...  ...infrastructure for all LinkedIn AI Agents.As a Staff Engineer, you will own the end-to-end technical vision,... 
    For contractors
    Work experience placement
    Work at office
    Flexible hours

    Linkedin

    Mountain View, CA
    3 days ago
  • YO AI Labs seeks experienced Civil Engineers for a remote contractor role contributing to AI evaluation. You will craft realistic engineering tasks, generate supporting drawings and reports, and assess code compliance across structural, transportation, and water systems... 
    Remote job
    For contractors

    YO AI Labs

    Dallas, TX
    23 hours ago
  • YO AI Labs seeks an experienced Civil Engineer to contribute to AI evaluation projects by creating realistic civil tasks, applying structural design and regulatory codes, and developing precise solutions. As a contractor, you will produce engineering drawings, reports,... 
    Remote job
    For contractors

    YO AI Labs

    Miami, FL
    23 hours ago
  • YO AI Labs is seeking experienced Civil Engineers for a remote, output-based evaluation project focused on AI-driven engineering tasks. You will apply engineering expertise to create realistic scenarios, develop accurate solutions, and assess AI-generated outputs. No prior... 
    Remote job

    YO AI Labs

    Los Angeles, CA
    23 hours ago
  • YO AI Labs is seeking experienced civil engineers to contribute to an AI evaluation project. You will apply engineering expertise to create realistic scenarios, develop solutions...  ...on code compliance and technical accuracy. Remote-friendly engagement with quick onboarding and... 
    Remote job

    YO AI Labs

    Raleigh, NC
    23 hours ago
  • $55 per hour

    A leading AI innovation firm is seeking QAs for autonomous AI agents to improve evaluation frameworks. Candidates should possess excellent analytical thinking and attention...  ...tasks and define clear standards. This remote, flexible opportunity offers rates up to $55/... 
    Remote job
    Part time
    Flexible hours

    Mindrift

    Raleigh, NC
    3 days ago
  •  ...in the United States seeks experienced Mechanical Engineering professionals to contribute to AI evaluation and training projects. You will shape enterprise-grade...  ..., and regulatory processes in Fortune 500 environments. Remote, flexible schedule. #J-18808-Ljbffr Synthires
    Remote job
    Flexible hours

    Synthires

    San Francisco, CA
    23 hours ago
  • $80 per hour

    A tech company specializing in AI is seeking QAs for autonomous AI agents to validate and enhance task structures...  ...allows contributors to work remotely while engaging in a complex AI project...  ...detail, facilitating AI testing and evaluation without needing a coding... 
    Remote job
    Flexible hours

    Mind Rift

    Houston, TX
    1 day ago
  • YO AI Labs in the United States is seeking experienced Civil Engineers to contribute to a high-impact AI evaluation project. You will apply your engineering expertise to create realistic scenarios, develop accurate solutions, and evaluate AI-generated engineering outputs... 
    Remote job
    For contractors

    YO AI Labs

    Annapolis, MD
    23 hours ago
  • YO AI Labs is seeking experienced Civil Engineers to contribute to a high-impact project focused on advancing AI evaluation. You will apply your engineering expertise to create realistic technical...  ...with rapid onboarding and flexible remote collaboration, enabling experts to... 
    Remote job
    Flexible hours

    YO AI Labs

    Houston, TX
    23 hours ago
  • YO AI Labs is seeking experienced Civil Engineers for remote, contractor-type work focused on AI evaluation. You will craft realistic civil engineering tasks, prepare technical materials, and assess code compliance across design domains. The role emphasizes constraint... 
    Remote job
    For contractors

    YO AI Labs

    San Francisco, CA
    4 days ago
  • $60 per hour

    A leading AI firm in Austin is looking for QA experts to validate...  ...and improve AI systems. This remote, freelance role requires...  ...detail. Candidates will review AI evaluation tasks, identify inconsistencies...  ...define expected behaviors for agents. Ideal applicants have experience... 
    Remote job
    Freelance

    Mindrift

    Austin, TX
    3 days ago
  • YO AI Labs is seeking experienced Civil Engineers to contribute to a high-impact AI evaluation project. You will apply civil engineering expertise to create realistic evaluation tasks, develop technical materials, and assess AI-generated outputs for accuracy and code compliance... 
    Remote job

    YO AI Labs

    Chicago, IL
    23 hours ago
  •  ...Role Overview Shape evaluation tasks for AI systems used in enterprise electrical systems and product...  ...Create electrical engineering scenarios covering complex circuit and...  ...design documentation. Work Terms ~ Remote, hourly engagement. Compensation... 
    Remote work
    Hourly pay

    SaidGig

    United States
    2 days ago
  • $80 per hour

     ...ethically shape the future of AI. Our platform connects...  ...realistic and structured evaluation scenarios for LLM-based agents. You'll create test cases...  ...Science, Software Engineering, Data Science / Data Analytics...  ...Take part in a flexible, remote, freelance project that... 
    Remote work
    Part time
    Freelance
    Flexible hours

    Mindrift

    Houston, TX
    3 days ago
  •  ...Cohere's cutting-edge AI workspace platform, designed...  ...that connects AI agents with workplace tools and...  ...You’ll work with product engineering, modelling and design...  ...users create, configure, evaluate, and improve workflows...  ...other offices if you are remote, plus an annual company... 
    Remote work
    Full time
    Local area

    Cohere

    Remote
    a month ago
  • $55 per hour

    A leading AI innovation firm in Dallas is seeking QAs for autonomous AI agents to ensure the quality of complex systems and scenarios. This flexible, remote project is ideal for those with excellent analytical...  ...should be adept at evaluating scenarios and documenting findings... 
    Remote job
    Flexible hours

    Mindrift

    Dallas, TX
    2 days ago
  • $70 - $80 per hour

     ...Role Overview Shape high-stakes AI evaluation tasks grounded in enterprise product design...  .... You will create realistic mechanical engineering scenarios, reference outputs, and assessment...  ...is preferred. Work Terms Remote, hourly engagement. Experience with... 
    Remote work
    Hourly pay

    SaidGig

    United States
    2 days ago
  •  ...Tyndale is expanding our AI and automation capabilities...  ...seeking an experienced AI Agent & Automation Engineer to build production solutions...  ...support.   HYBRID/REMOTE: Tyndale supports a strong...  ...platforms, AI agent testing/evaluation, containerization and operational... 
    Remote work
    Casual work
    1 day per week

    Tyndale

    Pipersville, PA
    9 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Agent Evaluation Engineer - Remote. Be the first to apply!